# https://voxel51.com llms-full.txt <|firecrawl-page-1-lllmstxt|> ## Visual AI Data Platform [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # Maximize AI performance with better data FiftyOne is the most powerful Visual AI and computer vision data platform. [Book a demo](https://voxel51.com/sales) [Access dev docs](https://docs.voxel51.com/) ![](https://cdn.sanity.io/images/h6toihm1/production/243b047fc84839f1fa1b585aa5557fb1189c2d6b-1704x1780.png?auto=format&dpr=2&fit=crop&fp-x=0.453&fp-y=0.641&h=433&q=75&rect=0,0,1704,1780&w=375) ![](https://cdn.sanity.io/images/h6toihm1/production/d68cfc6eab9b59bc59005d8e5b6ae6ba1f81405d-1682x1600.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=237&q=75&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/b35accefa5d99d017dbe4ed7c55c28535fb4230c-1600x1234.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=192&q=75&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/66b60f27821149efb7b7c3adaf0d0c59bc46e216-1334x1564.png?auto=format&dpr=2&fit=crop&fp-x=0.629&fp-y=0.695&h=433&q=75&rect=0,0,1334,1564&w=375) ![](https://cdn.sanity.io/images/h6toihm1/production/eaa44eeb403349df4caffcf9efb3d110b98a6119-1770x2062.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.397&h=237&q=75&rect=0,0,1770,2062&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/529e71d39bd8862a79d22ee44310cc8347fede5f-1140x944.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=237&q=75&rect=0,30,1140,892&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/58c472bddabbd00630cb25bc83409fa4e9dc48da-1140x944.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.497&h=192&q=75&rect=0,0,1140,944&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/243b047fc84839f1fa1b585aa5557fb1189c2d6b-1704x1780.png?auto=format&dpr=2&fit=crop&fp-x=0.453&fp-y=0.641&h=433&q=75&rect=0,0,1704,1780&w=375) ![](https://cdn.sanity.io/images/h6toihm1/production/d68cfc6eab9b59bc59005d8e5b6ae6ba1f81405d-1682x1600.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=237&q=75&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/b35accefa5d99d017dbe4ed7c55c28535fb4230c-1600x1234.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=192&q=75&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/66b60f27821149efb7b7c3adaf0d0c59bc46e216-1334x1564.png?auto=format&dpr=2&fit=crop&fp-x=0.629&fp-y=0.695&h=433&q=75&rect=0,0,1334,1564&w=375) ![](https://cdn.sanity.io/images/h6toihm1/production/eaa44eeb403349df4caffcf9efb3d110b98a6119-1770x2062.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.397&h=237&q=75&rect=0,0,1770,2062&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/529e71d39bd8862a79d22ee44310cc8347fede5f-1140x944.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=237&q=75&rect=0,30,1140,892&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/58c472bddabbd00630cb25bc83409fa4e9dc48da-1140x944.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.497&h=192&q=75&rect=0,0,1140,944&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/243b047fc84839f1fa1b585aa5557fb1189c2d6b-1704x1780.png?auto=format&dpr=2&fit=crop&fp-x=0.453&fp-y=0.641&h=433&q=75&rect=0,0,1704,1780&w=375) ![](https://cdn.sanity.io/images/h6toihm1/production/d68cfc6eab9b59bc59005d8e5b6ae6ba1f81405d-1682x1600.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=237&q=75&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/b35accefa5d99d017dbe4ed7c55c28535fb4230c-1600x1234.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=192&q=75&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/66b60f27821149efb7b7c3adaf0d0c59bc46e216-1334x1564.png?auto=format&dpr=2&fit=crop&fp-x=0.629&fp-y=0.695&h=433&q=75&rect=0,0,1334,1564&w=375) ![](https://cdn.sanity.io/images/h6toihm1/production/eaa44eeb403349df4caffcf9efb3d110b98a6119-1770x2062.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.397&h=237&q=75&rect=0,0,1770,2062&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/529e71d39bd8862a79d22ee44310cc8347fede5f-1140x944.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=237&q=75&rect=0,30,1140,892&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/58c472bddabbd00630cb25bc83409fa4e9dc48da-1140x944.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.497&h=192&q=75&rect=0,0,1140,944&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/243b047fc84839f1fa1b585aa5557fb1189c2d6b-1704x1780.png?auto=format&dpr=2&fit=crop&fp-x=0.453&fp-y=0.641&h=433&q=75&rect=0,0,1704,1780&w=375) ![](https://cdn.sanity.io/images/h6toihm1/production/d68cfc6eab9b59bc59005d8e5b6ae6ba1f81405d-1682x1600.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=237&q=75&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/b35accefa5d99d017dbe4ed7c55c28535fb4230c-1600x1234.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=192&q=75&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/66b60f27821149efb7b7c3adaf0d0c59bc46e216-1334x1564.png?auto=format&dpr=2&fit=crop&fp-x=0.629&fp-y=0.695&h=433&q=75&rect=0,0,1334,1564&w=375) ![](https://cdn.sanity.io/images/h6toihm1/production/eaa44eeb403349df4caffcf9efb3d110b98a6119-1770x2062.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.397&h=237&q=75&rect=0,0,1770,2062&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/529e71d39bd8862a79d22ee44310cc8347fede5f-1140x944.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=237&q=75&rect=0,30,1140,892&w=249) ![](https://cdn.sanity.io/images/h6toihm1/production/58c472bddabbd00630cb25bc83409fa4e9dc48da-1140x944.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.497&h=192&q=75&rect=0,0,1140,944&w=249) ![Walmart is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/55608739e687dbd7b498b77c13136e2b9ed611a3-304x72.png?auto=format&dpr=2&fit=max&q=75&w=120) ![](https://cdn.sanity.io/images/h6toihm1/production/4aea181cb46420e25282c4cb8f58c590e33053ba-250x158.png?auto=format&dpr=2&fit=max&q=75&w=120) ![University of Michigan is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/c50a1d5574718bbcec04edd8b3573e06e2e7a8f6-192x120.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,192,106&w=120) ![NVIDIA is a Voxel51 partner](https://cdn.sanity.io/images/h6toihm1/production/65e5748bfccdc518f0c0553056775e2e73165202-317x60.png?auto=format&dpr=2&fit=max&q=75&w=120) ![](https://cdn.sanity.io/images/h6toihm1/production/99639f3cbc1a2d4373a2c928e4d7fd09027ede00-1660x115.png?auto=format&dpr=2&fit=max&q=75&w=120) ![](https://cdn.sanity.io/images/h6toihm1/production/4a8d9cb4f0bfaa910010c81385db0ce9ba38d33e-360x180.png?auto=format&dpr=2&fit=max&q=75&rect=44,28,272,125&w=120) ![GM is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/f5359c53fcc5a8686e2a0ce71daa9862b4825ccd-102x102.png?auto=format&dpr=2&fit=max&q=75&w=102) ![Medtronic is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/a726afe74f6fef37caff91a77c0c59c4c1b554ab-280x48.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Bosch is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/18af41d884df95c147bfc7481a8b012380383875-310x66.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Microsoft is a Voxel51 partner](https://cdn.sanity.io/images/h6toihm1/production/40e7d53a7e32d74092a36210c35f19b4800cadaa-310x66.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Safely You is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/c366c172654df3da16b39ea6a408a55219aea57a-316x72.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Raytheon is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/13a9402984dd0a566cf65f7ea64c9164f9973970-300x90.png?auto=format&dpr=2&fit=max&q=75&w=120) ![](https://cdn.sanity.io/images/h6toihm1/production/174595171c0e067cdfc0a2704181a289db6e7060-702x300.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Google is a Voxel51 partner](https://cdn.sanity.io/images/h6toihm1/production/e70d344bc09d2cd1fb85ed80e39baadccb0a7f60-270x180.png?auto=format&dpr=2&fit=max&q=75&rect=19,53,230,75&w=120) ![RIOS is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/eb6156f36de826b6d65dbb53876c53b961530543-256x102.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Berkshire Grey is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/bdfe35dc56d9ba21e83468ed3a42a352f6f2884f-330x78.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Ford is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/0554cbeb43e083b5436bc2c99b74c30e73783fea-250x90.png?auto=format&dpr=2&fit=max&q=75&w=120) ![ArcelorMittal is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/2d8c0ac21b002de5c952f0889fd48acad327c3d1-294x150.png?auto=format&dpr=2&fit=max&q=75&rect=16,87,231,46&w=120) ![LG is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/b11dab534591c8e12e8af89ebe588e8d5b66ef22-196x90.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Sony is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/f2d5501398d625863e53957f3d87b3d74c03d245-311x54.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Meta used Voxel51 in the development of SAM2](https://cdn.sanity.io/images/h6toihm1/production/4a38deb361b13f2b8696b696068001298a2e00aa-299x60.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Walmart is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/55608739e687dbd7b498b77c13136e2b9ed611a3-304x72.png?auto=format&dpr=2&fit=max&q=75&w=120) ![](https://cdn.sanity.io/images/h6toihm1/production/4aea181cb46420e25282c4cb8f58c590e33053ba-250x158.png?auto=format&dpr=2&fit=max&q=75&w=120) ![University of Michigan is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/c50a1d5574718bbcec04edd8b3573e06e2e7a8f6-192x120.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,192,106&w=120) ![NVIDIA is a Voxel51 partner](https://cdn.sanity.io/images/h6toihm1/production/65e5748bfccdc518f0c0553056775e2e73165202-317x60.png?auto=format&dpr=2&fit=max&q=75&w=120) ![](https://cdn.sanity.io/images/h6toihm1/production/99639f3cbc1a2d4373a2c928e4d7fd09027ede00-1660x115.png?auto=format&dpr=2&fit=max&q=75&w=120) ![](https://cdn.sanity.io/images/h6toihm1/production/4a8d9cb4f0bfaa910010c81385db0ce9ba38d33e-360x180.png?auto=format&dpr=2&fit=max&q=75&rect=44,28,272,125&w=120) ![GM is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/f5359c53fcc5a8686e2a0ce71daa9862b4825ccd-102x102.png?auto=format&dpr=2&fit=max&q=75&w=102) ![Medtronic is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/a726afe74f6fef37caff91a77c0c59c4c1b554ab-280x48.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Bosch is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/18af41d884df95c147bfc7481a8b012380383875-310x66.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Microsoft is a Voxel51 partner](https://cdn.sanity.io/images/h6toihm1/production/40e7d53a7e32d74092a36210c35f19b4800cadaa-310x66.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Safely You is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/c366c172654df3da16b39ea6a408a55219aea57a-316x72.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Raytheon is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/13a9402984dd0a566cf65f7ea64c9164f9973970-300x90.png?auto=format&dpr=2&fit=max&q=75&w=120) ![](https://cdn.sanity.io/images/h6toihm1/production/174595171c0e067cdfc0a2704181a289db6e7060-702x300.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Google is a Voxel51 partner](https://cdn.sanity.io/images/h6toihm1/production/e70d344bc09d2cd1fb85ed80e39baadccb0a7f60-270x180.png?auto=format&dpr=2&fit=max&q=75&rect=19,53,230,75&w=120) ![RIOS is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/eb6156f36de826b6d65dbb53876c53b961530543-256x102.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Berkshire Grey is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/bdfe35dc56d9ba21e83468ed3a42a352f6f2884f-330x78.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Ford is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/0554cbeb43e083b5436bc2c99b74c30e73783fea-250x90.png?auto=format&dpr=2&fit=max&q=75&w=120) ![ArcelorMittal is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/2d8c0ac21b002de5c952f0889fd48acad327c3d1-294x150.png?auto=format&dpr=2&fit=max&q=75&rect=16,87,231,46&w=120) ![LG is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/b11dab534591c8e12e8af89ebe588e8d5b66ef22-196x90.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Sony is a Voxel51 customer](https://cdn.sanity.io/images/h6toihm1/production/f2d5501398d625863e53957f3d87b3d74c03d245-311x54.png?auto=format&dpr=2&fit=max&q=75&w=120) ![Meta used Voxel51 in the development of SAM2](https://cdn.sanity.io/images/h6toihm1/production/4a38deb361b13f2b8696b696068001298a2e00aa-299x60.png?auto=format&dpr=2&fit=max&q=75&w=120) ML Workflows ## Unlock the value of your data In a world where data fuels AI innovation, FiftyOne puts data at the center of your workflow—helping you exploit its full potential to gain a competitive edge. [Read the whitepaper](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation) ![](https://cdn.sanity.io/images/h6toihm1/production/a4132ca20a189c49a09bf2f22b8fd92b896bbd18-2560x1200.png?auto=format&dpr=2&fit=max&q=75&w=1280) [Data Curation & Management](https://voxel51.com/curation) [Smarter Annotation](https://voxel51.com/annotation) [Model Evaluation](https://voxel51.com/evaluation) Benefits & ROI ## Leading enterprises build using FiftyOne 0% increase in model accuracy 0+ months of development time saved 0% boost in team productivity ![](https://cdn.sanity.io/images/h6toihm1/production/cd7fe79ef465aa75489f4c049ce2af346f0985b4-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/720f6702611411baf6a274b1010856c78c07235d-240x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/57564abbae97325ddeb3bc3b8b7b5690d61cddc5-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/be26ca6233b020b9201ddd482f7d34f53555b51d-276x90.png?auto=format&dpr=2&fit=max&q=75&w=100)](https://voxel51.com/customers/safelyyou) Maintained 99% fall detection rates for model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/636eaa49ef62ea3d07ca264f7baa4fe0e494e125-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/3911d157da27a0e996468bf45393d4e01e522934-384x39.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/a54f5857074a8e45684334b38c3e4f1f1528274e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/e845660699edf4deed110e60425659ba8ae73e6e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Foundation for Florence-2 VLM development](https://voxel51.com/plugins/?search=florence) ![](https://cdn.sanity.io/images/h6toihm1/production/c8be8bb13c2be64cd70dda28334501a812eb0bc1-520x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/8c596efcf17de5a9bc02b05b44f55474802f0abe-90x90.png?auto=format&dpr=2&fit=max&q=75&w=90) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cb778b338fac262a1c8a9de0334b256747cfbc5b-500x261.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Eliminated repetitive manual transformations on 20 TB+ of visual data](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/1f4c7beab6c76544152d4f3d990059bc62eb90cb-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/aee0fde76666dd2d134875eb8d5247ee75dab896-153x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/32f1bd14467c47df8f6311bb1304f179c3e184be-1024x614.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/30b0acc56eeba0fd43bf86f4a390dd985d1916cd-640x152.png?auto=format&dpr=2&fit=max&q=75&w=100)](https://voxel51.com/customers/berkshire-grey) Sped up investigations of robotic arms by 3x. ![](https://cdn.sanity.io/images/h6toihm1/production/2eed779bf08ffa8ea43cb822cb013674cc05e546-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/6f935624676c89371a1b094385b7b60439e09824-351x60.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/09890f00ed8e1cb9be4a30a0a0a52a80b2d74fee-800x1220.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/249883257dad5df54dbe20949b7c6954032e0f2c-640x361.png?auto=format&dpr=2&fit=max&q=75&rect=32,159,571,56&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg)\\ \\ Official partner for visualizing Open Images Dataset V7](https://voxel51.com/blog/exploring-google-open-images-v7/) ![](https://cdn.sanity.io/images/h6toihm1/production/c5ba8bdc3b7afa8cb9d7a3e1d00aa0d82bf592de-1638x2048.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0e2f149ba6e28a4e307e55a8ad52f81f790571df-921x96.svg) ![](https://cdn.sanity.io/images/h6toihm1/production/cd7fe79ef465aa75489f4c049ce2af346f0985b4-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/720f6702611411baf6a274b1010856c78c07235d-240x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/57564abbae97325ddeb3bc3b8b7b5690d61cddc5-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/be26ca6233b020b9201ddd482f7d34f53555b51d-276x90.png?auto=format&dpr=2&fit=max&q=75&w=100)](https://voxel51.com/customers/safelyyou) Maintained 99% fall detection rates for model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/636eaa49ef62ea3d07ca264f7baa4fe0e494e125-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/3911d157da27a0e996468bf45393d4e01e522934-384x39.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/a54f5857074a8e45684334b38c3e4f1f1528274e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/e845660699edf4deed110e60425659ba8ae73e6e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Foundation for Florence-2 VLM development](https://voxel51.com/plugins/?search=florence) ![](https://cdn.sanity.io/images/h6toihm1/production/c8be8bb13c2be64cd70dda28334501a812eb0bc1-520x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/8c596efcf17de5a9bc02b05b44f55474802f0abe-90x90.png?auto=format&dpr=2&fit=max&q=75&w=90) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cb778b338fac262a1c8a9de0334b256747cfbc5b-500x261.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Eliminated repetitive manual transformations on 20 TB+ of visual data](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/1f4c7beab6c76544152d4f3d990059bc62eb90cb-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/aee0fde76666dd2d134875eb8d5247ee75dab896-153x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/32f1bd14467c47df8f6311bb1304f179c3e184be-1024x614.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/30b0acc56eeba0fd43bf86f4a390dd985d1916cd-640x152.png?auto=format&dpr=2&fit=max&q=75&w=100)](https://voxel51.com/customers/berkshire-grey) Sped up investigations of robotic arms by 3x. ![](https://cdn.sanity.io/images/h6toihm1/production/2eed779bf08ffa8ea43cb822cb013674cc05e546-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/6f935624676c89371a1b094385b7b60439e09824-351x60.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/09890f00ed8e1cb9be4a30a0a0a52a80b2d74fee-800x1220.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/249883257dad5df54dbe20949b7c6954032e0f2c-640x361.png?auto=format&dpr=2&fit=max&q=75&rect=32,159,571,56&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg)\\ \\ Official partner for visualizing Open Images Dataset V7](https://voxel51.com/blog/exploring-google-open-images-v7/) ![](https://cdn.sanity.io/images/h6toihm1/production/c5ba8bdc3b7afa8cb9d7a3e1d00aa0d82bf592de-1638x2048.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0e2f149ba6e28a4e307e55a8ad52f81f790571df-921x96.svg) Annotation ## Reduce annotation costs by 100,000x with Verified Auto Labeling Get AI-assisted labeling with built-in confidence scoring to prioritize labels for human review — reducing manual QA and accelerating annotation workflows. [Explore Verified Auto Labeling](https://voxel51.com/annotation) DATA CURATION & MANAGEMENT ## Visualize your data like never before Identify your best performing samples, weed out low-quality data, and organize dataset views intuitively — so you can focus on building better models, faster. [Explore data curation](https://voxel51.com/curation) Unify multimodal data Slice massive datasetsAnalyze data patternsImprove data qualityStreamline data explorationVersion datasets ### Multimodal data support Work seamlessly across data types with a unified interface. Images Video 3D models, point cloud data LIDAR, radar, GPS, sensor and geospatial data DICOM, CT scans, X-ray, infrared, NPY, SAR Audio [Learn more](https://docs.voxel51.com/user_guide/groups.html) ### Manage massive datasets Filter, query, sort, and slice across billions of samples to make sense of complex datasets quickly. [Learn more](https://docs.voxel51.com/user_guide/using_views.html) ### Analyze data patterns Visualize class imbalances, annotation gaps, and clusters using embeddings, histograms, and heatmap visualizations. [Visualize embeddings](https://voxel51.com/resources/learn/how-image-embeddings-transform-computer-vision-capabilities/) ![](https://cdn.sanity.io/images/h6toihm1/production/8ad693a738bd2157a1c30788e71881053c98eef2-1257x708.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1257,708&w=600) ### Improve data quality Identify edge cases, outliers, duplicates, and mislabeled samples with precision filtering and dynamic slices—so you can build cleaner, leaner datasets that drive better model results [Learn more](https://voxel51.com/blog/build-better-visual-ai-datasets-with-the-fiftyone-data-quality-workflow/) ![](https://cdn.sanity.io/images/h6toihm1/production/29c60b6b46b9df3e38fe4fef378a2ae29bebfe8c-1257x708.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1257,708&w=600) ### Streamline data exploration Query your data lake and retrieve relevant samples in seconds. Search across metadata, predictions, and more—without needing to load everything into memory. [Learn more](https://voxel51.com/blog/streamline-visual-data-discovery-with-fiftyone-data-lens/) ### Robust dataset versioning Snapshot your dataset at key moments—like training runs or annotation updates—to simplify reproducibility and protect against accidental data loss. [Learn more](https://docs.voxel51.com/enterprise/dataset_versioning.html?highlight=versioning) MODEL EVALUATION ## Evaluate data and models side-by-side Assess overall model performance and inspect individual samples — all in one interactive workflow. Identify failure modes, bias, and blind spots with ease. [Explore model evaluation](https://voxel51.com/evaluation) ![](https://cdn.sanity.io/images/h6toihm1/production/1b844e20fcd3fe215c76bb5f9901c5d5f119c453-3542x1661.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,3542,1661&w=1771) #### Scenario analysis Compare models across key metrics and specific data slices to quickly spot performance differences, edge-case failures, and areas for targeted improvement. #### Sample-level analysis Analyze model predictions on individual data samples to reveal hidden errors and critical insights to guide targeted model improvements. COMPLIANCE & GOVERNANCE ## Enterprise-grade security, scale, and extensibility FiftyOne is built to meet the demands of the most sophisticated ML stacks. Deploy anywhere ![](https://cdn.sanity.io/images/h6toihm1/production/79fa40b25cb38d7c9c85b3f565dcda26356a4f15-1276x634.png?auto=format&dpr=2&fit=max&q=75&w=640) Fully customizable and extensible ![](https://cdn.sanity.io/images/h6toihm1/production/c5d011a5d71e23096ce07385b12ddebe5e66cb42-1268x634.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1268,634&w=640) Support for billions of samples ![FiftyOne supports billions of samples and metadata](https://cdn.sanity.io/images/h6toihm1/production/9c891a7d08584ac04cb8f033bae7bc7c3941778b-951x1152.png?auto=format&dpr=2&fit=max&q=75&w=317) Dataset versioning ![FiftyOne includes robust data versioning capabilities](https://cdn.sanity.io/images/h6toihm1/production/e9f5d5d86653a3ccf36dc923a034c091cdf99ff1-951x1152.png?auto=format&dpr=2&fit=max&q=75&w=317) Role-based access controls ![](https://cdn.sanity.io/images/h6toihm1/production/ad606a31d0b9b15441e8528aa3d8836e307083ee-951x1152.png?auto=format&dpr=2&fit=max&q=75&w=317) ISO 27001 certification ![FiftyOne is ISO 27001 certified](https://cdn.sanity.io/images/h6toihm1/production/c498202dd43ae6504021355074de1d97950b5daa-951x1152.png?auto=format&dpr=2&fit=max&q=75&w=317) INTEGRATIONS ## Integrate with your existing ML stack [Explore integrations](https://voxel51.com/integrations) ![](https://cdn.sanity.io/images/h6toihm1/production/4262375676fd146768064041982c53f191063e6d-2560x1200.jpg?auto=format&dpr=2&fit=max&q=75&w=1280) Developer resources ## Loved by ML engineers ### More than 3 million installs [Access dev docs](https://docs.voxel51.com/) ![](https://cdn.sanity.io/images/h6toihm1/production/7abe956c642b40cc310bbeb3e8a2f704cbf88328-1276x754.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1276,754&w=640) ![](https://cdn.sanity.io/images/h6toihm1/production/dbacd2e2d7a316c4763ba21c03c0ef5b5651d35f-638x282.jpg?auto=format&dpr=2&fit=max&q=75&w=638) ### Built on open source standards [View on GitHub](https://github.com/voxel51/) ![](https://cdn.sanity.io/images/h6toihm1/production/dbacd2e2d7a316c4763ba21c03c0ef5b5651d35f-638x282.jpg?auto=format&dpr=2&fit=max&q=75&w=638) ![](https://cdn.sanity.io/images/h6toihm1/production/60fa23477f02dbee187d052370aa4905d4eb8731-606x548.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=350&q=75&rect=36,0,570,548&w=350) ### 22K+ computer vision community members [Join the community](https://voxel51.com/events) ![Voxel51 hosts one of the largest computer vision communities in the world](https://cdn.sanity.io/images/h6toihm1/production/527003753c0bce5b582f54c0733d7aa0da8d5f90-1005x894.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=350&q=75&w=350) ## Enough data wrangling.
 Request a demo. [Book a demo](https://voxel51.com/sales) [Unlock the value of data](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation) ![](https://cdn.sanity.io/images/h6toihm1/production/ea42e9b26f49f1cb54bb8aca31dc10e7f74fe11f-3024x960.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-2-lllmstxt|> ## Join Voxel51 Team [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # Voxel51Careers At Voxel51, we’re not just building technology—we’re redefining it. Our team thrives on intellectual curiosity, embraces bold challenges, and pushes the boundaries of what’s possible in computer vision and visual AI. If you’re driven to revolutionize the next generation of AI solutions and want to work alongside passionate innovators, your journey starts here. [Apply for open roles](https://voxel51.com/careers#open-positions) [About Voxel51](https://voxel51.com/about) ![](https://cdn.sanity.io/images/h6toihm1/production/061b90b942ecd6716bcfe06c097ba9039c7ade19-600x546.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) We meet twice a year in person for company offsites. ![](https://cdn.sanity.io/images/h6toihm1/production/456cd43a68b25ca60c89240765f23d2a84576c5f-768x788.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)\\ \\ 3M open source installs and 250+ enterprise customers](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/c8a82cb55270ec0dc68960fec331ca2136ad7c6e-768x594.jpg?auto=format&dpr=2&fit=crop&fp-x=0.469&fp-y=0.5&h=340&q=75&rect=29,0,739,594&w=260) Founded by PhDs at the University of Michigan. ![](https://cdn.sanity.io/images/h6toihm1/production/638e0e67af4c0b3b69b21657927857a1d3253ccb-768x576.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) ![](https://cdn.sanity.io/images/h6toihm1/production/8abeb200a76a49defd2366b12de7dd6eeed896e6-768x576.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=321,0,447,576&w=260) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)\\ \\ Series B with $45M in total funding](https://voxel51.com/blog/exploring-google-open-images-v7/) Backed by Bessemer, Drive Capital, Tru Arrow, Top Harvest, Shasta Ventures, and ID Ventures ![](https://cdn.sanity.io/images/h6toihm1/production/6aff7fb4ed058ed5dff92b5b712c8fc84508b25e-1512x2016.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) ![](https://cdn.sanity.io/images/h6toihm1/production/061b90b942ecd6716bcfe06c097ba9039c7ade19-600x546.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) We meet twice a year in person for company offsites. ![](https://cdn.sanity.io/images/h6toihm1/production/456cd43a68b25ca60c89240765f23d2a84576c5f-768x788.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)\\ \\ 3M open source installs and 250+ enterprise customers](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/c8a82cb55270ec0dc68960fec331ca2136ad7c6e-768x594.jpg?auto=format&dpr=2&fit=crop&fp-x=0.469&fp-y=0.5&h=340&q=75&rect=29,0,739,594&w=260) Founded by PhDs at the University of Michigan. ![](https://cdn.sanity.io/images/h6toihm1/production/638e0e67af4c0b3b69b21657927857a1d3253ccb-768x576.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) ![](https://cdn.sanity.io/images/h6toihm1/production/8abeb200a76a49defd2366b12de7dd6eeed896e6-768x576.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=321,0,447,576&w=260) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)\\ \\ Series B with $45M in total funding](https://voxel51.com/blog/exploring-google-open-images-v7/) Backed by Bessemer, Drive Capital, Tru Arrow, Top Harvest, Shasta Ventures, and ID Ventures ![](https://cdn.sanity.io/images/h6toihm1/production/6aff7fb4ed058ed5dff92b5b712c8fc84508b25e-1512x2016.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) ![](https://cdn.sanity.io/images/h6toihm1/production/061b90b942ecd6716bcfe06c097ba9039c7ade19-600x546.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) We meet twice a year in person for company offsites. ![](https://cdn.sanity.io/images/h6toihm1/production/456cd43a68b25ca60c89240765f23d2a84576c5f-768x788.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)\\ \\ 3M open source installs and 250+ enterprise customers](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/c8a82cb55270ec0dc68960fec331ca2136ad7c6e-768x594.jpg?auto=format&dpr=2&fit=crop&fp-x=0.469&fp-y=0.5&h=340&q=75&rect=29,0,739,594&w=260) Founded by PhDs at the University of Michigan. ![](https://cdn.sanity.io/images/h6toihm1/production/638e0e67af4c0b3b69b21657927857a1d3253ccb-768x576.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) ![](https://cdn.sanity.io/images/h6toihm1/production/8abeb200a76a49defd2366b12de7dd6eeed896e6-768x576.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=321,0,447,576&w=260) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)\\ \\ Series B with $45M in total funding](https://voxel51.com/blog/exploring-google-open-images-v7/) Backed by Bessemer, Drive Capital, Tru Arrow, Top Harvest, Shasta Ventures, and ID Ventures ![](https://cdn.sanity.io/images/h6toihm1/production/6aff7fb4ed058ed5dff92b5b712c8fc84508b25e-1512x2016.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260) Why Voxel51 ## Our commitment to the Voxel51 team $0M total funding 0M FiftyOne OSS downloads 0+ enterprise customers #### Mission driven We’re on a mission to bring transparency and clarity to the world’s data. #### Open source We believe in the power of open source and community to revolutionize data-centric AI/ML. #### Human first We treat everyone—our community, customers, and team—with the respect, care, and flexibility that all people deserve #### Remote first We embrace distributed work while maintaining connection through virtual and in-person team events., normal, theme-default-raised. #### Benefits We offer medical, dental, and vision benefit plans, as well as generous PTO, flexible work schedules, 401k plans, and more #### Learn and grow Work alongside others who love what they do, share ideas, ask questions, and grow your skills. Jobs at Voxel51 ![Voxel51 Logo](https://s5-recruiting.cdn.greenhouse.io/external_greenhouse_job_boards/logos/400/085/700/original/1x1__450da6ff.png?1646881346) # Current openings at Voxel51 Create a Job Alert Level-up your career by having opportunities at Voxel51 sent directly to your inbox. [Create alert](https://my.greenhouse.io/users/sign_in?job_board=voxel51&source=job_alert) Search ## 7 jobs ### Main Staff | Job | | --- | | [Machine Learning Customer Success Engineer\
\
Remote](https://voxel51.com/jd/?4367259005&gh_jid=4367259005) | | [Machine Learning Engineer\
\
Remote](https://voxel51.com/jd/?4550368005&gh_jid=4550368005) | | [Marketing Ops SpecialistNew\
\
Remote](https://voxel51.com/jd/?4598945005&gh_jid=4598945005) | | [Principal Infrastructure EngineerNew\
\
Remote](https://voxel51.com/jd/?4599511005&gh_jid=4599511005) | | [Principal Software Engineer (Frontend)\
\
Remote](https://voxel51.com/jd/?4559955005&gh_jid=4559955005) | | [Sales Development Representative\
\
United States](https://voxel51.com/jd/?4544024005&gh_jid=4544024005) | | [Senior Software Engineer (Visualization)\
\
Remote](https://voxel51.com/jd/?4593137005&gh_jid=4593137005) | ## We unlock breakthroughs in visual AI and transform industries from automotive to robotics ![](https://cdn.sanity.io/images/h6toihm1/production/22b434030668d87caf6a19635da5cab49dfe08d4-2048x1365.jpg?auto=format&dpr=2&fit=max&q=75&w=600) ### Our founding story Voxel51 was founded in 2018 at the University of Michigan when professor Jason Corso teamed up with his PhD Brian Moore to turn their research into developer-friendly tooling for computer-vision data. The explosive adoption of FiftyOne has made us the go-to platform for ML teams who work with multimodal data at scale. ![](https://cdn.sanity.io/images/h6toihm1/production/282a93d2b293247c11ae9f2b05bf0084ce5178d4-768x620.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,768,620&w=384) ![](https://cdn.sanity.io/images/h6toihm1/production/a7a1085305de96e5ec83621e54a23dfed5cb725c-2560x640.png?auto=format&dpr=2&fit=max&q=75&w=1280) Customer Testimonials ## Developers love FiftyOne > “From vehicle safety and autonomy to security systems to robotics, Bosch is a leader in artificial intelligence solutions utilizing computer vision. Voxel51’s solutions help us organize, evaluate and refine our data and models, enabling us to develop robust, reliable AI applications across multiple teams and projects. ” > > **Arvind Kumar Shekar** > > Lead Expert AI Validation, Bosch ![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) > “The biggest benefit of FiftyOne has been the speed of development. What used to take weeks or even months can now be done in days, with fewer people and a 7% increase in model performance. It’s freed up our team to focus on what they do best while accelerating our computer vision pipeline.”” > > **Kermal Eren** > > Lead Computer Vision Engineer, Ancera ![](https://cdn.sanity.io/images/h6toihm1/production/17832cc5c128034561711a5d43942252792462b4-479x101.webp?auto=format&dpr=2&fit=max&q=75&rect=0,4,479,95&w=100) > “As we dive into the development of Florence-5B, we’re relying on FiftyOne more than ever. The tool’s intuitive interface and rich feature set are essential for effectively managing our large datasets and gaining critical insights. ” > > **Bin Xiao** > > AI Researcher, Florence Visual Language Model ![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/17832cc5c128034561711a5d43942252792462b4-479x101.webp?auto=format&dpr=2&fit=max&q=75&rect=0,4,479,95&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-3-lllmstxt|> ## Evaluating Multimodal AI [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Rethinking How We Evaluate Multimodal AI Jun 12, 2025 • 16 min read Article content In this article [CVPR 2025 reveals why spatial reasoning and subjective ‘vibes’ are redefining how we benchmark AI systems](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#6fb465ae6a3b) [Andre Araujo: Multimodal AI is Amazing... Yet Deeply Flawed](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#3022e134293c) [TIPS: Engineering Spatial Understanding](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#2fe347b23e02) [Saining Xie: Language Shortcuts Undermine Visual Intelligence](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#69039529203c) [Lisa Dunlap: The Problem with Single-Number Leaderboards](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#3d00da618f5d) [The Path Forward](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#204fd62a31a3) In this article [CVPR 2025 reveals why spatial reasoning and subjective ‘vibes’ are redefining how we benchmark AI systems](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#6fb465ae6a3b) [Andre Araujo: Multimodal AI is Amazing... Yet Deeply Flawed](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#3022e134293c) [TIPS: Engineering Spatial Understanding](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#2fe347b23e02) [Saining Xie: Language Shortcuts Undermine Visual Intelligence](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#69039529203c) [Lisa Dunlap: The Problem with Single-Number Leaderboards](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#3d00da618f5d) [The Path Forward](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai#204fd62a31a3) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## CVPR 2025 reveals why spatial reasoning and subjective ‘vibes’ are redefining how we benchmark AI systems At CVPR 2025, I spent my first day attending talks on benchmarking and evaluation of multimodal AI systems. Despite the impressive capabilities showcased throughout the conference, these sessions revealed critical gaps between how we evaluate these models and their actual performance in real-world scenarios. What emerged was a fascinating narrative about the disconnect between public perception and technical reality. Our current multimodal models — despite their seemingly magical abilities — struggle with spatial reasoning tasks that toddlers master effortlessly. Meanwhile, our evaluation systems often reward the wrong things: verbose responses over accurate ones, language shortcuts over visual understanding, and single metrics over nuanced capabilities. Three speakers particularly stood out, each touching on different dimensions of evaluation: - [Andre Araujo’s](https://andrefaraujo.github.io/) gave a comprehensive overview of the significant progress in multimodal AI, highlighting both its “magical” capabilities and critical limitations, while proposing innovative solutions for spatial awareness, effective tool use, and fine-grain understanding. - [Saining Xie](https://www.sainingxie.com/) addressed the challenges of benchmarking real intelligence, emphasizing the importance of visual-spatial intelligence in conjunction with linguistic skills. He notes advancements in self-supervised learning (SSL) in the multimodal era, particularly in optical character recognition (OCR). Xie also introduces VSI-Bench, a benchmark for evaluating a model's spatial understanding from videos, highlighting the limitations of current models in spatial reasoning compared to humans. - [Lisa Dunlap](https://www.lisabdunlap.com/) discussed the challenges of evaluating Large Multimodal Models (LMMs) beyond traditional metrics. She discusses Chatbot Arena, a platform that gathers real user conversations and votes to develop a live leaderboard. Emphasizing the need for detailed quality scores, she focuses on subjective aspects like reasoning, tone, and style. The concept of a “vibe check” is introduced to assess these properties, advocating for customizable evaluation interfaces that offer tailored model recommendations based on user preferences, rather than generalized leaderboards. As you read through my takeaways from each talk, a common thread becomes clear: we’re entering an era where benchmarking must evolve beyond sterile leaderboards toward more human-centred, spatially-aware, and personalized evaluation frameworks. The future of AI depends on it. ## Andre Araujo: Multimodal AI is Amazing... Yet Deeply Flawed ![](https://cdn.sanity.io/images/h6toihm1/production/20bca88ed225e3cbf502b173b5467f7c96084744-1400x1058.webp?auto=format&dpr=2&fit=max&q=75&w=1400) Multimodal AI models are showing mind-blowing capabilities that seem almost magical. These systems can seamlessly process combinations of text, images, video, and audio to handle complex user requests. They excel at retrieving information from massive data collections, directly analyzing visual content without external help, and can even leverage specialized tools to expand their capabilities. The architectural innovation behind these systems involves complex fusion mechanisms that align representations across modalities, enabling them to understand relationships between visual elements and textual descriptions in ways previously impossible. Current models are just scratching the surface of true multimodal Intelligence. ### The Embarrassing Reality Check ![](https://cdn.sanity.io/images/h6toihm1/production/e9fffa7cf59edf02e1ed86cfdb92c50a38898eff-800x508.webp?auto=format&dpr=2&fit=max&q=75&w=800) Multimodal AI systems fail spectacularly at tasks that even toddlers can handle. These supposedly advanced systems struggle with basic spatial reasoning, often failing to count objects correctly or determine the direction of bird flight in an image. They demonstrate poor fine-grain visual understanding, regularly misidentifying specific bird species or artworks despite having access to specialized recognition tools. The fundamental issue lies in their representation learning — current visual encoders excel at global image understanding but lack the dense prediction capabilities needed for precise spatial reasoning and object localization, creating a significant gap between human-like perception and machine interpretation. The gap between flashy demos and real-world reliability remains frustratingly wide. ### HAMMR Multimodal ReACT ![](https://cdn.sanity.io/images/h6toihm1/production/cc185972e50cee2c03befc124144879407a65c08-1400x1058.webp?auto=format&dpr=2&fit=max&q=75&w=1400) [HAMMR](https://arxiv.org/html/2404.05465v2) introduces a paradigm shift in how AI systems leverage tools to solve complex problems. Unlike traditional approaches that ineffectively cram dozens of tools and examples into a single prompt, HAMMR implements a hierarchical, modular architecture of specialized agents. Each agent manages a carefully curated subset of tools and can dynamically call other specialized agents when needed, creating a compositional problem-solving approach. This architecture extends the React (Reasoning and Acting) framework beyond text to support truly multimodal variables — including images, video, and audio — as both inputs and outputs, with each observation step providing explicit descriptions of variable assignments. HAMMR’s iterative “thought, action, observation” loop enables unprecedented flexibility and error recovery. ## TIPS: Engineering Spatial Understanding ![](https://cdn.sanity.io/images/h6toihm1/production/d8205570f55368d93761dab28436efc7ed3526ad-1400x1058.webp?auto=format&dpr=2&fit=max&q=75&w=1400) [TIPS](https://arxiv.org/html/2410.16512v2) fundamentally reimagines how visual foundation models encode spatial information. The approach addresses the critical limitation of existing multimodal encoders by introducing a dual-CLS token architecture in vision transformers— one token aligns with noisy web captions for main object recognition. In contrast, a second token aligns with synthetic, spatially rich captions that describe backgrounds and spatial relationships. This innovation is complemented by self-supervised masked image modeling techniques like Ibot loss, which incentivize the model to learn location- sensitive features by reconstructing masked image patches, effectively teaching the system to understand “where” in addition to “what.” This architectural breakthrough enables robust performance across both global and dense vision tasks. ### UDON: Mastering Fine-Grain Visual Understanding ![](https://cdn.sanity.io/images/h6toihm1/production/336f63e8e9130ac10f4820381d9ba697eac9149d-1400x638.webp?auto=format&dpr=2&fit=max&q=75&w=1400) [UDON](https://arxiv.org/html/2406.08332v2) addresses the persistent challenge of multi-domain, fine-grained visual recognition. The technique employs a multi-teacher distillation approach that first trains separate “teacher” models, specialized for distinct domains such as landmarks, food, or artwork, each capturing domain-specific visual cues and taxonomies. It then consolidates this specialized knowledge into a universal “student” model through dynamic sampling strategies that balance domains with vastly different class distributions. The resulting unified model maintains domain expertise without forcing contradictory visual cues to compete, enabling unprecedented accuracy in identifying subtle visual distinctions across diverse categories. This elegant solution demonstrates how carefully designed knowledge transfer can overcome fundamental limitations in visual representation learning. ### Benchmarking is an Important Reality Check Rigorous benchmarking exposes the real capabilities and limitations of multimodal AI. Effective evaluation requires testing both global understanding (like image-text retrieval) and dense prediction tasks (such as semantic segmentation and depth estimation). Challenging benchmarks push models to demonstrate fine-grain visual understanding through “single-hop” and “two-hop” questions that require identifying visual elements and then retrieving related factual knowledge. Performance gaps on these benchmarks reveal fundamental limitations in current architectures — HAMMR shows nearly 20% improvement over standard tool-use approaches, yet still falls short on complex spatial reasoning tasks that require integrated understanding across modalities. Only through systematic and multifaceted evaluation can we drive the architectural innovations necessary for truly robust multimodal intelligence. ## Saining Xie: Language Shortcuts Undermine Visual Intelligence ![](https://cdn.sanity.io/images/h6toihm1/production/7958c902e5e27e46bf22005bc91415030bbbcc6e-1400x1058.webp?auto=format&dpr=2&fit=max&q=75&w=1400) Current multimodal models are cheating with language instead of truly understanding what they see. Language intelligence is a powerful shortcut that masks significant gaps in visual understanding. These models often just associate visual symbols with pre-existing knowledge rather than developing genuine visual intelligence. The resulting systems might perform well on benchmarks but fail catastrophically when deployed in real-world scenarios that require robust visual reasoning. We need benchmarks that force models to develop actual visual intelligence, not just leverage language priors. ### Self-Supervised Learning Makes a Comeback Self-supervised learning models are finally closing the gap with language-supervised approaches, such as CLIP. Previous performance differences weren’t due to inherent weaknesses in SSL methodology but simply insufficient scale. By training on billion-scale web data and scaling parameters beyond 1 billion, SSL models now outperform CLIP on average visual perception benchmarks. The most dramatic improvements are observed in challenging domains, such as OCR and character recognition, suggesting that SSL’s true potential was previously underestimated. SSL’s remarkable responsiveness to data distribution makes it an incredibly promising path forward. ### Visual Search Is Non-Negotiable Deliberate visual processing isn’t optional — it’s a fundamental capability that all multimodal models must possess. [The V\* model](https://arxiv.org/abs/2312.14135) demonstrates how integrating methodical visual search enables AI to focus on crucial details when answering complex, high-resolution questions. This approach mirrors human cognition, where we allocate additional processing power to difficult visual tasks rather than relying on obvious but potentially misleading cues. OpenAI’s “thinking with image” feature validates this direction by achieving near-perfect scores on the V\* benchmark. This isn’t just an engineering trick — it’s a core cognitive mechanism that future models cannot succeed without. ### Video Benchmarks Miss the Point ![](https://cdn.sanity.io/images/h6toihm1/production/489b714b99f09dca548a64449b4a17a0c93c5903-1400x468.webp?auto=format&dpr=2&fit=max&q=75&w=1400) Most current video understanding benchmarks are fundamentally broken and test the wrong things. These benchmarks inadvertently reward knowledge retrieval and linguistic understanding instead of genuine visual-spatial reasoning. Questions like “Why are objects flying?” in scientifically incorrect videos or queries about astronaut equipment don’t test visual understanding — they’re glorified trivia contests. This misalignment creates models with impressive benchmark scores but crippling real-world limitations. We’re heading down a dangerous path if we continue optimizing for these flawed metrics. ### VSI-Bench Forces Spatial Thinking ![](https://cdn.sanity.io/images/h6toihm1/production/978815455d3e842c7dd326a3b673090ba40563eb-1400x645.webp?auto=format&dpr=2&fit=max&q=75&w=1400) [VSI-Bench](https://vision-x-nyu.github.io/thinking-in-space.github.io/) represents a new approach that makes models genuinely think in three-dimensional space. Unlike traditional benchmarks focused on recognition or storytelling, VSI-Bench evaluates mental imagery and spatial manipulation abilities. The tasks — from counting objects and determining relative directions to planning routes and estimating dimensions — require models to track spatial relationships across extended video sequences. By repurposing existing 3D datasets to automatically generate high-quality video-question pairs, this approach becomes both rigorous and practical. The massive performance gap between humans and state-of-the-art models like Gemini 1.5 Pro on VSI-Bench should serve as a wake-up call. ### Models Fail at Spatial Logic AI models recognize objects perfectly but can’t figure out where they are in relation to each other. Analysis shows that 71% of errors on VSI-Bench stem from spatial reasoning failures rather than visual perception or language understanding problems. Surprisingly, common linguistic reasoning techniques like chain-of-thought prompting actually degrade performance on spatial tasks. Current models can handle objects appearing together in single frames but collapse when tracking relationships across time. We need fundamentally new mechanisms for spatial reasoning, not just more data or parameter scaling. ### Spatial Supersensing Is the Future The ultimate goal is for AI to understand physical space as effortlessly as humans navigate the world. This vision extends far beyond current chatbot interactions to encompass always-on spatial intelligence, integrated into devices such as AI glasses. Achieving this capability requires abandoning brute-force encoding of every pixel and frame in favour of more efficient mechanisms that can process unlimited visual information. Current models remain “definitely worse than cats” at this crucial capability despite their impressive performance in other domains. This is a massively exciting, open frontier that demands novel approaches to truly **ground AI in the real world**. ## Lisa Dunlap: The Problem with Single-Number Leaderboards ![](https://cdn.sanity.io/images/h6toihm1/production/e24960046b9c5afb304dafbefe2935cc401060e1-1400x1058.webp?auto=format&dpr=2&fit=max&q=75&w=1400) Traditional LLM leaderboards are completely missing the point. They reduce complex AI systems to a single number, which is like rating a chef solely on how fast they cook. Generative AI quality isn’t just about correctness — it’s about tone, style, explanation approach, and countless subjective properties that users actually care about. These nuanced characteristics (or “vibes”) are what make people prefer one model over another in real-world usage. Single metrics just don’t cut it anymore. ![](https://cdn.sanity.io/images/h6toihm1/production/d540e7f86f24585fc02fb21bded5d5127d44b44a-1400x1058.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ### The Chatbot Arena Revolution Chatbot Arena ( [now known as LMArena](https://lmarena.ai/)) is changing the evaluation game entirely. It collects millions of real user conversations and pairwise preferences, allowing people to directly compare anonymous models side by side. With over 100 million queries and 3 million votes, it’s become the go-to platform for understanding how models perform “in the wild” across countless languages and tasks. The battle mode brilliantly forces users to make explicit choices about which response they prefer. This is the evaluation that actually matters. ### Style Matters More Than We Thought ![](https://cdn.sanity.io/images/h6toihm1/production/b9cb4732be581c18aacbe26c7914465d29dc63a5-1400x1058.webp?auto=format&dpr=2&fit=max&q=75&w=1400) Users are being tricked by verbose models. Longer responses consistently win user preferences, even when they’re not better answers. Models like GPT-4 sometimes exploit this by generating unnecessarily lengthy responses (like writing paragraphs to answer “what’s the scientific name for octopus?”). When researchers control for style factors, some models’ rankings change dramatically, revealing they’ve been “style hacking” rather than providing genuinely better answers. Style is the hidden influencer of user preference. ### The “Vibe Check” Methodology ![](https://cdn.sanity.io/images/h6toihm1/production/b9cb4732be581c18aacbe26c7914465d29dc63a5-1400x1058.webp?auto=format&dpr=2&fit=max&q=75&w=1400) Manually identifying all the subjective qualities that matter to users is impossible, so we desperately need to automate “vibe” discovery. The traditional approach of hand-defining categories for evaluation is limiting, as there are numerous qualitative differences in how models respond that we may not consider beforehand. This is where the “vibe check” comes in, leveraging LLMs themselves as a tool to automatically discover and quantify these subjective properties (or “vibes”) directly from model outputs. It’s about finding what’s _user-aligned, well-defined, and differentiating_ between models. [The “vibe check” approach](https://arxiv.org/abs/2410.12851) uses LLMs themselves to analyze and score responses on subjective qualities like friendliness, humor, or formality. It was developed to move beyond the limitations of hand-defined features for evaluating qualitative differences in models, seeking to uncover properties users truly care about that might be hard to anticipate beforehand. A “good vibe” in this context is defined by three criteria: - **User aligned:** It’s a property relevant to the user. - **Well-defined:** Different judges (human or LLM) would consistently agree on its presence (e.g., if a response is “friendly”). - **Differentiating:** Models being evaluated should show clear differences in this property. The method operates through a **four-step process**: 1. **Discovering Vibes:** An LLM is prompted to identify differences between pairwise model responses from a subset of data (like Chatbot Arena battles), compiling a list of frequently appearing “candidate vibes”. 2. **Scoring Outputs:** A panel of smaller LLM judges then scores each model output for the presence of these discovered vibes (e.g., “which output is more friendly?”), providing fine-grained subjective analysis. 3. **Filtering Properties:** Vibes are filtered out if there is low agreement among judges (meaning it’s not well-defined) or if the vibe is perceived equally across all model outputs (meaning it does not differentiate). 4. **Quantifying Utility:** Finally, logistic regression is used to predict which model generated an output or which output a user would prefer, based _solely_ on these identified “vibes.” This reveals the impact each vibe has on a model’s identity and user preference. The “vibe check” has proven incredibly useful, particularly in explaining why models like **Llama 3 ranked highly on preference benchmarks like Chatbot Arena** despite potentially lower performance on traditional objective metrics. It revealed Llama 3 was perceived as **more friendly, funny, and interactive**, which positively correlated with user preference, unlike GPT-4 and Claude which were more formal or ethics-focused. The method is **agnostic to Chatbot Arena**, meaning it can be applied to any benchmark to uncover distinct, context-specific differences between models that traditional metrics might miss. It can even **inform how to adjust a model’s behavior** to influence its perceived quality, as shown by re-prompting Gemini to include subjective interpretations, which improved its preference among judges. Vibes are the missing piece of AI evaluation. ### The Future is Personalized Evaluation ![](https://cdn.sanity.io/images/h6toihm1/production/6b636a58d2d5a61c73d938f315671d4fa35ba748-1400x1058.webp?auto=format&dpr=2&fit=max&q=75&w=1400) One-size-fits-all rankings need to die. The future of LLM evaluation is undeniably moving towards **personalization**, recognizing that a “one-size-fits-all” leaderboard with a single numerical score is insufficient for real-world applications. The ultimate goal of benchmarks is to evaluate general-purpose agents designed to cater to the diverse personal needs of everyone, which generates a vast amount of information that is currently difficult to interpret. While platforms like **Chatbot Arena** and **Helm** provide extensive data and decompose quality beyond a single number, they still struggle to make this information genuinely useful and actionable for individual users. - Helm, despite its comprehensive array of benchmarks and aggregated results, can present a “light” leaderboard that is difficult for a user to interpret and decide which model to use. - Chatbot Arena’s category leaderboards, while offering more detail, are still “somewhat reductive” and complex to understand, especially when considering the many nuances within specific tasks. - Even with methods like “vibe checks” that offer personalized insights into subjective properties, the challenge remains in scaling and easily understanding these highly personalized results. The key to unlocking truly useful evaluation lies in **providing interfaces for benchmarks that can customize results for each user.** This doesn’t mean creating entirely new evaluations for every person, but rather intelligently presenting existing benchmark data to offer a customized recommendation experience. This customization can occur in two main ways: **Customizing to a Specific Task:** - The idea is to allow users to input a specific task they care about and receive a predicted leaderboard tailored to that task. - An example of preliminary work in this area is **“Prompt to Leaderboard”**, a model trained on Arena battles that predicts what the leaderboard will look like based on a user’s specific prompt input. **Customizing to User Preferences:** - This focuses on the more subjective aspects of model preference. - The **“vibe check”** method, as discussed previously, is a prime example, automatically discovering and quantifying subjective properties that influence user preference. - Other preliminary work includes papers like **“Report Cards”** and **“Inverse Constitutional AI”**, which aim to generate natural language descriptions of models’ properties. Despite these advancements, a method that seamlessly combines both **task-specific customization** and **user-preference customization** is still an area with significant work to be done. ### Why Personalization is the Evolution of Evaluation Leaderboards exist to guide users and developers in choosing the best model or checkpoint for their needs. However, the definition of “best” is highly individual: - For a developer, it’s about selecting the right model for production. - For a user, it’s about deciding which subscription to buy. - Crucially, **what model to use depends fundamentally on “what you’re using it for and who you are as a person”**. The biggest question in evaluation, therefore, is how to present metrics so that two people with the same task but different preferences (or vice versa) can find the optimal model for _them_. To facilitate this personalized future, model providers should consider: - **Enhanced Customization:** Implementing ideas like **“memory”** (e.g., based on concepts like MemGPT), which learns specific user context (like being a PhD student in AI) through conversations to better customize responses. - **Utilizing Implicit Feedback:** There are many points in multi-turn conversations where users implicitly signal their preferences or what they truly desire. Finding better ways to leverage this feedback signal in the training process or personalization techniques will be incredibly important. Personalization is the endgame of AI evaluation. ## The Path Forward As Day 1 of CVPR 2025 made abundantly clear, the gap between impressive demos and fundamental limitations can no longer be ignored. Araujo’s architectural innovations, Xie’s spatial reasoning benchmarks, and Dunlap’s subjective “vibe checks” collectively point to a new evaluation paradigm — one that prizes human-like spatial understanding and user-relevant qualities over misleading metrics. The next generation of truly capable multimodal systems won’t emerge from chasing leaderboard positions, but from confronting these uncomfortable truths about what our models still can’t do. The question isn’t just “which model ranks highest?” but “which model thinks spatially, understands deeply, and resonates personally?” That’s the benchmark that matters. [CVPR](https://voxel51.com/blog/tag/cvpr) [benchmark](https://voxel51.com/blog/tag/benchmark) [multimodal](https://voxel51.com/blog/tag/multimodal) [model evaluation](https://voxel51.com/blog/tag/model-evaluation) ![](https://cdn.sanity.io/images/h6toihm1/production/a41a0477c7a98264f600772e9568607d070eea59-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Sahota Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a558b86370f2f17212fb2f2c894d590101458a85-5760x3241.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Multimodal Frontier in Computer Vision, Medicine, and Agriculture— CVPR 2025 Reflections\\ \\ Event Recaps, Industry Solutions\\ \\ • \\ \\ Jun 24, 2025](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections) [![](https://cdn.sanity.io/images/h6toihm1/production/97cf3d887735ab9574b2b3e2d3825146016d875e-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Embodied Computer Vision at CVPR 2025: The Next AI Frontier\\ \\ Event Recaps\\ \\ • \\ \\ Jun 30, 2025](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier) [![](https://cdn.sanity.io/images/h6toihm1/production/e468545aa08daf6c6829d2593ffd8b5457c7dee5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ NVIDIA’s C-RADIOv3 is the Vision Encoder You Should Be Using\\ \\ Event Recaps, Integrations\\ \\ • \\ \\ Jun 23, 2025](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-4-lllmstxt|> ## CVPR 2024 Datasets Overview [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) CVPR 2024 Datasets and Benchmarks – Part 1: Datasets Apr 23, 2024 • 17 min read Article content In this article [Panda-70M](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets#40862fd600ac) [360 + x](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets#72d1bd1ca882) [TSP6K](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets#e23563ec7016) [Conclusion](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets#949e0665ad4b) [Visit Voxel51 at CVPR 2024!](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets#04d4d215f558) In this article [Panda-70M](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets#40862fd600ac) [360 + x](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets#72d1bd1ca882) [TSP6K](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets#e23563ec7016) [Conclusion](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets#949e0665ad4b) [Visit Voxel51 at CVPR 2024!](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets#04d4d215f558) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) CVPR, the IEEE/CVF Conference on Computer Vision and Pattern Recognition, is the Coachella Festival of computer vision. Just like any music festival needs its headliners, deep learning needs its rockstars: datasets and benchmarks. Over the years, these have played a massive role as the driving force in advancing computer vision and deep learning in general, and CVPR has consistently been a platform for introducing new ones. Without a platform like CVPR for researchers to present new datasets, progress in deep learning wouldn’t be as fast as it has been over the last decade. But what sets a dataset apart from a benchmark? A dataset is a collection of data samples, such as images, videos, or annotations, used to train and evaluate deep learning models. It provides the raw data that models learn from during the training process. Datasets are often labeled or annotated to provide ground truth information for supervised learning tasks. In computer vision, famous datasets from previous CVPRs include 2009’s [ImageNet](https://ieeexplore.ieee.org/document/5206848), 2016’s [Cityscapes Dataset](https://openaccess.thecvf.com/content_cvpr_2016/html/Cordts_The_Cityscapes_Dataset_CVPR_2016_paper.html), 2017’s [Kinetics Dataset](https://openaccess.thecvf.com/content_cvpr_2017/html/Carreira_Quo_Vadis_Action_CVPR_2017_paper.html), and 2020’s [nuScenes](https://openaccess.thecvf.com/content_CVPR_2020/html/Caesar_nuScenes_A_Multimodal_Dataset_for_Autonomous_Driving_CVPR_2020_paper.html). These large-scale datasets have enabled the development of state-of-the-art deep models for tasks like image classification, object detection, and semantic segmentation. On the other hand, benchmarks are standardized tasks or challenges used to evaluate and compare the performance of different models or algorithms. Benchmarks typically consist of a dataset, a well-defined evaluation metric, and a leaderboard ranking the performance of different models. They are the yardstick against which models are measured, allowing researchers to gauge performance against state-of-the-art approaches and track progress in a specific task or domain. Examples include the 2009 [Pedestrian Detection Benchmark](https://ieeexplore.ieee.org/document/5206631) and 2012’s [KITTI Vision Benchmark Suite](https://www.cvlibs.net/datasets/kitti/). The importance of datasets and benchmarks in deep learning can’t be overstated. Datasets serve as the foundation for training deep learning models. At the same time, benchmarks allow for the objective assessment and comparison of model performance. This is why introducing new research in this direction pushes research boundaries and addresses the limitations of current approaches. New datasets are essential to exploring novel tasks, accommodating emerging model architectures, and capturing diverse real-world scenarios. As models achieve higher - and eventually human-level - performance on established benchmarks, researchers create new, more challenging ones to push the boundaries of what's possible. Several new datasets and benchmarks at CVPR 2024 have caught my eye, each with the potential to advance progress in computer vision and deep learning. Do a ctrl-f for “datasets” on the [list of accepted papers](https://cvpr.thecvf.com/Conferences/2024/AcceptedPapers), and you’ll see a whopping 72 results. These datasets cover a spectrum of computer vision tasks, including multimodal, multi-view, room layout, 3D//4D/6D, robotic perception datasets, and more. Admittedly, the topic I’m most interested in is vision language and multimodality, so my picks are biased towards that. I’ll explore both in this two-part blog series, with this blog focusing on three datasets that I found interesting (presented in no particular order): - Panda-70M - 360 + _x_ - TSP6K Similarly, in part two, I’ll discuss three benchmarks that captured my interest: - ImageNet-D - Polaris - VBench In the following sections, I'll focus on the following aspects of each dataset: - **Task and Domain:** Pinpoint each dataset's specific problem and area within the computer vision landscape. - **Dataset Curation, Size, and Composition:** Explore the scale of each dataset, its creation process, and the diversity of data it encompasses. - **Unique Features and Innovations:** Uncover what sets each dataset apart and any novel approaches employed in its development. - **Potential Applications and Impact:** Discuss how each dataset can move the field forward and its potential real-world applications. - **Comparison to Existing Datasets:** Provide context by comparing each dataset to similar ones, showcasing how it fills existing gaps or offers unique advantages. - **Accessibility and Usability:** Share information about each dataset's availability, licensing, and where you can find it. - **Experiments and Results:** When applicable, I'll summarize key experimental findings for each dataset. # Panda-70M ### tl;dr - **Task:** Video captioning - **Size:** The full training set has 70.7M samples at 720p resolution, an average duration of 8.5s (totalling 167 Khrs), and an average caption of 13.2 words. - **License:** [Non-commercial, research only](https://raw.githubusercontent.com/microsoft/XPretrain/main/hd-vila-100m/LICENSE) - [Paper](https://arxiv.org/html/2402.19479v1) ### Task and Domain [Panda-70M](https://snap-research.github.io/Panda-70M/) is for multimodal learning tasks that involve video and text, specifically targeting video captioning, video and text retrieval, and text-driven video generation. The domain is open, meaning the dataset includes a wide range of subjects and is not limited to any specific field or type of content. ### Dataset Curation, Size, and Composition [Panda-70M](https://snap-research.github.io/Panda-70M/) contains 70.8 million video clips with automatically generated captions. It was created by curating 3.8 million high-resolution videos from the publicly available [HD-VILA-100M](https://github.com/microsoft/XPretrain/tree/main/hd-vila-100m) dataset. The dataset is created following these three steps: 1. Split videos into semantically coherent clips (see section 3.1 of the paper for details). 2. Caption each clip using various cross-modality teacher models, like image/video visual question-answering models, with additional inputs/metadata about the video (discussed below in the next section). 3. Take a 100k subset of this data and have human annotators generate ground truth captions. 1. Fine-tune a retrieval model on this subset (see section 3.3 of the paper for more details). 2. Select the most precise annotation describing each scene of the videos. The result is a diverse dataset, with videos that are semantically coherent, high-resolution, and free of watermarks. The captions describe each video's main object and action, with an average of 13.2 words per caption. The authors have pointed out that the dataset has a limitation, which is an artifact of taking a subset of HD-VILA-100M: most of the samples are videos with a lot of speech. Therefore, the main categories in our dataset are news, television shows, documentary films, egocentric videos, and instructional and narrative videos. ### Unique Features and Innovations The authors' core insight is that a video typically contains information from several modalities that can assist an automatic captioning model. These include the title, description, subtitles, static frames, and video. These modalities should be used to create a large-scale video-language dataset with rich captions more efficiently than manual annotation. - **Automatic Annotation:** Utilizes multiple cross-modality teacher models to generate captions, eliminating the need for expensive manual annotation. - **Semantics-aware Splitting:** Ensures video clips are semantically consistent while maintaining sufficient length for capturing motion content. - **Fine-grained Retrieval:** Employs a retrieval model specifically fine-tuned to select the most accurate and detailed caption from multiple candidates. - **Multimodal Input:** Takes advantage of various modalities and metadata like video frames, titles, descriptions, and subtitles to generate captions. ### Potential Applications and Impact [Panda-70M](https://snap-research.github.io/Panda-70M/) can enable more effective pretraining of video-language models, driving progress on video understanding tasks. By pretraining on Panda-70 M, the authors demonstrate substantial improvements in video captioning, retrieval, and text-to-video generation. Potential applications: - **Video Captioning:** Training and evaluating video captioning models with improved accuracy and detail. - **Video and Text Retrieval:** Facilitating accurate retrieval of videos based on text queries and vice versa. - **Text-driven Video Generation:** This dataset will enable the generation of videos from textual descriptions, potentially leading to advancements in video editing and content creation. - **Multimodal Understanding:** Advancing research in multimodal learning and understanding the relationships between vision and language. ### Comparison to Existing Datasets The [Panda-70M](https://snap-research.github.io/Panda-70M/) video captioning dataset is much larger than other datasets that rely on manual labelling. It overcomes the limitations of ASR annotations, which often fail to capture the main content and actions, and provides a solution to manual annotations' cost and scalability issues. - **Larger Scale:** Compared to existing manually annotated video captioning datasets, Panda-70M offers a significantly larger scale. - **Richer Captions:** Provides more detailed and informative captions than datasets annotated with ASR-generated subtitles. - **Open Domain:** This dataset covers a broader range of video content than datasets focused on specific domains, like cooking or actions. ### Accessibility and Usability The authors make the [Panda-70M](https://snap-research.github.io/Panda-70M/) dataset and code publicly available. This enables the research community to build upon this resource and utilize it for various video-language tasks. However, [according to the license](https://raw.githubusercontent.com/microsoft/XPretrain/main/hd-vila-100m/LICENSE), the dataset cannot be used commercially. ### Experiments and Results The paper provides benchmark results of models pretrained on [Panda-70M](https://snap-research.github.io/Panda-70M/) on 3 downstream tasks: - Video captioning (evaluated using metrics like [BLEU-4, ROUGE-L, METEOR, CIDEr](https://www.sciencedirect.com/science/article/pii/S2666827023000415)) - Video-text retrieval (R@1, R@5, R@10) - Text-to-video generation ( [FVD](https://openreview.net/pdf?id=rylgEULtdN), [CLIPSIM](https://dl.acm.org/doi/pdf/10.1145/3588707)) The author pretrained models on Panda-70M and presented impressive results, even beating SOTA models like Video-LLaMA and BLIP-2. # 360 + _x_ ### tl;dr - **Task:** Panoptic multimodal scene understanding - **Size:** 2,152 videos representing 232 data examples (464 from 360 camera, 1,688 from Spectacles) - **License:** [Creative Commons Attribution-NonCommercial-ShareAlike 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en) (non-commercial) - [Paper](https://arxiv.org/html/2404.00989v2) - [Dataset on HuggingFace](https://huggingface.co/datasets/quchenyuan/360x_dataset) ### Task and Domain [360+x](https://x360dataset.github.io/) is a panoptic multimodal scene understanding dataset that provides a holistic view of scenes from multiple viewpoints and modalities. The dataset focuses on **real-world, everyday scenes** with diverse activities and environments. It’s especially suited for tasks like: - **Video classification:** Identifying the scene category (e.g., park, restaurant) - **Temporal action localization:** Detecting and classifying actions within a video (e.g., eating, walking) - **Self-supervised representation learning:** Learning meaningful representations from unlabeled data - **Cross-modality retrieval:** Retrieving relevant information across different modalities (e.g., finding video segments based on audio cues) - **Dataset adaptation:** Adapting pre-trained models to new datasets ### Dataset Size and Composition [360 + _x_](https://x360dataset.github.io/) was curated using two main devices: the Insta 360 One X2 and Snapchat Spectacles 3 cameras, ensuring high resolution and frame rate for both video and audio modalities. The dataset contains 360° panoramic videos, third-person front view videos, egocentric monocular/binocular videos, multi-channel audio, directional binaural delay information, GPS location data, and textual scene descriptions. - **Size:** 2,152 videos representing 232 data examples (464 from 360 camera, 1,688 from Spectacles) - **Composition:** - **Viewpoints:** 360° panoramic, third-person front view, egocentric monocular, egocentric binocular - **Modalities:** Video, multi-channel audio, directional binaural delay, location data, textual scene descriptions - **Curation:** - **Scene selection:** Based on comprehensiveness, real-life relevance, diverse weather/lighting, and rich sound sources - **Annotation:** Scene categories (28) based on Places Database and large language models; temporal action segmentation (38 action instances) with consensus among annotators ### Unique Features and Innovations [360+x](https://x360dataset.github.io/) is the first dataset to cover multiple viewpoints (panoramic, third-person, egocentric) with multiple modalities (video, audio, location, text) to mimic real-world perception. Using novel data collection techniques, including 360-degree camera technology for panoramic video, it captures the following: - **Multiple viewpoints and modalities:** Offers a more holistic understanding of scenes than unimodal datasets. - **Real-world complexity:** Captures diverse activities and interactions in various environments, providing a more realistic challenge. - **Directional binaural delay:** Enables sound source localization and richer audio analysis. - **Egocentric and third-person views:** Allows for studying both participant and observer perspectives. ### Potential Applications and Impact This dataset's rich multimodal data opens up new avenues for research in multimodal deep learning, particularly in multi-view, multimodal data for video understanding tasks. This will undoubtedly have a huge positive impact in real-world applications, such as enhancing robot navigation systems and enriching user interaction within smart environments. Here are some ways I think it could impact research: - **Multiple Viewpoints and Modalities**: By offering a more comprehensive understanding of scenes compared to unimodal datasets, this feature allows for a holistic interpretation of complex environments. - **Real-World Complexity**: The dataset captures a wide array of activities and interactions across diverse environments, providing a more realistic and challenging context for research. - **Directional Binaural Delay**: This feature enables precise sound source localization and facilitates more in-depth audio analysis. - **Egocentric and Third-Person Views**: This unique combination allows researchers to study participant and observer perspectives, offering a complete understanding of human-environment interaction. ### Comparison to Existing Datasets Unlike datasets that focus on specific viewpoints, such as egocentric or third-person views alone, or those that lack audio-visual correlation, 360+ _x_ combines multiple viewpoints with multimodal data. Below are some of the limitations associated with existing datasets - **UCF101, Kinetics, HMDB51, ActivityNet:** Primarily focus on video classification with limited scene complexity and single viewpoints. - **EPIC-Kitchens, Ego4D:** Focus on egocentric videos lacking other viewpoints and modalities. - **AVA, AudioSet, VGGSound:** Focus on audio-visual analysis but lack multiple viewpoints and directional audio information. 360+x addresses these limitations by providing a more comprehensive and realistic dataset with multiple viewpoints, modalities, and rich annotations. It also has longer 6-minute videos capturing more complex co-occurring activities than short single-action clips in most datasets. ### Accessibility and Usability The dataset, code, and associated tools are publicly available for research under the [Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License](http://creativecommons.org/licenses/by-nc-sa/4.0/), which means they are unavailable for commercial use. ### Experiments and Results The paper establishes several baselines on 360+x for tasks such as video classification, temporal action localization, and cross-modal retrieval. Some interesting findings include: **Video Scene Classification** - A 360° panoramic view outperforms an egocentric binocular or third-person front view alone. - Utilizing all three views (360°, egocentric binocular, third-person front) leads to the best performance. - Including additional modalities (audio, directional binaural information) and video improves the average precision. **Temporal Action Localization (TAL)** - Adding audio and directional binaural delay modalities improves baseline and extractors pre-trained on 360+x dataset performance. - Using custom extractors pre-trained on 360+x dataset provides additional improvements over baseline extractors. **Cross-modality Retrieval** - The intermodality retrieval results (Query-to-Video, Query-to-Audio, Query-to-Directional binaural delay) show the 360+x dataset's modality compliance quality. - Collaborative training of audio and directional binaural features as query features leads to better performance than treating them independently. **Self-supervised Representation Learning** - Using self-supervised learning (SSL) pre-trained models consistently improves precision for video classification, with a combination of video pace and clip order SSL techniques resulting in ~7% improvement. - For the TAL task, pre-training with video pace or clip order individually improves average performance by ~1.2% and ~0.9%, respectively, compared to the supervised baseline. Combining both SSL methods yields the highest performance gain of ~1.9%. # TSP6K ### tl;dr - **Task:** Traffic scene parsing and understanding - **Size:** 2,999 training images, 1,207 validation images, and 1,794 test images, which include: - **License:** Repo is under Apache 2.0 with [a warning about the SegFormer license](https://github.com/PengtaoJiang/TSP6K/blob/main/LICENSES.md). - [GitHub Repository](https://github.com/PengtaoJiang/TSP6K) ### Task and Domain TSP6K is a dataset for traffic scene parsing and understanding, focusing on semantic and instance segmentation. Scene parsing involves segmenting and understanding various components within an image, identifying semantic objects and the surrounding environment. This dataset is tailored for traffic monitoring rather than general driving scenarios, providing detailed pixel-level and instance-level annotations for urban traffic scenes. ### Dataset Size and Composition The TSP6K dataset includes 6,000 finely annotated images, which are specifically collected from urban road shooting platforms at different locations and times of the day. It’s split into 2,999 training images, 1,207 validation images, and 1,794 test images, which include: - High-resolution images were captured from traffic monitoring cameras positioned across 10 Chinese provinces. - Careful selection of images ensures diversity in scene type (e.g., crossings, pedestrian crossings), weather conditions (sunny, cloudy, rain, fog, snow), and time of day. - Pixel-level annotations for 21 semantic classes (e.g., road, vehicles, pedestrians, traffic signs, lanes) and instance-level annotations for all traffic participants. This approach ensures a rich diversity in the dataset, capturing various traffic conditions and lighting environments. Each image in the dataset comes with high-quality semantic and instance-level annotations, which are necessary for detailed scene analysis in traffic monitoring applications. ### Unique Features and Innovations Some interesting features of this dataset are: - Crowded traffic scenes with significantly more traffic participants than existing driving scene datasets present a more challenging and realistic scenario for model training and evaluation. - An average of 42 traffic participants per image compared to 5-19 in driving datasets. - 30% of images contain over 50 instances. - It captures a wide range of object scales from the perspective of the monitoring cameras, forcing models to handle significant size variations. ### Potential Applications and Impact By providing a dataset specifically for traffic monitoring, the TSP6K dataset opens up new research directions, like: - **Traffic Flow Analysis:** Improved scene understanding can lead to better traffic flow monitoring, congestion prediction, and anomaly detection. - It enables the training and evaluation of scene parsing, instance segmentation, and unsupervised domain adaptation methods for traffic monitoring applications. - **Smart City Development:** The dataset can contribute to developing intelligent traffic management systems and infrastructure planning. - **Domain Adaptation Research:** The domain gap between TSP6K and existing datasets provides a valuable benchmark for evaluating and advancing domain adaptation techniques. ### Comparison to Existing Datasets Datasets like Cityscapes, KITTI, and BDD100K have advanced traffic scene understanding. However, they primarily focused on autonomous driving use cases and scenarios. This means their contents are mainly collected from drivers' perspectives inside vehicles. Deep learning models trained on driver-centric datasets underperform when applied to traffic monitoring scenes. TSP6K, however, is curated from a traffic monitoring perspective. This involves higher vantage points in denser and more diverse traffic scenarios from various urban locations captured at different times. This addresses a gap in existing datasets by focusing on the unique challenges presented by traffic monitoring scenes, such as the varied sizes of instances and the complex interactions in crowded urban settings. ### Accessibility and Usability The code and dataset are publicly available to researchers at the GitHub repository. The repository is under an Apache 2.0 license; however, [the SegFormer license is mentioned](https://github.com/PengtaoJiang/TSP6K/blob/main/LICENSES.md), and users should be careful using the dataset for commercial purposes. ### Experiments and Results The paper details experiments using the TSP6K dataset to evaluate various scene parsing, instance segmentation, and unsupervised domain adaptation methods. - Evaluated the performance of existing scene parsing, instance segmentation, and domain adaptation methods on TSP6K. - Proposed a detailed refining decoder architecture that achieves 75.8% mIoU and 58.4% iIoU on the validation set. - Used mIoU and iIoU (for traffic instances) as evaluation metrics. These results set a new benchmark for traffic scene parsing performance, particularly in monitoring scenarios, and demonstrate the dataset's utility in developing robust advanced parsing algorithms across different urban traffic environments. ## Conclusion CVPR 2024 has once again showcased the vital role that datasets and benchmarks play in advancing computer vision and deep learning. The three datasets explored in this blog post—Panda-70M, 360+x, and TSP6K—each bring unique innovations and potential for impact in their respective domains. - Panda-70M's massive scale and automatic annotation approach pave the way for more effective pretraining of video-language models, enabling advancements in video captioning, retrieval, and generation. - 360+x's multi-view, multimodal data captures real-world complexity and opens up new avenues for research in holistic scene understanding. - TSP6K fills a critical gap by focusing on traffic monitoring scenarios. These scenarios present challenges distinct from driver-centric datasets and will help advance traffic scene parsing and domain adaptation. As we've seen, these datasets address limitations in existing resources, provide rich and diverse data, and establish new performance benchmarks. They are poised to drive research forward and contribute to real-world applications like content creation, smart environments, and intelligent traffic management. However, it's important to recognize that these datasets represent just a snapshot of the many groundbreaking contributions presented at CVPR 2024. At least 72 papers introduced new datasets spanning a wide range of computer vision tasks, and the conference continues to be a hub for innovation and collaboration. Stay tuned for part two of this series, where we'll dive into three exciting benchmarks from CVPR 2024: ImageNet-D, LaMPilot, and Polaris. [In the meantime, be sure to check out the Awesome CVPR 2024 Challenges, Datasets, Papers, and Workshops repository on GitHub, and feel free to open an issue if there's a specific dataset or benchmark you'd like to see covered.](http://github.com/harpreetsahota204/awesome-cvpr-2024) ## Visit Voxel51 at CVPR 2024! And if you're attending CVPR 2024, don't forget to stop by the Voxel51 booth #1519 to connect with the team, discuss the latest in computer vision and NLP, and grab some of the most sought-after swag at the event. See you there! [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [CVPR](https://voxel51.com/blog/tag/cvpr) [DreamBooth](https://voxel51.com/blog/tag/dreambooth) [F2-NeRF](https://voxel51.com/blog/tag/f2-nerf) [GLIGEN](https://voxel51.com/blog/tag/gligen) [ImageBind](https://voxel51.com/blog/tag/imagebind) [Mask DINO](https://voxel51.com/blog/tag/mask-dino) [MobileNeRF](https://voxel51.com/blog/tag/mobilenerf) [SadTalker](https://voxel51.com/blog/tag/sadtalker) [SOLDIER](https://voxel51.com/blog/tag/soldier) [VideoFusion](https://voxel51.com/blog/tag/videofusion) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/ff87d65e5b4e5ef50c5732e905f3f16aff8b0a4e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 Survival Guide\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 25, 2023](https://voxel51.com/blog/cvpr-2023-survival-guide) [![](https://cdn.sanity.io/images/h6toihm1/production/0aa3f8dad8ae1464d05d81ac4a301bd92aea55e3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2024 Survival Guide: Five Vision-Language Papers You Don’t Want to Miss\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Apr 15, 2024](https://voxel51.com/blog/cvpr-2024-survival-guide-five-vision-language-papers-you-dont-want-to-miss) [![](https://cdn.sanity.io/images/h6toihm1/production/770b8cfdbd7944916b1195dc11e5b173dbab8e97-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 and the State of Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 18, 2023](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-5-lllmstxt|> ## Exciting Computer Vision Developments [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) Why 2022 was the most exciting year in computer vision history (so far) Dec 14, 2022 • 10 min read Article content In this article [Computer Vision trends](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#70417dff40f4) [Computer Vision buzz from big tech](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#1a823bb1e9b5) [Electrifying new applications of Computer Vision](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#a5b9b1922547) [Prominent Computer Vision papers you can’t pass up](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#d99f308c0a3c) [CV tooling startups grow in size and impact](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#756a60aba6b9) [Conclusion](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#5c4eff2d6dc8) [FiftyOne Computer Vision toolset](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#18037b097ae4) In this article [Computer Vision trends](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#70417dff40f4) [Computer Vision buzz from big tech](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#1a823bb1e9b5) [Electrifying new applications of Computer Vision](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#a5b9b1922547) [Prominent Computer Vision papers you can’t pass up](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#d99f308c0a3c) [CV tooling startups grow in size and impact](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#756a60aba6b9) [Conclusion](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#5c4eff2d6dc8) [FiftyOne Computer Vision toolset](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#18037b097ae4) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/a735267ad7effa9f799f850ab7c8ffa241088710-1024x1024.png?auto=format&dpr=2&fit=max&q=75&w=1024) The past 12 months have seen rapid advances in computer vision, from the enabling infrastructure, to new applications across industries, to algorithmic breakthroughs in research, to the explosion of AI-generated art. It would be impossible to cover all of these developments in full detail in a single blog post. Nevertheless, it is worth taking a look back to highlight some of the biggest and most exciting developments in the field This post is broken into five parts: - [Big-picture trends in computer vision](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#trends) - [The latest and greatest announcements from the tech giants](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#techgiants) - [Industry-specific computer vision applications](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#newapps) - [Computer vision papers you won’t want to miss](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#cvpapers) - [Computer vision tooling news](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far#cvtooling) ## Computer Vision trends ![](https://cdn.sanity.io/images/h6toihm1/production/e3b066a6740ef365738e0f76fefb4f777e4f8a1b-1400x834.png?auto=format&dpr=2&fit=max&q=75&w=1400) ### Transformers take hold of computer vision Transformer models exploded onto the deep learning scene in 2017 with [Attention is All You Need](https://arxiv.org/pdf/1706.03762.pdf), setting the standard for a variety of NLP tasks and ushering in the era of large language models (LLMs). The [Vision Transformer](https://arxiv.org/pdf/2010.11929.pdf) (ViT), introduced in late 2020, marked the first application of these self-attention based models in a computer vision context. This year has seen research push transformer models to the forefront in computer vision, achieving state-of-the-art performance on a variety of tasks. Just check out the [panoply of vision transformer models](https://huggingface.co/docs/transformers/index) in Hugging Face’s model zoo, including [DETR](https://huggingface.co/docs/transformers/model_doc/detr), [SegFormer](https://huggingface.co/docs/transformers/model_doc/segformer), [Swin Transformer](https://huggingface.co/docs/transformers/model_doc/swin), and [ViT](https://huggingface.co/docs/transformers/model_doc/vit)! [This GitHub page](https://github.com/Yangzhangcst/Transformer-in-Computer-Vision) also provides a fairly comprehensive list of transformers in vision. ### Data-centric computer vision gains traction As computer vision matures, an increasingly large portion of machine learning development pipelines is focused on wrangling, cleaning, and augmenting data. Data quality is becoming a bottleneck for performance, and the industry is moving towards data-model co-design. The [data-centric ML](https://datacentricai.org/) movement is growing in popularity. At the helm of this effort are a new wave of startups — synthetic data generation companies ( [gretel](https://gretel.ai/), [Datagen](https://datagen.tech/), [Tonic](https://www.tonic.ai/)) and evaluation, observability, and experiment tracking tools ( [Voxel51](https://voxel51.com/), [Weights & Biases](https://wandb.ai/site), [CleanLab](https://cleanlab.ai/)) — joining existing labeling and annotation services ( [Labelbox](https://labelbox.com/), [Label Studio](https://labelstud.io/), [CVAT](https://www.cvat.ai/), [Scale](https://scale.com/), [V7](https://www.v7labs.com/)) in the effort. ### AI-generated artwork gets (too?) good Between improvements in Generative Adversarial Networks (GANs) and the rapid development and iteration in diffusion models, AI-generated art is having what can only be described as a renaissance. With tools like [Stable Diffusion](https://huggingface.co/spaces/stabilityai/stable-diffusion), [Nightcafe](https://nightcafe.studio/), [Midjourney](https://www.midjourney.com/showcase/recent/), and OpenAI’s [DALL-E2](https://openai.com/dall-e-2/) it is now possible to generate incredibly nuanced images from user-input text prompts. [Artbreeder](https://www.artbreeder.com/) allows users to “breed” multiple images into new creations, Meta’s [Make-A-Video](https://makeavideo.studio/) generates videos from text, and [RunwayML](https://runwayml.com/) has changed the game when it comes to creating animations and editing videos. Many of these tools also support [inpainting](https://huggingface.co/spaces/multimodalart/stable-diffusion-inpainting) and [outpainting](https://openai.com/blog/dall-e-introducing-outpainting/), which can be used to edit and extend the scope of images. With all of these tools revolutionizing AI art capabilities, controversy was all but inevitable, and there has been plenty of it. In September, an [AI-generated image won a fine art competition](https://www.creativebloq.com/news/ai-art-wins-competition), igniting heated debate about what counts as art, as well as how ownership, attribution, and copyrights will work for this new class of content. Expect this debate to intensify! ### Multi-modal AI matures In addition to AI-generated artwork, 2022 has seen a ton of research and applications at the intersection of multiple modalities. Models and pipelines that deal with multiple types of data, including language, audio, and vision, are becoming increasingly popular. The lines between these disciplines have never been more blurred, and cross-pollination has never been more fruitful. At the heart of this collision of contexts is [contrastive learning](https://www.v7labs.com/blog/contrastive-learning-guide#:~:text=Contrastive%20Learning%20is%20a%20technique,a%20data%20class%20from%20another.), which revamps the embedding of multiple types of data into the same space, the seminal example being Open AI’s Contrastive Language-Image Pretraining ( [CLIP](https://openai.com/blog/clip/)) model. One consequence of this is the ability to semantically search through sets of images based on input that can either text or another image. This has spurred a boom in vector search engines, with [Qdrant](https://qdrant.tech/), [Pinecone](https://www.pinecone.io/), [Weaviate](https://weaviate.io/), [Milvus](https://milvus.io/), and others leading the way. In a similar vein, the systematic connection between modalities is strengthening visual question answering and zero-shot and few-shot image classification. ## Computer Vision buzz from big tech ![](https://cdn.sanity.io/images/h6toihm1/production/a1bd3c66320a72ce9a272c4bbe7f5532da935e07-1400x1018.png?auto=format&dpr=2&fit=max&q=75&w=1400) As dataset sizes continue to grow, the computational and financial resources required to train large, high quality models from scratch has risen dramatically. As a result, many of the most broadly applicable advances this year were either led or supported by scientists from big tech research groups. Here’s some of the highlights. ### Alphabet Alphabet was active in computer vision this year, which saw the Google Brain team study the [scaling of Vision Transformers](https://arxiv.org/pdf/2106.04560.pdf), and Google research develop [contrastive captioners](https://arxiv.org/pdf/2205.01917.pdf) (CoCa). The Google Brain team also extended their text-to-image diffusion model [Imagen](https://imagen.research.google/) to the video domain with [Imagen Video](https://imagen.research.google/video/paper.pdf). DeepMind introduced a [new paradigm for self-supervised learning](https://arxiv.org/pdf/2203.08777.pdf), achieving state-of-the-art performance in a variety of transfer learning tasks. Finally, Google released [Open Images V7](https://storage.googleapis.com/openimages/web/index.html), which adds keypoint data to more than a million images. ### Amazon Amazon was prolific to say the least, with 40 papers accepted to just CVPR and ECCV. Highlighting this veritable barrage of research were a paper on [translating images into maps](https://www.amazon.science/publications/translating-images-into-maps), which won the best paper award at ICRA 2022, [a method for assessing bias in face verification systems](https://assets.amazon.science/80/92/a715a3c947a1a59ac9c7c7830ee1/unsupervised-and-semi-supervised-bias-benchmarking-in-face-recognition.pdf) without complete (or any) labels, and a systematic prescription for [modifying specific features in images generated by GANs](https://assets.amazon.science/9e/a5/34f3992c45b8a8d0f733b6857a0b/rayleigh-eigendirections-reds-nonlinear-gan-latent-space-traversals-for-multidimensional-features.pdf), which works by recasting the problem in the language of [Rayleigh quotients](https://en.wikipedia.org/wiki/Rayleigh_quotient). ### Microsoft Microsoft did a lot of work with Transformer models. It was just January when Microsoft’s paper introducing BEiT ( [BERT Pre-Training of Image Transformers](https://arxiv.org/pdf/2106.08254.pdf)) was accepted at ICLR, and the the ensuing family of models has become a staple of the Transformer model landscape, with the base model [accruing 1.4M+ downloads](https://huggingface.co/microsoft/beit-base-patch16-224-pt22k-ft22k) from Hugging Face in the past month alone. The BEiT family is blossoming, with papers on [generative vision-language pretraining](https://arxiv.org/pdf/2206.01127.pdf) (VL-BEiT), [masked image modeling with vector quantized visual tokenizers](https://arxiv.org/pdf/2208.06366.pdf) (BEiT V2), and modeling [image as a foreign language](https://arxiv.org/pdf/2208.10442v2.pdf). Beyond BEiT, Microsoft has been riding the Swin Transformer wave they created last year with [StyleSwin](https://arxiv.org/abs/2112.10762) and [Swin Transformer V2](https://arxiv.org/abs/2111.09883). Other notable works from 2022 include [MiniViT: Compressing Vision Transformers with Weight Multiplexing](https://arxiv.org/pdf/2204.07154.pdf), [RegionCLIP: Region-based Language-Image Pretraining](https://arxiv.org/pdf/2112.09106.pdf), and [NICE-SLAM: Neural Implicit Scalable Encoding for SLAM](https://arxiv.org/pdf/2112.12130.pdf). ### Meta Meta maintained a strong focus on multi-modal machine learning at the crossroads of language and vision. [Audio-visual HuBERT](https://github.com/facebookresearch/av_hubert) achieved state-of-the art results in lip reading and audio-visual speech recognition. [Visual Speech Recognition for Multiple Languages in the Wild](https://arxiv.org/pdf/2202.13084.pdf) demonstrates that adding auxiliary tasks to a Visual Speech Recognition (VSR) model can dramatically improve performance. [FLAVA: A Foundational Language And Vision Alignment Model](https://arxiv.org/pdf/2112.04482.pdf) presents a single model that performs well across 35 distinct language and vision tasks. And [data2vec](https://arxiv.org/pdf/2202.03555.pdf) introduces a unified framework for self-supervised learning that spans vision, speech, and language. With [DEiT III](https://arxiv.org/pdf/2204.07118.pdf), researchers at Meta AI revisit the training step for Vision Transformers and show that a model trained with basic data augmentation can significantly outperform fully supervised ViTs. Meta also made progress in [continual learning for reconstructing signed distance fields](https://arxiv.org/pdf/2204.02296.pdf) (SDFs), and a group of researchers including Yann LeCun shared theoretical insights into [why contrastive learning works](https://arxiv.org/pdf/2110.09348.pdf). Read this. Really. Finally, in September Meta AI spun out PyTorch into the vendor-agnostic [PyTorch Foundation](https://pytorch.org/foundation), which shortly thereafter released [PyTorch 2.0](https://pytorch.org/get-started/pytorch-2.0/). ### Adobe In 2022, Adobe took the sophisticated machinery of modern computer vision and turned it to artistic tasks of manipulation like editing, re-styling, and rearranging. [Third Time’s the Charm?](https://arxiv.org/abs/2201.13433) puts Nvidia’s StyleGAN3 to work editing images and videos, introducing a video inversion scheme that reduces [texture sticking](https://www.youtube.com/watch?v=-2hLdOonvK0). [BlobGAN](https://arxiv.org/abs/2205.02837) models scenes as collections of mid-level (between pixel-level and image-level) “blobs”, which become associated with objects in the scene without supervision, allowing for editing of scenes on the object-level. [ARF: Artistic Radiance Fields](https://arxiv.org/abs/2206.06360) accelerates the generation of artistic 3D content by combining style transfer with [neural radiance fields](https://arxiv.org/pdf/2003.08934.pdf) (NeRFs). ### Nvidia Nvidia made contributions across the board, including multiple works on performing three dimensional computer vision tasks with single-view (monocular) images and videos. [CenterPose](https://github.com/NVlabs/CenterPose) sets the standard for category-level 6 degree of freedom (DoF) pose estimation using only a single-stage network; [GLAMR](https://arxiv.org/pdf/2112.01524.pdf) globally situates humans in 3D space from videos recorded with dynamic (moving) cameras; and by separating the tasks of feature generation and neural rendering, [EG3D](https://nvlabs.github.io/eg3d/media/eg3d.pdf) can produce high-quality 3D geometry from single images. Other works of note include [GroupViT](https://arxiv.org/abs/2202.11094), [FreeSOLO](https://arxiv.org/abs/2202.12181), and ICLR spotlight paper [Tackling the Generative Learning Trilemma with Denoising Diffusion GANs](https://arxiv.org/abs/2112.07804) ## Electrifying new applications of Computer Vision ![](https://cdn.sanity.io/images/h6toihm1/production/5898b7cb4242c2f036e010a21a87e361e0e6525a-1400x943.png?auto=format&dpr=2&fit=max&q=75&w=1400) Computer vision now plays a role in everything from sports and entertainment to construction, to security, to agriculture, and within each of these industries there are far too many companies employing computer vision to count. This section highlights some of the key developments in some of the industries where computer vision is becoming deeply embedded. ### Sports Computer vision featured on the biggest of stages when [FIFA employed a semi-automated system for offsides detection](https://www.fifa.com/technical/football-technology/football-technologies-and-innovations-at-the-fifa-world-cup-2022/semi-automated-offside-technology) at the World Cup in Qatar. They also used computer vision to prevent stampedes at the stadium. Other noteworthy developments include [Sportsbox AI raising a $5.5M Series A](https://www.sportspromedia.com/news/sportsbox-ai-seed-round-golf/) led by EP Golf Ventures to bring motion tracking to golf (and other sports), and new company [Jabbr tailoring computer vision for combat sports](https://jabbr.ai/about), starting with DeepStrike, a model that automatically counts punches and edits boxing videos. ### Climate and Conservation Circular economy startup [Greyparrot](https://www.greyparrot.ai/) raised an $11M Series A round for its computer vision-driven waste monitoring system. Carbon marketplace NCX, which uses cutting edge computer vision models with satellite imagery to deliver precision assessment of timber and carbon potential, raised a $50M Series B. And Microsoft [announced the Microsoft Climate Research Initiative (MCRI)](https://www.microsoft.com/en-us/research/blog/introducing-the-microsoft-climate-research-initiative/), which will house their computer vision for climate efforts in renewable energy mapping, land cover mapping, and glacier mapping. ### Autonomous vehicles 2022 was a bit of a mixed bag for the autonomous vehicles industry as a whole, with self-driving car company [Argo AI shutting down operations](https://techcrunch.com/2022/10/26/ford-vw-backed-argo-ai-is-shutting-down/) in October, and Ford and [Rivian](https://techcrunch.com/2022/10/19/rivians-rj-scaringe-on-the-future-of-micromobility-avs-and-the-supply-chain/) shifting their focus from L4 (highly automated) to L2 (partial) and L3 (conditional) automation. Apple also recently announced that it was [scaling back its self-driving efforts](https://www.reuters.com/business/autos-transportation/apple-scale-back-self-driving-car-ambitions-delay-car-launch-2026-bloomberg-news-2022-12-06/), “Project Titan”, and pushing launch back until 2026. Nevertheless, there were some notable wins for computer vision. Researchers at MIT [released](https://news.mit.edu/2022/researchers-release-open-source-photorealistic-simulator-autonomous-driving-0621) the first open-source, photorealistic simulator for autonomous driving. Driver-assist unit [Mobileye raised an $861M IPO](https://www.cnn.com/2022/10/26/business/mobileye-ipo-intel) after spinning out of Intel. [Google acquired spatial AI and mobility startup Phiar](https://meet-global.bnext.com.tw/articles/view/47794). And [Waymo launched an autonomous vehicle service](https://www.statepress.com/article/2022/11/waymo-released-to-public-downtown-phoenix) in downtown Phoenix. ### Health and Medicine In Australia, engineers devised a promising no-contact [computer vision-based approach for blood pressure detection](https://www.sciencedaily.com/releases/2022/12/221205104242.htm), which may offer an alternative to the traditional inflatable cuffs. Additionally Google began licensing its computer vision based breast cancer detection tool to cancer detection and therapy provider [iCAD](https://www.icadmed.com/). ## Prominent Computer Vision papers you can’t pass up - [Tackling the generative learning trilemma with denoising diffusion GANs](https://arxiv.org/abs/2112.07804) - [Understanding dimensional collapse in contrastive self-supervised learning](https://arxiv.org/pdf/2110.09348.pdf) - [InternImage: Exploring large-scale vision foundation models with deformable convolutions](https://arxiv.org/pdf/2211.05778v2.pdf) - [YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors](https://arxiv.org/pdf/2207.02696.pdf) ## CV tooling startups grow in size and impact - Annotation startup [Labelbox raised a $110 Million Series D](https://www.globenewswire.com/en/news-release/2022/01/06/2362270/0/en/Labelbox-Raises-110-Million-Series-D-Led-by-SoftBank-Vision-Fund-2.html) - [V7 raised a $33M Series A](https://www.v7labs.com/news/v7-raises-33m-series-a) to help teams build robust AI - Roboflow released [Roboflow 100](https://www.rf100.org/), a new object detection benchmark - [Voxel51 raised a $12.5M Series A](https://www.prnewswire.com/news-releases/voxel51-raises-12-5m-series-a-to-bring-transparency-and-clarity-to-computer-vision-data-301629679.html) to help bring clarity and transparency to the world’s data ## Conclusion 2022 was extremely lively for machine learning, and especially so for computer vision. The crazy thing is the rapid pace of development in research, growth in number practitioners, and adoption in industry appear to be accelerating. Let’s see what 2023 has in store! ## FiftyOne Computer Vision toolset [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset developed by Voxel51 that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [AI generated artwork](https://voxel51.com/blog/tag/ai-generated-artwork) [computer vision applications](https://voxel51.com/blog/tag/computer-vision-applications) [computer vision research](https://voxel51.com/blog/tag/computer-vision-research) [computer vision trends](https://voxel51.com/blog/tag/computer-vision-trends) [data-centric computer vision](https://voxel51.com/blog/tag/data-centric-computer-vision) [multi-modal AI](https://voxel51.com/blog/tag/multi-modal-ai) [transformers](https://voxel51.com/blog/tag/transformers) [Year in Review](https://voxel51.com/blog/tag/year-in-review) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/e6ca14f73ab3cebac9a67e469fc0cb4f87f7b08b-960x640.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Making of Avatar: The Way of Water\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jan 19, 2023](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water) [![](https://cdn.sanity.io/images/h6toihm1/production/770b8cfdbd7944916b1195dc11e5b173dbab8e97-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 and the State of Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 18, 2023](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/99dc871d875865fc200f14931a3cfb3118ded624-1020x1007.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Tunnel vision in computer vision: can ChatGPT see?\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 16, 2022](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) Why 2022 was the most exciting year in computer vision history (so far) - Voxel51 <|firecrawl-page-6-lllmstxt|> ## Computer Vision in Retail [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Industry Solutions](https://voxel51.com/blog/category/industry-solutions) How Computer Vision Is Changing Retail Feb 6, 2024 • 12 min read Article content In this article [Retail Industry Overview](https://voxel51.com/blog/how-computer-vision-is-changing-retail#1144a3c8b218) [Key Industry Challenges in Retail](https://voxel51.com/blog/how-computer-vision-is-changing-retail#75bdbdf5728e) [Computer Vision Applications in Retail](https://voxel51.com/blog/how-computer-vision-is-changing-retail#c15fd149a408) [Companies at the Cutting Edge of Computer Vision in Retail](https://voxel51.com/blog/how-computer-vision-is-changing-retail#552c4b9f984a) [Retail Datasets](https://voxel51.com/blog/how-computer-vision-is-changing-retail#d16e2035034a) [Join the FiftyOne Community!](https://voxel51.com/blog/how-computer-vision-is-changing-retail#bb69874139ef) In this article [Retail Industry Overview](https://voxel51.com/blog/how-computer-vision-is-changing-retail#1144a3c8b218) [Key Industry Challenges in Retail](https://voxel51.com/blog/how-computer-vision-is-changing-retail#75bdbdf5728e) [Computer Vision Applications in Retail](https://voxel51.com/blog/how-computer-vision-is-changing-retail#c15fd149a408) [Companies at the Cutting Edge of Computer Vision in Retail](https://voxel51.com/blog/how-computer-vision-is-changing-retail#552c4b9f984a) [Retail Datasets](https://voxel51.com/blog/how-computer-vision-is-changing-retail#d16e2035034a) [Join the FiftyOne Community!](https://voxel51.com/blog/how-computer-vision-is-changing-retail#bb69874139ef) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/45ce2fbf373f5804f31dccde3ce35058a5332348-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=1600) Welcome to the fifth installment of Voxel51’s computer vision industry spotlight blog series. Each edition, we highlight how different industries – from construction to climate tech, from medicine to robotics, and more – are using computer vision, machine learning, and artificial intelligence to drive innovation. We’ll dive deep into the main computer vision tasks being put to use, current and future challenges, and companies at the forefront. In this edition, we’ll focus on retail! Read on to learn about computer vision in the retail and ecommerce industry. ## Retail Industry Overview Retail is an important pillar of modern economies, bridging the gap between producers and consumers. It's a sector that facilitates commerce and reflects and shapes societal trends and consumer preferences. Key facts and figures: - The global retail market size grew [from $26.18 trillion in 2022 to $28.34 trillion in 2023](https://www.globenewswire.com/news-release/2023/04/21/2652108/0/en/Retail-Global-Market-Report-2023.html), with a Compound Annual Growth Rate (CAGR) of 8.3%. - In the United States, retail sales were expected to grow moderately yet positively [between 4% and 6%](https://nrf.com/media-center/press-releases/nrf-forecasts-2023-retail-sales-grow-between-4-and-6) in 2023. - Ecommerce will continue to grow in popularity, with [~21% of retail purchases expected to take place online in 2023, rising to 24% by 2026](https://www.forbes.com/advisor/business/ecommerce-statistics/). - The value of AI in the retail market is estimated to be around $7.1 billion in 2023, [marking a 29% increase from the previous year](https://chisw.com/blog/ai-in-retail-2023/). Applying computer vision and AI to retail opens up a whole new world of possibilities. Recent innovations in both retail and retail technologies have set the stage for today's advancements. Thanks to vision- and AI-powered technologies, retailers can get a better read on what shoppers want, meet those needs more sustainably, and continue making shopping an enjoyable omnichannel affair. Before we dive into various popular applications of computer vision-based AI technologies in retail, it’s important to highlight the key challenges facing the industry. ## Key Industry Challenges in Retail - **Supply Chain Issues:** [Excess inventory has emerged as a significant challenge](https://hbr.org/2023/09/the-next-supply-chain-challenge-isnt-a-shortage-its-inventory-glut) globally, including extra inventory of high-tech electronics components totaling more than $250 billion in the US alone. - **Loss Prevention and Security:** The retail sector continues to experience losses due to theft and other criminal activities. [In FY 2022, the average shrink rate increased to 1.6% from 1.4% in FY 2021, amounting to $112.1 billion in losses](https://nrf.com/media-center/press-releases/shrink-accounted-over-112-billion-industry-losses-2022-according-nrf). - **Evolving Customer Expectations:** According to a report by Zendesk, [73% of consumers will switch to a competitor following multiple bad experiences](https://www.zendesk.com/in/blog/customer-expectations-meet-rising-demands/), indicating a high level of expectation throughout the [customer experience](https://voxel51.com/computer-vision-use-cases/retail/using-computer-vision-to-enhance-customer-experience-in-retail/). - **Evolving Ecommerce & Omnichannel Expectations:** Evolving consumer expectations aren’t just limited to brick-and-mortar shopping experiences; they extend to [ecommerce and omnichannel experiences](https://www.techtarget.com/searchcustomerexperience/post/Top-e-commerce-challenges-for-2023-and-how-to-overcome-them), too. - **Sustainability Initiatives:** Consumers are increasingly holding brands responsible for sustainability, [with 46% looking to brands to lead in creating sustainable change](https://nielseniq.com/global/en/insights/analysis/2022/trend-watch-2023-sustainability/). Continue reading to learn about several ways in which computer vision applications are helping organizations in the retail industry. ## Computer Vision Applications in Retail ### Inventory Management & Supply Chain Optimization ![](https://cdn.sanity.io/images/h6toihm1/production/b50b5de8b4b3ff378f975b877fe4188fca8562eb-1280x720.png?auto=format&dpr=2&fit=max&q=75&w=1280) Inventory management is all about getting the right products to customers at the right place and time, avoiding frustrating stockouts and wasteful overstocking. Computer vision is an ideal ingredient in inventory management systems due to the availability of visual data from cameras and other sensors that make it possible to monitor and analyze inventory levels in real time. Computer vision techniques at the core of AI-powered inventory management systems include object detection, object recognition, image classification, and more. Visual search is increasing in popularity, enabling consumers to search for products using images instead of (or in addition to) text. Vision-based inventory management systems bring a variety of benefits to retailers and consumers, including enhanced operational efficiency, increased customer satisfaction, cost savings, and streamlined supply chains. The automation of inventory processes also frees up valuable staff time, allowing workforces to focus on other value-added tasks and continuing to improve the overall shopping experience for customers. Computer vision is also paving the way for a new suite of AI-powered tools to help retailers optimize inventory management, including [smart shelf solutions](https://streetfightmag.com/2018/03/19/5-smart-shelf-solutions-for-retailers/), [AI route optimization](https://loconav.com/blog/ai-route-planning/) for deliveries, [store layout optimization](https://link.springer.com/article/10.1007/s10462-022-10142-3), and in-store shopper analytics. For further reading on the use of computer vision and AI technologies for automated inventory management at popular retailers, visit the following resources: - [How Walmart is using A.I. to make shopping better for its millions of customers](https://www.cnbc.com/2023/03/27/how-walmart-is-using-ai-to-make-shopping-better.html) - [American Eagle to deploy AI-based inventory tracking in stores](https://chainstoreage.com/american-eagle-deploy-ai-based-inventory-tracking-stores) Additionally, here are a few academic papers related to using computer vision for real-time inventory management: - [Product Stock Management Using Computer Vision](https://ieeexplore.ieee.org/abstract/document/9231673) - [A computer vision pipeline for automatic large-scale inventory tracking](https://dl.acm.org/doi/abs/10.1145/3409334.3452063) - [Vision Based Object Counting Using Speeded Up Robust Features for Inventory Control](https://ieeexplore.ieee.org/abstract/document/7881432) - [Banana Sub-Family Classification and Quality Prediction using Computer Vision](https://arxiv.org/pdf/2204.02581.pdf) ### Autonomous Checkouts & Smart Carts ![](https://cdn.sanity.io/images/h6toihm1/production/e2ad7ad0e77e2eec31a4efaeb883e02fef009807-960x720.png?auto=format&dpr=2&fit=max&q=75&w=960) Autonomous checkout systems and smart carts in retail utilize computer vision to deliver a fast, efficient shopping experience. By employing cameras, scanners, sensors, and object recognition concepts, these systems can instantly recognize and tally products. Shoppers can simply place their items in a designated area, and the system automatically calculates the total, facilitating a seamless and rapid checkout experience. For retailers, self-checkout systems increase checkout speeds, accuracy, and efficiency, while also reducing labor costs, which is especially important in a tight labor market. For shoppers, contactless checkout systems offer a smooth grab-and-go shopping experience, while reducing the time spent at checkout counters and enhancing overall convenience. Visit these resources for further reading on automated checkout systems at popular retailers: - [How the Amazon Go Store’s AI Works](https://towardsdatascience.com/how-the-amazon-go-store-works-a-deep-dive-3fde9d9939e9) - [Grocery smart carts aim to be the saving grace for self-checkout hate](https://www.retailbrew.com/stories/2022/08/23/grocery-smart-carts-aim-to-be-the-saving-grace-for-self-checkout-hate) Check out these papers about using computer vision for automated checkout systems: - [Automated Checkout for Stores: A Computer Vision Approach](https://www.researchgate.net/profile/Jyotsna-A/publication/353244285_Automated_Checkout_for_Stores_A_Computer_Vision_Approach/links/60ef07e99541032c6d3e79d4/Automated-Checkout-for-Stores-A-Computer-Vision-Approach.pdf) - [AI-based machine vision for retail self-checkout system](https://lup.lub.lu.se/luur/download?func=downloadFile&recordOId=8985308&fileOId=8985340) - [Enhancing Retail Checkout Through Video Inpainting, YOLOv8 Detection, and DeepSort Tracking](https://openaccess.thecvf.com/content/CVPR2023W/AICity/html/Vats_Enhancing_Retail_Checkout_Through_Video_Inpainting_YOLOv8_Detection_and_DeepSort_CVPRW_2023_paper.html) ### Virtual Try-Ons ![](https://cdn.sanity.io/images/h6toihm1/production/80495d44b2bc6039a051df9c5b5a282236959fba-1440x1080.jpg?auto=format&dpr=2&fit=max&q=75&w=1440) Virtual try-ons in retail allow customers to virtually "wear" clothing, accessories, and makeup from the comfort of their own homes using digital overlays on their images or live feeds. These systems analyze the user's physical features using pose estimation, image segmentation, and 3D modeling, and superimpose products on them, providing a realistic virtual representation of how the items would look worn in real life. Using virtual try-ons makes shopping a breeze and fun, letting folks visualize how the products would look on them without having to try things on in real life. Virtual try-on technologies reduce the number of returns, increase online engagement, and offer a competitive edge in the ecommerce landscape, as well as create unique and compelling reasons for consumers to make in-person visits to stores. Check out these resources on virtual try-on technologies at popular retailers: - [Top 6 Virtual Try-On Examples which Enhance Personal Shopping Experience](https://www.netguru.com/blog/virtual-try-on-examples) - [Walmart introduces virtual try-on tech which uses customers’ own photos to model the clothing](https://techcrunch.com/2022/09/14/walmart-introduces-virtual-try-on-tech-which-uses-customers-own-photos-to-model-the-clothing/) Here are a few papers related to using computer vision for virtual try-ons: - [VTNFP: An Image-Based Virtual Try-On Network With Body and Clothing Feature Preservation](https://openaccess.thecvf.com/content_ICCV_2019/html/Yu_VTNFP_An_Image-Based_Virtual_Try-On_Network_With_Body_and_Clothing_ICCV_2019_paper.html) - [VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization](https://arxiv.org/abs/2103.16874#:~:text=The%20task%20of%20image,warped%20item%20with%20the%20person) - [Towards Detailed Characteristic-Preserving Virtual Try-On](https://research.google/pubs/pub51586/#:~:text=Proceedings%20of%20the%20IEEE%2FCVF%20Conference,to%20faithfully%20represent%20various) ### Customer Behavior Analysis ![](https://cdn.sanity.io/images/h6toihm1/production/526abc53d7cc5cd4c99454c1d70b4ec4c41925f1-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) Retail stores can leverage existing CCTV cameras to analyze customer behavior. Based on video footage from these cameras, object tracking algorithms can track customer movements, dwell times, and interactions within the store. This provides insights into shopping patterns, popular areas, and product preferences. The benefits of understanding customer behavior are substantial. For retailers, it offers actionable insights to optimize store layouts, enhance product placements, and tailor marketing strategies. It also aids in predicting shopping trends, allowing for better inventory management. It results in a more personalized shopping experience, as stores can adjust their offerings and layouts based on observed preferences. Here are a few papers related to using computer vision for customer behavior analytics: - [Real Time Retail Analytics with Computer Vision](https://ieeexplore.ieee.org/document/9925538) - [How Computer Vision Provides Physical Retail with a Better View](https://www.researchgate.net/publication/335361426_How_Computer_Vision_Provides_Physical_Retail_with_a_Better_View_on_Customers) - [Customer Behavior Recognition in Retail Store from Surveillance Camera](https://www.researchgate.net/publication/304415146_Customer_Behavior_Recognition_in_Retail_Store_from_Surveillance_Camera) ### Product Recommendations ![](https://cdn.sanity.io/images/h6toihm1/production/c3ce5c4b8a77433c2cc0851973652ff5c26e6012-1320x743.jpg?auto=format&dpr=2&fit=max&q=75&w=1320) Computer vision enhances product recommendation systems, opening up new possibilities for customer engagement and personalization. For example, visual search adds a convenient way for consumers to discover new products and information, beyond text searches alone. A growing number of retailers, including [IKEA](https://www.ikea.com/us/en/) and [Amazon through its multimodal (image and text) search](https://techcrunch.com/2023/09/14/amazon-updates-visual-search-ar-search-and-more-in-challenge-to-google/), offer the ability for consumers to search for a product by uploading an image. Recommender systems can present items based on visual similarity to uploaded images, as well additional factors such as recent browsing history, wishlist items, and past purchases, to tailor the shopping experience to an individual consumer’s style and preferences. Check out these academic papers on computer vision in recommender systems: - [CorrEmbed: Evaluating Pre-trained Model Image Similarity Efficacy with a Novel Metric](https://arxiv.org/abs/2308.16126) - [VICTOR: Visual Incompatibility Detection with Transformers and Fashion-specific contrastive pre-training](https://arxiv.org/pdf/2207.13458v2.pdf) - [Graph Neural Networks for Social Recommendation](https://arxiv.org/pdf/1902.07243v2.pdf) ## Companies at the Cutting Edge of Computer Vision in Retail ### Trigo ![](https://cdn.sanity.io/images/h6toihm1/production/074c8e8be780d56a89eb425861ad5e1f6c50bb9e-1908x1072.png?auto=format&dpr=2&fit=max&q=75&w=1600) [Trigo](https://www.trigoretail.com/) combines ceiling-based cameras, shelf sensors, and machine vision algorithms to create a digital twin of the retail space. This digital representation allows for real-time analysis of shoppers' journeys and product choices, enabling better shopping experiences and business outcomes. This setup also enables [computer-vision-based autonomous checkout systems](https://magazine.retail-today.com/nrf_2023/trigo) that accurately identify and capture the shopping items selected by customers, and make checking out entirely automated. Trigo Vision has attracted significant investments, notably from [the German retail giant REWE Group, pushing Trigo's total fundraising to over $100 million.](https://www.israelhayom.com/2021/06/16/computer-vision-startup-trigo-raises-100m-after-investments-from-rewe-group-viola-growth/) ### Trax Retail ![](https://cdn.sanity.io/images/h6toihm1/production/ac3cf6fb7ca81b3fe39a8ae52578a392424442e3-631x381.png?auto=format&dpr=2&fit=max&q=75&w=631) [Trax Retail](https://traxretail.com/)’s mission is to enable brands and retailers to harness the power of digital technologies to produce the best shopping experiences imaginable. Trax’s retail platform allows customers to understand what is happening on-shelf, in every store, all the time so they can focus on what they do best – delighting shoppers. As pioneers in computer vision, Trax continues to lead the industry in innovation and excellence through development of advanced technologies and scalable data collection methods. Many of the [world’s top CPG companies and retailers](https://traxretail.com/customer-success-stories/) use Trax’s dynamic merchandising, in-store execution, shopper engagement, market measurement, analytics, and shelf monitoring solutions at scale to drive positive shopper experiences and unlock revenue opportunities at all points of sale. ### Link Retail ![](https://cdn.sanity.io/images/h6toihm1/production/c78eb00a8c181aa59367bbff838895cd0b8ee37b-1283x1622.png?auto=format&dpr=2&fit=max&q=75&w=1283) [Link Retail](https://linkretail.com/), based in Oslo, Norway, uses AI powered techniques to help brick-and-mortar retailers strategically boost sales and optimize operations. The company builds a variety of solutions, including: - Food Waste Management: AI software that optimizes grocery product ordering procedures and reduces retail food waste - Video Analytics: A high accuracy footfall counting system that turns in-store CCTV camera footage into rich operational and shopping data, such as real-time occupancy analysis, shopper flow, and queue analytics - Space Management: An AI analytics tool that employs Point of Sales (POS) data and generates actionable insights on optimizing retail space including floor, shelf, and sales activities Link Retail helps retailers navigate the challenges of the physical retail environment, making strides toward creating data-driven retail spaces. ### Dayta AI ![](https://cdn.sanity.io/images/h6toihm1/production/ef407ccbb9fb3091ffd2db9b2e4319d6ba9f04f6-640x426.png?auto=format&dpr=2&fit=max&q=75&w=640) [Dayta AI](https://www.dayta.ai/) is a retail analytics Software as a Service (SaaS) company that uses computer vision and AI to turn camera footage from retail stores into useful insights. Dayta AI's solution, [Cyclops](https://www.dayta.ai/why-cyclops), is engineered to work with any RTSP-supported cameras, allowing retailers to use their existing video cameras without additional equipment costs. Cyclops can monitor and analyze customer traffic, zone-specific activities, footfall, engagement count, dwell time, queue time, and even emotions, among other metrics. These data points help retailers understand customer behavior, optimize store layouts, and improve the overall customer experience. ### RetailNext ![](https://cdn.sanity.io/images/h6toihm1/production/c8175bb9c1e4d5656269c578f4d5a86cfb9baa49-800x500.png?auto=format&dpr=2&fit=max&q=75&w=800) [RetailNext](https://retailnext.net/) was founded to address challenges faced by modern retailers and bring e-commerce style shopper analytics to brick-and-mortar stores, brands, and malls. Through its centralized SaaS platform, RetailNext automatically collects and analyzes shopper behavior data, providing retailers with the critical insights they need to improve the shopper experience in real time. This platform helps retailers optimize store operations, store layouts and marketing strategies, and improve customer experiences​​. The company also offers a next-generation IoT sensor, [Aurora](https://retailnext.net/product/aurora), which is powered by an patented algorithm that uses 3D imagery and deep learning. RetailNext’s technology is trusted by 400+ top retailers and brands globally. ## Retail Datasets If you are interested in exploring applications of computer vision in retail, check out these datasets: - [RPC: A Large-Scale and Fine-Grained Retail Product Checkout Dataset](https://rpc-dataset.github.io/): This dataset is designed to advance automatic checkout solutions. It is a collection of 53,739 single-product images in the training set and a combined total of 30,000 multi-product images in the validation and test sets, categorized finely across various product types. Explore the validation set of [6,000 checkout images with FiftyOne in your browser](https://try.fiftyone.ai/datasets/retail-product-checkout/samples). - [Supermarket Shelves Dataset](https://humansintheloop.org/resources/datasets/supermarket-shelves-dataset/): This dataset can be used to enhance product detection on supermarket shelves. It encompasses 45 images from global supermarkets, amounting to 11,743 bounding boxes, averaging 260 boxes per image. - [MERL Shopping](https://paperswithcode.com/dataset/merl-shopping): This dataset contains 106 videos of around 2 minutes each from an overhead camera in a grocery setting. It highlights five actions: "Reach To Shelf," "Retract From Shelf," "Hand In Shelf," "Inspect Product," and "Inspect Shelf." - [Zalando Fashion MNIST:](https://www.kaggle.com/datasets/zalando-research/fashionmnist) This dataset consists of 70,000 28x28 grayscale images of fashion products categorized into 10 classes like T-shirt/top, trouser, pullover, etc. - [GroZi-120](http://grozi.calit2.net/grozi.html): This dataset comprises 120 grocery products captured in both isolated (in vitro) and real-world settings (in situ). - [Clothing Coparsing Dataset](https://github.com/bearpaw/clothing-co-parsing): This dataset shows the semantic segmentation of different outfits from multiple street fashion models. [Explore this dataset with FiftyOne in your browser](https://try.fiftyone.ai/datasets/clothing-segmentation/samples). - [Fashion Product Images](https://try.fiftyone.ai/datasets/fashion-product-images/samples): This dataset is a fantastic ecommerce dataset that stores tons of metadata about each article of clothing. The dataset was built for recommender systems that want to take in many different features. [Explore this dataset with FiftyOne in your browser](https://try.fiftyone.ai/datasets/fashion-product-images/samples). If you would like to see any of these or other computer vision retail datasets added to the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html), get in touch, and we can work together to make this happen! ## Join the FiftyOne Community! Developers of retail applications can benefit from FiftyOne’s ability to easily filter through the huge amounts of visual data collected daily from stores and other sources. Using [open source FiftyOne](https://github.com/voxel51/fiftyone), this data can be curated into datasets for model training or to share with experts for annotation or analysis of CV models. Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 2,400+ [FiftyOne Slack](https://slack.voxel51.com/) members - 6,350+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 18,000+ [CV Meetup](https://www.meetup.com/pro/computer-vision-meetups/) & [AI Meetup](https://www.meetup.com/pro/ai-machine-learning-data-science-network/) members - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 470+ repositories - 80+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [autonomous checkout](https://voxel51.com/blog/tag/autonomous-checkout) [customer behavior analysis](https://voxel51.com/blog/tag/customer-behavior-analysis) [industry use cases](https://voxel51.com/blog/tag/industry-use-cases) [inventory management](https://voxel51.com/blog/tag/inventory-management) [product recommendations](https://voxel51.com/blog/tag/product-recommendations) [recommender engines](https://voxel51.com/blog/tag/recommender-engines) [retail](https://voxel51.com/blog/tag/retail) [retail industry](https://voxel51.com/blog/tag/retail-industry) [retail use case](https://voxel51.com/blog/tag/retail-use-case) [supply chain](https://voxel51.com/blog/tag/supply-chain) [virtual try-ons](https://voxel51.com/blog/tag/virtual-try-ons) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/b4fd054bf7574e4060ffc7a9a8f201c31a66c327-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Using Computer Vision to Enhance Customer Experience in Retail\\ \\ Learn\\ \\ • \\ \\ Jan 23, 2025](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail) [![](https://cdn.sanity.io/images/h6toihm1/production/bbb1d9add0b0b9aa12682acac795df7c2ba760a9-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — November 2022\\ \\ Event Recaps\\ \\ • \\ \\ Nov 16, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/d94b00bbfadfd5cf59c4debb3d48da6eea3a3a4f-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Computer Vision Is Changing Sports\\ \\ Computer Vision, Industry Solutions, Product & News\\ \\ • \\ \\ Jan 16, 2024](https://voxel51.com/blog/how-computer-vision-is-changing-sports) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-7-lllmstxt|> ## CVPR 2024 Benchmarks Overview [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) CVPR 2024 Datasets and Benchmarks – Part 2: Benchmarks Apr 30, 2024 • 18 min read Article content In this article [ImageNet-D](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks#5c04f0900b9d) [Polaris](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks#88795ff3627c) [VBench](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks#c2ab96f5ed92) [Conclusion](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks#a7a7aba3adbf) [Visit Voxel51 at CVPR 2024!](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks#2e9641d6657c) In this article [ImageNet-D](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks#5c04f0900b9d) [Polaris](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks#88795ff3627c) [VBench](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks#c2ab96f5ed92) [Conclusion](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks#a7a7aba3adbf) [Visit Voxel51 at CVPR 2024!](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks#2e9641d6657c) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) In [part one of this series](https://docs.google.com/document/d/15QM8Iu-0VUzJmWcMlNV1kzRYmZrnKEBJdFqso-EUchA/edit?usp=sharing), I explored some interesting datasets presented at CVPR 2024, highlighting how they’ll help advance computer vision and deep learning. Now, it's time to turn our attention to the other side of the coin: benchmarks. Just as musicians need stages to showcase their talent, deep learning models need benchmarks to demonstrate their capabilities and push the boundaries of what's possible. These standardized tasks and challenges provide a crucial yardstick for evaluating and comparing different models, driving healthy competition and accelerating progress. CVPR 2024 has once again delivered a collection of innovative benchmarks that address existing limitations and explore new frontiers in computer vision. In this second part of the series, I’ll highlight three benchmarks I found interesting: - **ImageNet-D:** Testing the robustness of image classifiers against real-world perturbations. - **Polaris:** Assessing the ability of vision-language models to follow natural language instructions in interactive environments. - **VBench:** Comprehensive Benchmark Suite for Video Generative Models Each of these benchmarks presents unique challenges and opportunities for researchers, pushing the field towards more robust models. In the following sections, I'll focus on the following aspects of each dataset: **Task and Objective:** Clearly define the specific task or problem the benchmark evaluates. **Dataset and Evaluation Metric:** Provide details about the benchmark, including its size, composition, and the evaluation metrics employed to measure model performance. **Benchmark Design and Protocol:** Explain the benchmark's design and protocol, including how the dataset is split into training, validation, and test sets. **Comparison to Existing Benchmarks:** Compare the new benchmark to existing ones in the same domain, highlighting its unique challenges, evaluation criteria, and/or how the benchmark complements or improves upon existing benchmarks. **State-of-the-Art Results:** Showcase the leaderboard on the benchmark, if it exists, and what the top-performing models are. If available, discuss the model’s key architectural features or training strategies. **Impact and Future Directions:** Discuss the benchmark's potential impact, how it can drive research in new directions, and address important challenges in existing benchmarks. # ImageNet-D ![](https://cdn.sanity.io/images/h6toihm1/production/4734e7c9f026949888a0c2849f21fa82f38e3001-1600x712.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### tl;dr - **Task:** Object recognition on synthetic images - **Metric:** Top-1 Accuracy - [Paper](https://arxiv.org/html/2403.18775v1) - [GitHub](https://github.com/chenshuang-zhang/imagenet_d) - [Dataset on Hugging Face](https://huggingface.co/datasets/zcs15/ImageNet-D) ### Task and Domain The [ImageNet-D](https://chenshuang-zhang.github.io/imagenet_d/) benchmark evaluates the robustness of neural networks in object recognition tasks using synthetic images generated by diffusion models. - It assesses the performance of various vision models, ranging from standard visual classifiers to foundation models like CLIP and MiniGPT-4. - The primary objective is to rigorously test the robustness of these models in correctly identifying objects under challenging conditions. - The benchmark focuses explicitly on "hard" images designed to test the models' perception abilities. - Using synthetic images generated by diffusion models, ImageNet-D provides a rigorous evaluation of how well neural networks can handle variations in object representation. ### Dataset Curation, Size, and Composition The synthetic images were generated using Stable Diffusion models steered by language prompts. They tested the robustness of visual recognition systems by using diverse backgrounds, textures, and materials to challenge the models' perception capabilities. - It comprises 4,835 challenging images across 113 overlapping categories between ImageNet and ObjectNet. - The images feature a diverse array of backgrounds (3,764), textures (498), and materials (573) to push the limits of object recognition models. - The dataset is generated by pairing each object with 547 nuisance candidates from the Broden dataset, resulting in various realistic and challenging synthetic images. - The primary evaluation metric is top-1 accuracy in object recognition, which measures the proportion of correctly classified images. - Compared to standard datasets, ImageNet-D proves to be significantly more challenging, as evidenced by the notable drop in accuracy percentages for various state-of-the-art models. ### Benchmark Design and Protocol ![](https://cdn.sanity.io/images/h6toihm1/production/f16d725bef4d1435407526e180e180e63959056c-1600x450.png?auto=format&dpr=2&fit=max&q=75&w=1600) The benchmark's construction follows a rigorous process that involves image generation, labeling, hard image mining, human verification, and quality control: - The image generation process is formulated as Image(C, N) = Stable Diffusion(Prompt(C, N)), where C and N refer to the object category and nuisance, respectively. - The nuisance N includes background, material, and texture. - For example, images of backpacks with various backgrounds, materials, and textures are generated, offering a broader range of combinations than existing test sets. - Each generated image is labeled with its prompt category C as the ground truth for classification. - An image is misclassified if the model's predicted label does not match the ground truth C. - After creating a large image pool with all object categories and nuisance pairs, the CLIP (ViT-L/14) model is evaluated on these images. - Hard images are selected based on shared perception failure, defined as an image that leads multiple models to predict the object's label incorrectly. The test set is constructed using shared failures of known surrogate models. - The test set is challenging if these failures lead to low accuracy in unknown models. This property is called transferable failure. - **Human Labeling:** Since ImageNet-D includes images with diverse object and nuisance pairs that may be rare in the real world, human labeling is performed using Amazon Mechanical Turk (MTurk). Workers are asked to answer two questions for each image: - _Can you recognize the desired object (\[ground truth category\]) in the image?_ - _Can the object in the image be used as the desired object (\[ground truth category\])?_ - To ensure workers understand the labeling criteria, they are asked to label two example images for practice, providing the correct answers. After the practice session, workers must label up to 20 images in one task, answering both questions for each image by selecting 'yes' or 'no'. - Sentinels ensure high-quality annotations. Workers' annotations are removed if they fail to select the correct answers for positive sentinels, select 'yes' for negative sentinels, or provide inconsistent answers for consistent sentinels. - Positive sentinels are images that belong to the desired category and are correctly classified by multiple models. - Negative sentinels are images that do not belong to the desired category. - Consistent sentinels are images that appear twice in a random order. ### Comparison to Existing Benchmarks The big difference with [ImageNet-D](https://chenshuang-zhang.github.io/imagenet_d/) is that it creates entirely new. synthetic images. Here’s how it’s different from existing benchmarks that try to do the same thing: - Unlike ObjectNet, which collects real-world object images with controlled factors like background, or ImageNet-C, which introduces low-level visual corruptions, ImageNet-D generates entirely new images with diverse backgrounds, textures, and materials. - While ImageNet-9 combines foreground and background from different images, it is limited by poor image fidelity. Similarly, Stylized-ImageNet alters the textures of ImageNet images but cannot control global factors like backgrounds. In contrast, ImageNet-D allows for specific control over the image space, which is crucial for robustness benchmarks. - Compared to DREAM-OOD, which finds outliers by decoding sampled latent embeddings to images but lacks control over the image space, ImageNet-D focuses on hard images with a single attribute. - By generating new images and mining the most challenging ones as the test set, ImageNet-D achieves a greater accuracy drop compared to methods that modify existing datasets. The results show that ImageNet-D causes a significant accuracy drop, up to 60%, in a range of vision models, from standard visual classifiers to the latest foundation models like CLIP and MiniGPT-4. - The approach utilized in ImageNet-D demonstrates the potential for using generative models to evaluate model robustness, and its effectiveness is expected to grow further with advancements in generative models. ### State-of-the-Art Results [ImageNet-D](https://chenshuang-zhang.github.io/imagenet_d/) is a challenging benchmark for various state-of-the-art models, causing significant drops in their object recognition accuracy. Here are some key findings: - CLIP experiences a substantial accuracy reduction of 46.05% on ImageNet-D compared to its performance on ImageNet. - LLaVa's accuracy drops by 29.67% when evaluated on the ImageNet-D benchmark. - Despite being a more recent model, MiniGPT -4 still has a 16.81% decrease in accuracy on ImageNet-D. - All tested models show an accuracy drop of more than 16% on ImageNet-D compared to their performance on the standard ImageNet dataset. - Even the latest models, such as LLaVa-1.5 and LLaVa-NeXT, are not immune to the challenges posed by ImageNet-D, experiencing significant accuracy drops. ### Impact and Future Directions [ImageNet-D](https://chenshuang-zhang.github.io/imagenet_d/) demonstrates the effectiveness of using generative models to evaluate the robustness of neural networks. The authors suggest that their approach is general and has the potential for greater effectiveness as generative models improve. They aim to create more diverse and challenging test images in the future by capitalizing on advancements in generative models. # Polaris ![](https://cdn.sanity.io/images/h6toihm1/production/d47442cf0bcbe1bb708369c37287825a74bcfe6d-1548x1600.png?auto=format&dpr=2&fit=max&q=75&w=1548) ### tl;dr - **Task:** Image Captioning - **Metric:** Polos, based on the novel Multimodal Metric Learning from Human Feedback (M²LHF) framework - [Project Page](https://yuiga.dev/polos/) - [Paper](https://ar5iv.labs.arxiv.org/html/2402.18091) - [GitHub](https://github.com/keio-smilab24/polos?tab=readme-ov-file) - [Dataset on Hugging Face](https://huggingface.co/datasets/yuwd/Polaris) This paper was interesting because it introduces [Polaris](https://yuiga.dev/polos/), a new large-scale benchmark dataset for evaluating image captioning models, and Polos, a state-of-the-art (SOTA) metric trained on this dataset. Let’s quickly discuss the concepts of _benchmarks_ and _metrics._ A benchmark is a standardized dataset or suite of datasets used to evaluate and compare the performance of different models or algorithms on a specific task. In image captioning, a benchmark typically consists of images, associated human-written captions, and human judgments of caption quality for a subset of the data. The benchmark provides a common ground for comparing different captioning models or evaluation metrics. A metric, on the other hand, is a method or function used to measure a model's performance on a specific task. In image captioning, a metric takes an image, a candidate caption, and possibly one or more reference captions as input. It outputs a score indicating the quality of the candidate caption. The metric's performance is evaluated by measuring how well its scores correlate with human judgments on a benchmark dataset. ### Task and Objective Automatic evaluation of image captioning models is essential for accelerating progress in image captioning, as it enables researchers to quickly and objectively compare different models and architectures without the need for time-consuming and expensive human evaluations. This research aims to create an effective evaluation metric, Polos, designed explicitly for image captioning models, that closely mirrors human judgment of image caption quality. The main goal is the development of a metric that closely aligns with human judgments of caption quality, fluency, relevance, and descriptiveness. This paper closely intertwines the development of the Polos metric and the Polaris benchmark dataset. The authors introduce the Multimodal Metric Learning from Human Feedback (M²LHF) framework, which is used to develop the Polos metric. ### Dataset and Evaluation Metric ![](https://cdn.sanity.io/images/h6toihm1/production/c76d04a0f91b25113b318d7edd5ac5519bec69a6-1026x1406.png?auto=format&dpr=2&fit=max&q=75&w=1026) Evaluating image captioning models accurately requires metrics that align with human judgment. However, existing datasets often lack the scale and diversity needed to train such metrics effectively. This paper addresses this challenge by introducing the Polaris dataset and the Polos evaluation metric. #### Dataset - The Polaris dataset includes 13,691 images and 131,020 generated captions from 10 diverse image captioning models, providing a wide range of caption quality and style. The images used in Polaris are drawn from the MS-COCO and nocaps datasets, chosen for their widespread use in image captioning tasks and their diverse range of image content Additionally, it contains 262,040 human-written reference captions, which serve as a gold standard for comparison. - The Polaris dataset is split into training (78,631 samples), validation (26,269 samples), and test sets (26,123 samples). - The generated captions in the Polaris dataset encompass 3,154 unique words, totalling 1,177,512 words. On average, each generated caption is composed of 8.99 words. - The reference captions have a vocabulary of 22,275 unique words and a word count of 8,309,300. On average, each reference caption consists of 10.7 words. - The authors collected 131,020 human judgments from 550 evaluators on the image-caption pairs in the Polaris dataset to obtain a comprehensive assessment of caption quality. - Human evaluators rated each caption on a 5-point scale, considering factors such as fluency, image relevance, and detail level. These ratings were then normalized to a range of \[0, 1\] to facilitate comparison and evaluation. #### Metric - Polos uses a parallel feature extraction mechanism that combines features from the CLIP model, which captures image-text similarity, and a RoBERTa model pretrained with SimCSE, which provides high-quality textual representations. - The extracted features are then passed through a multilayer perceptron (MLP) to predict the human evaluation score. - The primary objective of the Polos metric is to achieve a high correlation with human judgments, demonstrating its ability to assess caption quality accurately. - The effectiveness of Polos is quantified using Kendall's Tau correlation coefficients, specifically Tau-b for the Flickr8K-CF dataset and Tau-c for other datasets. These coefficients measure the alignment between the rankings produced by the Polos metric and those derived from human judgments, with a higher correlation indicating better performance. - The authors also introduce the Multimodal Metric Learning from Human Feedback (M²LHF) framework, a general approach for developing metrics that learn from human judgments on multimodal inputs. ### Benchmark Design and Protocol - Each caption in Polaris received an average of eight judgments from different evaluators. This novel approach aims to closely mimic human judgment by considering various aspects of caption quality, providing a more comprehensive and reliable assessment of caption quality compared to datasets with fewer judgments per caption. - Evaluators rated captions on a 5-point scale based on fluency, image relevance, and detail level. These ratings were then normalized to a \[0, 1\] range for training the Polos metric. - The proposed Polos metric integrates similarity-based and learning-based methods to evaluate the quality of captions. By leveraging the large-scale Polaris dataset for training and evaluation, the authors ensure the Polos metric is robust and generalizable across different image captioning models and datasets. ### Comparison to Existing Benchmarks Unlike traditional datasets, Polaris includes a much larger volume of human judgments and evaluations from a diverse set of evaluators. This benchmark addresses the gap in existing metrics - poor correlation with human judgment. - Standard datasets commonly used to evaluate image captioning include Flickr8K-Expert, Flickr8K-CF, Composite, and PASCAL-50S. - Flickr8K-Expert and Flickr8K-CF datasets: - Comprise a significant amount of human judgments on captions provided by humans. - Do not contain any captions generated by models, which presents an issue from the perspective of the domain gap when using them for training metrics. - Composite dataset: - Contains 12K human judgments across images collected from MSCOCO, Flickr8k, and Flickr30k. - Although each image initially contains five references, only one was selected for human judgments within the dataset. - CapEval1k dataset: - Introduced for training automatic evaluation metrics. - Has several limitations: it is a closed dataset, uses outdated models, and includes only 1K human judgments. ### State-of-the-Art Results The Polos metric performs state-of-the-art on the Polaris benchmark and the Composite, Flickr8K (Expert and CF), PASCAL-50S and FOIL benchmarks. It outperforms the previous best metric RefPAC-S by decent margins ranging from 0.2 to 1.8 Kendall's Tau points on the different sets. These results demonstrate the benefit of large-scale supervised training and improved text representations compared to unsupervised methods relying on CLIP alone. However, substantial room for improvement is still needed to close the gap with human-level consistency. ### Impact and Future Directions The Polaris dataset and Polos metric aim to spur the development of more accurate automatic evaluation methods for image captioning. A reliable automatic metric that correlates well with human judgment is crucial for accelerating progress in this area, as it enables faster iteration and comparison of captioning models. The authors note some limitations and future directions. Polos tends to overemphasize identifying the most noticeable objects while missing fine-grained details and contextual information. Techniques like RegionCLIP could potentially help improve its fine-grained alignment capabilities. Another direction is to extend the dataset with a harsher or multi-step scoring system. Overall, this work takes important steps toward a more discriminative and human-like evaluation of image captioning models. # VBench ![](https://cdn.sanity.io/images/h6toihm1/production/c1f5871ba7bb08e4dad4943f5f4208af826c87e2-1013x516.png?auto=format&dpr=2&fit=max&q=75&w=1013) ### tl;dr - **Task:** Text-to-Vide (T2V) Generation - **Metric:** Instead of reducing the qality of video generation to one number, this benchmark looks at 16 dimensions in video generation, with fine-grained levels that reveal a models strengths and weaknessed - [Project Page](https://vchitect.github.io/VBench-project/) - [Paper](https://ar5iv.labs.arxiv.org/html/2311.17982) - [GitHub](https://github.com/Vchitect/VBench) - [Leaderboard on Hugging Face](https://huggingface.co/spaces/Vchitect/VBench_Leaderboard) ### Task and Objective This benchmark evaluates the performance of text-to-video (T2V) models. It assesses the quality and consistency of generated videos across multiple dimensions and content categories. - Provides a standardized and reliable framework for comparing different T2V models - Evaluates video quality, consistency, and fulfillment of conditions ### Dataset and Evaluation Metric VBench introduces a diverse and carefully curated prompt suite to evaluate T2V models. The benchmark employs a multi-dimensional evaluation framework with human preference annotations. - Prompt suite per dimension: - Contains around 100 prompts for each of the 16 evaluation dimensions - Evaluation dimensions: Subject Consistency, Background Consistency, Temporal Flickering, Motion Smoothness, Dynamic Degree, Object Class, Human Action, Color, Spatial Relationship, Scene, Appearance Style, and Temporal Style - Prompt suite per category: - Covers various content domains, such as animals, objects, humans, and scenes - Evaluation perspectives: Video Quality, Video-Condition Consistency, and Fulfillment of Conditions - Human preference annotations: collected through a data preparation procedure involving video generation, video pair sampling, annotation interface design, and data quality control measures ### Benchmark Design and Protocol The authors laid out an extensive procedure for data preparation for human preference annotations, as they hope to capture a models alignment with human perception. Data preparation procedure: - **Video Generation:** T2V models generate videos based on the prompts in the VBench prompt suite. - **Video Pair Sampling:** Pairs of generated videos sharing the same input prompt are sampled for each evaluation dimension. - **Annotation Interface Design:** A user-friendly interface allows human annotators to compare and express preferences between video pairs. - **Annotation Collection:** Human annotators review video pairs and select their preferred video based on the specified evaluation dimension. - **Data Quality Control:** Quality control measures, such as attention checks and consistency checks, are implemented to ensure the reliability of the collected preference data. - **Dimension-Specific Focus:** Annotators focus solely on the specific evaluation dimension being assessed when expressing their preference. - **Comparative Judgment:** Annotators judge between video pairs rather than providing absolute ratings. - **Attention to Detail:** Annotators pay close attention to relevant details and nuances in the generated videos. - **Neutral and Unbiased Assessment:** Annotators maintain a neutral and unbiased perspective when comparing and selecting preferred videos. - **Consistency and Reproducibility:** Annotators consistently apply the same criteria and standards across different video pairs and evaluation dimensions. - **Handling Ambiguous or Challenging Cases:** Annotators make their best judgment based on available information and specific criteria, with the option to indicate a "tie" or "unsure" response if necessary. ### Comparison to Existing Benchmarks ![](https://cdn.sanity.io/images/h6toihm1/production/5435ec85bb08b174d3934f26c7d391c9b0bcabc8-453x546.png?auto=format&dpr=2&fit=max&q=75&w=453) The paper does not explicitly compare it to existing benchmarks, they do discuss existing metrics that are commonly used for evaluating video generation models and their limitations: - Inception Score (IS) - Fréchet Inception Distance (FID) - Fréchet Video Distance (FVD) - CLIPSIM The main issues with these metrics are: - **Inconsistency with human judgment:** The paper states that these existing metrics "are inconsistent with human judgement" when evaluating the quality of generated videos. - **Lack of diversity and specificity in prompts:** The prompts used for these metrics, such as class labels from UCF-101 dataset for IS, FID, and FVD, and human-labeled video captions from MSR-VTT for CLIPSIM, "lack diversity and specificity, limiting accurate and fine-grained evaluation of video generation." - **Not tailored to the unique challenges of video generation:** The paper also mentions that generic Video Quality Assessment (VQA) methods "are primarily designed for real videos, thereby neglecting the unique challenges posed by generative models, such as artifacts in synthesized videos." - **Oversimplification of evaluation:** Existing metrics often reduce video generation model performance to a single number, which "oversimplifies the evaluation" and fails to provide insights into individual models' specific strengths and weaknesses - VBench aims to address these limitations by providing a comprehensive and standardized evaluation framework that aligns with human perception and captures the nuanced aspects of video quality and consistency. ### State-of-the-Art Results The leaderboard is on [Hugging Face](https://huggingface.co/spaces/Vchitect/VBench_Leaderboard). \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ### Impact and Future Directions In my opinion, VBench has the potential to significantly impact the field of text-to-video generation and shape its future directions. By providing a comprehensive and standardized evaluation framework, VBench addresses the limitations of existing metrics and offers a more reliable and insightful way to assess the performance of T2V models. One of the key strengths of VBench is its alignment with human perception. By incorporating human preference annotations and considering multiple evaluation dimensions, VBench captures the nuanced aspects of video quality and consistency important to human viewers. This human-centric approach ensures the benchmark reflects real-world expectations and requirements for generated videos. By assessing performance across various dimensions and content categories, researchers can identify areas where models excel and where improvements are needed. This granular analysis facilitates targeted research efforts and drives the development of more advanced and specialized T2V models. The insights provided by VBench, such as the trade-offs between temporal consistency and dynamic degree, highlight the challenges and opportunities in the field. These findings can guide researchers in developing techniques to mitigate trade-offs and optimize model performance across different dimensions. The identification of hidden potential in specific content categories encourages the exploration of specialized models tailored to particular domains, such as humans or animals. ## Conclusion CVPR 2024 has once again demonstrated the incredible pace of innovation in computer vision and deep learning. From robust image classifiers tested against real-world perturbations to vision-language models navigating interactive environments, the benchmarks presented this year are pushing the boundaries of what's possible. ImageNet-D, Polaris, and VBench each offer unique challenges and opportunities for researchers, driving the development of more robust, versatile, and human-aligned models. As these benchmarks continue to evolve and inspire new research directions, we can expect even more groundbreaking advancements in the field of computer vision. The future is bright, and I, for one, am excited to see what incredible innovations emerge next! ## Visit Voxel51 at CVPR 2024! And if you're attending CVPR 2024, don't forget to stop by the Voxel51 booth #1519 to connect with the team, discuss the latest in computer vision and NLP, and grab some of the most sought-after swag at the event. See you there! [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [CVPR](https://voxel51.com/blog/tag/cvpr) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Loading related posts... [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-8-lllmstxt|> ## FiftyOne Integrations Overview [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/37d1c19b637336d3c51085364fa43eb57e5c809b-3024x640.png?auto=format&dpr=2&fit=max&q=75&w=1512) # Integrations FiftyOne integrates with the AI stack you love and use. Category Vector Search Cloud + Deployments Models SDK + Notebooks Datasets Annotation Tools [![](https://cdn.sanity.io/images/h6toihm1/production/0a8a0fa5359280ef4e395c36e52009ad1b38e06a-500x500.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ ActivityNet\\ \\ Use FiftyOne to download, visualize, and evaluate ActivityNet dataset, a large-scale video benchmark for human activity understanding.](https://docs.voxel51.com/integrations/activitynet.html) [![](https://cdn.sanity.io/images/h6toihm1/production/250dbb8408bc1a7c13370ccb0acd323073c990a9-2160x2160.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Albumentations\\ \\ Use Albumentations transformation pipelines with FiftyOne to apply and visualize augmentations, and test their effectiveness on your data.](https://docs.voxel51.com/integrations/albumentations.html) [![](https://cdn.sanity.io/images/h6toihm1/production/854e37f0c96ac7f1ec289398843b004294237feb-500x500.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Build & Deploy\\ \\ Deploy FiftyOne in your enterprise via Docker or Kubernetes.](https://docs.voxel51.com/enterprise/installation.html) [![](https://cdn.sanity.io/images/h6toihm1/production/c1cd696907bab290c1cca8063c674880e5875995-500x500.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ COCO\\ \\ Download, visualize, and evaluate the COCO dataset, a large-scale object detection, segmentation, and captioning dataset, directly in FiftyOne.](https://docs.voxel51.com/integrations/coco.html) [![](https://cdn.sanity.io/images/h6toihm1/production/9d11cb7ee2aa6a09ac3d15311af07ff7311de232-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ CVAT\\ \\ CVAT is a popular open-source image and video annotation tool. Upload data directly from FiftyOne to CVAT to add or edit labels.](https://docs.voxel51.com/integrations/cvat.html) [![](https://cdn.sanity.io/images/h6toihm1/production/199d32b9241cfc158793211af7a04d4e63c2661d-24x24.svg)\\ \\ Cloud\\ \\ Work with datasets residing in the public cloud (e.g. AWS, Google Cloud, Microsoft Azure …) or your own private cloud.](https://docs.voxel51.com/enterprise/cloud_media.html) [![](https://cdn.sanity.io/images/h6toihm1/production/e46488f3cd5781f1f4d2acecf898c8461e92b8f6-668x694.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=48&q=75&w=48)\\ \\ Detectron2\\ \\ Use FiftyOne datasets to train a model with Detectron2, Facebook AI Research’s library for detection and segmentation algorithms.](https://docs.voxel51.com/tutorials/detectron2.html) [![](https://cdn.sanity.io/images/h6toihm1/production/d0b532d947349018840d9236105d3585164609e3-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Elasticsearch\\ \\ Elasticsearch’s vector database offers an efficient way to create, store, and search vector embeddings at scale. Create Elasticsearch indexes, upload vectors, and run similarity queries programmatically and via the FiftyOne App.](https://docs.voxel51.com/integrations/elasticsearch.html) [![](https://cdn.sanity.io/images/h6toihm1/production/9a7a5994f17587dba5e0d18328fcfcf9a14e115b-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Hugging Face\\ \\ Run inference with Hugging Face Transformers models on your FiftyOne datasets, as well as directly push and load datasets from Hugging Face Hub.](https://docs.voxel51.com/integrations/huggingface.html) [![](https://cdn.sanity.io/images/h6toihm1/production/006b38e240590df1f23b0ddbfcf621456dca3e23-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Label Studio\\ \\ Label Studio is a data labeling tool supporting images, video, text, and other modalities. Upload data directly from FiftyOne to Label Studio for labeling.](https://docs.voxel51.com/integrations/labelstudio.html) [![](https://cdn.sanity.io/images/h6toihm1/production/8d6f72fe14f6d32a1f535c8568717f9ab00887d6-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Labelbox\\ \\ Labelbox is a cloud-based image and video annotation tool. Upload data directly from FiftyOne to Labelbox for labeling.](https://docs.voxel51.com/integrations/labelbox.html) [![](https://cdn.sanity.io/images/h6toihm1/production/660ac9e3ba562e7d930375ccf15eec7ee5695212-385x385.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ LanceDB\\ \\ LanceDB is a serverless vector database with deep integrations with the Python ecosystem. Create LanceDB tables and run similarity queries programmatically or via the FiftyOne App.](https://docs.voxel51.com/integrations/lancedb.html) [![](https://cdn.sanity.io/images/h6toihm1/production/616c2244d71f24e958f0ef41bfb71cd875769a0c-500x500.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Lightning Flash\\ \\ Train Flash models on FiftyOne datasets and improve your models](https://docs.voxel51.com/integrations/lightning_flash.html) [![](https://cdn.sanity.io/images/h6toihm1/production/2fa8096400e3f6ecec8b673aeda6e9624abcd361-1501x1501.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Milvus\\ \\ Milvus makes unstructured data search more accessible and provides a consistent user experience. Create Milvus collections, upload vectors, and run similarity queries programmatically and via the FiftyOne App.](https://docs.voxel51.com/integrations/milvus.html) [![](https://cdn.sanity.io/images/h6toihm1/production/8f68ed1664e58d56191b335399e4f8d4826b7eeb-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ MongoDB\\ \\ MongoDB
Atlas Vector Search enables building semantic search and AI-powered applications. Create MongoDB vector search indexes, add/remove vectors, and run similarity queries programmatically and via the FiftyOne App.](https://docs.voxel51.com/integrations/mongodb.html) [![](https://cdn.sanity.io/images/h6toihm1/production/0f7e8821fa3f304d0086858d8fe06b2d8957fb70-24x24.svg)\\ \\ Notebooks\\ \\ FiftyOne supports web-based or cloud interactive development environments such as Jupyter Notebook, Google Colab, Databricks Notebook, SageMaker Notebook, and more.](https://docs.voxel51.com/environments/index.html) [![](https://cdn.sanity.io/images/h6toihm1/production/3cfd407d56b68fc833078cbac5db5d9e0f32746c-2048x2048.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Open Images\\ \\ Use FiftyOne, the recommended integration by Google, to download and visualize their Open Images dataset.](https://docs.voxel51.com/integrations/open_images.html) [![](https://cdn.sanity.io/images/h6toihm1/production/2219f184ebdc669c5f5a7d913bb3e69483324a23-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ OpenCLIP\\ \\ Run inference with CLIP models on FiftyOne datasets](https://docs.voxel51.com/integrations/openclip.html) [![](https://cdn.sanity.io/images/h6toihm1/production/8ef4cbc492116c1e6b38aab527be5f6aabd49788-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Pinecone\\ \\ Pinecone provides a fully managed, developer-friendly, easily scalable vector database. Create Pinecone indexes, upload vectors, and run similarity queries programmatically and via the FiftyOne App.](https://docs.voxel51.com/integrations/pinecone.html) [![](https://cdn.sanity.io/images/h6toihm1/production/f34c7148236fcf4d654ca0af14c09acfbe31f223-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ PyTorch Hub\\ \\ Load any model from PyTorch hub and run inference on FiftyOne datasets](https://docs.voxel51.com/integrations/pytorch_hub.html) [![](https://cdn.sanity.io/images/h6toihm1/production/d39440530aa49b143ffb333a56311056ac7fe042-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Qdrant\\ \\ Qdrant vector database & vector similarity search engine provides search for the nearest high-dimensional vectors. Create Qdrant collections, upload vectors, and run similarity queries, both programmatically and via the FiftyOne App.](https://docs.voxel51.com/integrations/qdrant.html) [![](https://cdn.sanity.io/images/h6toihm1/production/c50e69a574941d3b4e6edfdafcb602db4fe3643f-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Redis\\ \\ Redis Vector Search enables the storage of vectors and the associated metadata within hashes or JSON to perform vector similarity searches. Create Redis vector search indexes, upload vectors, and run similarity queries programmatically or via the FiftyOne App.](https://docs.voxel51.com/integrations/redis.html) [![](https://cdn.sanity.io/images/h6toihm1/production/03d451b1067e3b985d44e0d2204b7d761502d628-24x24.svg)\\ \\ SDK\\ \\ FiftyOne Teams-specific methods for managing users, dataset permissions, plugins, API keys, and more.](https://docs.voxel51.com/enterprise/management_sdk.html) [![](https://cdn.sanity.io/images/h6toihm1/production/f7f765907c6fb8b1b82b0a226273def55109ef11-1099x1903.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=48&q=75&w=48)\\ \\ Segments.ai\\ \\ Segments.ai is a data labeling platform providing 2D, 3D point cloud, and multi-sensor labeling. Request Segments.ai annotations directly from FiftyOne.](https://github.com/segments-ai/segments-voxel51-plugin) [![](https://cdn.sanity.io/images/h6toihm1/production/fc641b66f14479486a2093ada022f498672f658c-342x285.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=48&q=75&w=48)\\ \\ SuperGradients\\ \\ Run inference with YOLO-NAS models on FiftyOne datasets.](https://docs.voxel51.com/integrations/super_gradients.html) [![](https://cdn.sanity.io/images/h6toihm1/production/15984d8836534da1c0d27f257adf8a22770cecf2-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ Ultralytics\\ \\ Load, fine-tune, and run inference with Ultralytics models on FiftyOne datasets.](https://docs.voxel51.com/integrations/ultralytics.html) [![](https://cdn.sanity.io/images/h6toihm1/production/8deb50d7201f6eb5aef710eebcb4abbdc42a9e36-546x546.png?auto=format&dpr=2&fit=max&q=75&w=48)\\ \\ V7\\ \\ V7 is a popular image and video annotation tool. Upload images or video data directly from FiftyOne to V7 to add or edit labels.](https://docs.voxel51.com/integrations/v7.html) [![](https://cdn.sanity.io/images/h6toihm1/production/33dd053a4aaae07fffd4d3ede5fba658ca8fa2e0-800x800.svg)\\ \\ Weights & Biases\\ \\ Track model experiments with Weights & Biases and visualize model predictions in FiftyOne to find the best-performing model.](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone/) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-9-lllmstxt|> ## Voxel51 Sitemap https://voxel51.com/curation2025-08-19T11:10:32Zmonthly0.9https://voxel51.com/industries/manufacturing2025-08-08T16:56:14Zmonthly0.9https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections2025-07-04T13:05:27Zmonthly0.8https://voxel51.com/customers/rios2025-08-22T16:52:10Zmonthly0.8https://voxel51.com/customers/metu2025-05-29T15:00:46Zmonthly0.8https://voxel51.com/customers/updata2025-05-29T14:16:20Zmonthly0.8https://voxel51.com/events/visual-ai-in-healthcare-june-27-20252025-08-04T13:15:38Zmonthly0.8https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately2025-07-21T07:52:58Zmonthly0.8https://voxel51.com/whitepapers/the-best-data-centric-computer-vision-tools-for-the-enterprise2025-08-14T01:30:08Zmonthly0.8https://voxel51.com/customers?category=av-physical-ai2025-05-31T18:26:02Zmonthly0.8https://voxel51.com/events/visual-ai-in-manufacturing-september-10-20252025-08-05T07:16:14Zmonthly0.8https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning2025-07-21T08:31:08Zmonthly0.8https://voxel51.com/customers/ibm2025-08-22T16:41:36Zmonthly0.8https://voxel51.com/customers/binit2025-05-29T14:01:32Zmonthly0.8https://voxel51.com/point-cloud2025-07-10T17:30:51Zmonthly0.9https://voxel51.com/events/from-research-to-reality-building-gui-agents-that-actually-work-august-22-20252025-07-16T19:03:54Zmonthly0.8https://voxel51.com/fiftyone2025-07-18T17:45:52Zmonthly0.9https://voxel51.com/industries/sports2025-06-06T08:02:32Zmonthly0.9https://voxel51.com/customers/fast-code-ai2025-05-29T14:55:38Zmonthly0.8https://voxel51.com/customers?category=retail-consumer2025-05-31T18:26:06Zmonthly0.8https://voxel51.com/careers2025-06-02T18:59:39Zmonthly0.9https://voxel51.com/events/boston-ai-ml-and-computer-vision-meetup-workshop-september-25-20252025-08-04T18:02:54Zmonthly0.8https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub2025-07-18T16:03:34Zmonthly0.8https://voxel51.com/about2025-06-18T01:42:24Zmonthly0.9https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-12025-06-09T14:59:15Zmonthly0.8https://voxel51.com/events/tokyo-ai-ml-and-computer-vision-meetup-july-312025-07-28T23:34:58Zmonthly0.8https://voxel51.com/industries/retail2025-06-06T08:01:46Zmonthly0.9https://voxel51.com/events/brussels-ai-ml-and-computer-vision-meetup-july-152025-06-24T22:49:48Zmonthly0.8https://voxel51.com/events/ai-ml-and-computer-vision-meetup-june-19-20252025-08-04T14:09:33Zmonthly0.8https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier2025-06-30T22:22:30Zmonthly0.8https://voxel51.com/industries/agriculture2025-06-02T06:46:31Zmonthly0.9https://voxel51.com/blog/cvpr-20252025-06-04T23:30:51Zmonthly0.8https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation2025-07-29T21:04:11Zmonthly0.8https://voxel51.com/whitepapers/why-vision-ai-models-fails2025-08-14T01:29:44Zmonthly0.8https://voxel51.com/customers/aquabyte2025-05-29T14:06:07Zmonthly0.8https://voxel51.com/industries/robotics2025-06-06T08:01:59Zmonthly0.9https://voxel51.com/blog/how-we-built-annotation-savings-estimator2025-06-18T19:07:22Zmonthly0.8https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities2025-08-12T01:05:34Zmonthly0.8https://voxel51.com/research2025-08-05T06:34:56Zmonthly0.9https://voxel51.com/industries/autonomous-vehicles-systems2025-06-06T07:50:19Zmonthly0.9https://voxel51.com/blog/comprehensive-guide-point-cloud-data2025-07-25T15:32:00Zmonthly0.8https://voxel51.com/events/scaling-computer-vision-ai-in-the-enterprise-28-august-20252025-08-18T17:53:30Zmonthly0.8https://voxel51.com/data-centric-visual-ai2025-05-29T10:25:17Zmonthly0.9https://voxel51.com/blog/nvidia-ai-podcast-adas-with-porsche2025-07-30T16:57:42Zmonthly0.8https://voxel51.com/customers/secury3602025-08-12T00:58:47Zmonthly0.8https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies2025-07-17T03:35:36Zmonthly0.8https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone2025-08-11T15:00:10Zmonthly0.8https://voxel51.com/events/ai-ml-and-computer-vision-meetup-july-17-20252025-08-04T12:55:04Zmonthly0.8https://voxel51.com/blog/tag/visual-agents2025-06-30T22:21:31Zmonthly0.8https://voxel51.com/blog/tag/auto-labeling2025-06-03T09:21:49Zmonthly0.8https://voxel51.com/events/from-research-to-reality-building-gui-agents-that-actually-work-august-15-20252025-07-16T19:04:14Zmonthly0.8https://voxel51.com/events/getting-started-with-fiftyone-for-healthcare-use-cases-july-23-20252025-08-04T12:53:48Zmonthly0.8https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation2025-07-21T08:55:03Zmonthly0.8https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape2025-07-22T14:38:58Zmonthly0.8https://voxel51.com/events/building-visual-ai-in-the-enterprise-workshop-june-4-20252025-08-06T23:36:35Zmonthly0.8https://voxel51.com/events/getting-started-with-fiftyone-workshop-june-18-20252025-08-04T14:08:17Zmonthly0.8https://voxel51.com/blog/a-guide-to-ai-image-segmentation2025-07-21T08:26:37Zmonthly0.8https://voxel51.com/blog/category/press2025-05-20T15:41:34Zmonthly0.8https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services2025-05-22T01:53:59Zmonthly0.8https://voxel51.com/customers/aidence2025-05-29T14:39:53Zmonthly0.8https://voxel51.com/blog/smarter-automotive-datasets-selection2025-07-24T16:51:40Zmonthly0.8https://voxel51.com/customers/kitro2025-05-29T14:10:02Zmonthly0.8https://voxel51.com/pricing2025-08-18T16:59:51Zmonthly0.9https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail2025-07-21T07:32:05Zmonthly0.8https://voxel51.com/customers/indian-institute-of-it-and-management2025-05-29T14:14:16Zmonthly0.8https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-32025-06-09T15:07:31Zmonthly0.8https://voxel51.com/events/raleigh-ai-ml-and-computer-vision-meetup-august-20-20252025-07-23T14:36:23Zmonthly0.8https://voxel51.com/events/best-of-cvpr-july-9-20252025-08-04T13:12:38Zmonthly0.8https://voxel51.com/customers/taranis2025-05-29T14:52:30Zmonthly0.8https://voxel51.com/blog/how-automated-data-labeling-enhances-computer-vision-efficiency-and-accuracy2025-07-21T09:10:23Zmonthly0.8https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs2025-06-05T16:44:08Zmonthly0.8https://voxel51.com/customers/forsight2025-08-22T16:44:18Zmonthly0.8https://voxel51.com/customers?category=dev-tools2025-05-27T22:05:37Zmonthly0.8https://voxel51.com/events/amsterdam-ai-ml-and-computer-vision-meetup-july-142025-06-25T13:44:05Zmonthly0.8https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai2025-06-12T16:07:55Zmonthly0.8https://voxel51.com/events/visual-ai-in-manufacturing-how-multimodal-data-powers-adaptive-process-control2025-08-04T23:47:14Zmonthly0.8https://voxel51.com/customers/rif-robotics2025-05-29T13:55:24Zmonthly0.8https://voxel51.com/events/ai-ml-and-computer-vision-meetup-en-espanol-october-23-20252025-08-18T17:09:13Zmonthly0.8https://voxel51.com/whitepapers/visual-ai-for-defect-detection-in-manufacturing2025-08-14T01:29:12Zmonthly0.8https://voxel51.com/customers/adt-commercial2025-08-22T16:39:05Zmonthly0.8https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness2025-07-21T08:51:21Zmonthly0.8https://voxel51.com/blog/what-makes-good-data-a-view-from-the-front-lines-of-ai2025-07-17T16:19:10Zmonthly0.8https://voxel51.com/events/from-research-to-reality-building-gui-agents-that-actually-work-august-29-20252025-07-16T19:03:37Zmonthly0.8https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation2025-06-27T19:15:55Zmonthly0.8https://voxel51.com/blog/the-complete-guide-to-auto-labeling2025-07-21T07:54:15Zmonthly0.8https://voxel51.com/customers/ai-fish2025-05-29T14:03:48Zmonthly0.8https://voxel51.com/voxelgpt2025-05-29T06:16:23Zmonthly0.9https://voxel51.com/events/advanced-car-damage-detection-with-fiftyone-and-the-cardd-dataset-july-122025-06-29T20:12:14Zmonthly0.8https://voxel51.com/events/valencia-ai-ml-and-computer-vision-meetup-september-25-20252025-08-05T16:40:35Zmonthly0.8https://voxel51.com/events/verified-auto-labeling-smarter-annotation-at-scale-june-24-20252025-08-04T13:17:50Zmonthly0.8https://voxel51.com/blog/vggt-is-a-pure-neural-approach-to-3d-vision2025-06-26T06:12:03Zmonthly0.8https://voxel51.com/customers/allstate2025-08-22T16:42:28Zmonthly0.8https://voxel51.com/blog/tag/video2025-06-03T09:56:20Zmonthly0.8https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision2025-07-21T09:07:06Zmonthly0.8https://voxel51.com/events/visual-ai-in-manufacturing-and-robotics-september-11-20252025-08-11T17:35:39Zmonthly0.8https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks2025-08-20T14:22:52Zmonthly0.8https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using2025-06-23T07:57:21Zmonthly0.8https://voxel51.com/events/boston-ai-ml-and-computer-vision-meetup-june-26-20252025-06-06T12:47:49Zmonthly0.8https://voxel51.com/events/virtual-how-porsche-uses-auto-labeling-to-supercharge-av-development-4-september-20252025-08-18T17:53:01Zmonthly0.8https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai2025-07-03T20:59:41Zmonthly0.8https://voxel51.com/events/ai-ml-and-computer-vision-meetup-aug-28-20252025-08-14T08:16:54Zmonthly0.8https://voxel51.com/events/visual-ai-in-manufacturing-and-robotics-september-12-20252025-08-21T19:15:12Zmonthly0.8https://voxel51.com/events/preventing-critical-misses-in-defect-detection-a-data-centric-approach-with-mongodb-and-voxel512025-08-18T17:53:15Zmonthly0.8https://voxel51.com/blog/tag/cost-estimation2025-06-03T09:42:20Zmonthly0.8https://voxel51.com/industries/security2025-06-06T08:02:12Zmonthly0.9https://voxel51.com/blog/deploy-computer-vision-in-manufacturing2025-08-06T19:40:53Zmonthly0.8https://voxel51.com/customers?category=security2025-05-27T21:57:12Zmonthly0.8https://voxel51.com/industries/defense2025-06-06T07:54:50Zmonthly0.9https://voxel51.com/events/getting-started-with-fiftyone-for-manufacturing-use-cases-sept-30-20252025-08-04T18:00:00Zmonthly0.8https://voxel51.com/blog/uncommon-objects-in-3d2025-06-18T21:17:05Zmonthly0.8https://voxel51.com/events/madrid-ai-ml-and-computer-vision-meetup-september-26-20252025-08-05T16:08:13Zmonthly0.8https://voxel51.com/blog/opening-remarks-from-cvpr-20252025-06-20T21:30:21Zmonthly0.8https://voxel51.com/customers/raytheon-technologies2025-08-22T16:51:21Zmonthly0.8https://voxel51.com/events/exposing-your-datas-blind-spots-scenario-mining-for-safer-av2025-08-05T14:23:15Zmonthly0.8https://voxel51.com/customers/vivint2025-05-29T14:42:59Zmonthly0.8https://voxel51.com/events/workshop-train-a-medical-ai-model-in-one-day-july-25-20252025-07-14T16:26:06Zmonthly0.8https://voxel51.com/events/women-in-ai-july-242025-08-04T12:51:07Zmonthly0.8https://voxel51.com/sales2025-05-22T10:22:48Zmonthly0.9https://voxel51.com/blog/tag/agi2025-06-17T15:49:05Zmonthly0.8https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition2025-07-21T09:21:02Zmonthly0.8https://voxel51.com/blog2025-07-30T16:41:44Zmonthly0.8https://voxel51.com/blog/category/learn2025-07-21T07:23:09Zmonthly0.8https://voxel51.com/community2025-07-17T16:47:53Zmonthly0.9https://voxel51.com/events/women-in-ai-october-2-20252025-07-31T23:06:57Zmonthly0.8https://voxel51.com/industries/healthcare2025-06-24T20:37:20Zmonthly0.9https://voxel51.com/customers/aisprid2025-05-30T17:45:02Zmonthly0.8https://voxel51.com/customers/g422025-05-29T14:21:20Zmonthly0.8https://voxel51.com/events/food-waste-estimation-hackathon-computer-vision-for-sustainability-1-august-20252025-07-31T15:27:43Zmonthly0.8https://voxel51.com/careers2025-06-02T06:57:52Zmonthly0.8https://voxel51.com/customers2025-06-26T22:19:00Zmonthly0.8https://voxel51.com/events/understanding-visual-agents-august-7-20252025-08-12T22:40:02Zmonthly0.8https://voxel51.com/customers/protex-ai2025-07-15T16:51:21Zmonthly0.8https://voxel51.com/customers/qinecsa2025-05-29T14:18:36Zmonthly0.8https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data2025-07-21T09:05:13Zmonthly0.8https://voxel51.com/link-catcher2025-05-12T20:01:57Zmonthly0.9https://voxel51.com/evaluation2025-08-13T19:41:15Zmonthly0.9https://voxel51.com/customers/smart-eye2025-08-22T16:38:01Zmonthly0.8https://voxel51.com/blog/tag/verified-auto-labeling2025-06-03T09:41:50Zmonthly0.8https://voxel51.com/customers/safelyyou2025-08-22T16:54:17Zmonthly0.8https://voxel51.com/events/best-of-cvpr-july-11-20252025-08-04T12:57:22Zmonthly0.8https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-22025-06-09T15:06:31Zmonthly0.8https://voxel51.com/customers/wildlife-ai2025-05-29T13:57:59Zmonthly0.8https://voxel51.com/customers/berkshire-grey2025-08-22T16:50:24Zmonthly0.8https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models2025-07-21T09:14:50Zmonthly0.8https://voxel51.com/customers/lancedb2025-05-29T14:57:24Zmonthly0.8https://voxel51.com/blog/tag/annotation-savings2025-06-03T09:42:06Zmonthly0.8https://voxel51.com/blog/visual-agents-at-cvpr-20252025-06-02T20:19:41Zmonthly0.8https://voxel51.com/customers?category=aerospace-and-defense2025-05-27T21:43:27Zmonthly0.8https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape2025-06-26T04:29:55Zmonthly0.8https://voxel51.com/customers?category=health-and-medicine2025-05-27T06:31:54Zmonthly0.8https://voxel51.com/customers/seafar2025-05-29T14:30:29Zmonthly0.8https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic2025-06-17T15:55:36Zmonthly0.8https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance2025-08-19T17:19:59Zmonthly0.8https://voxel51.com/whitepapers/auto-labeling-data-for-object-detection2025-08-14T01:30:27Zmonthly0.8https://voxel51.com/industries/aviation2025-06-06T07:50:30Zmonthly0.9https://voxel51.com/events2025-05-20T12:29:44Zmonthly0.8https://voxel51.com/customers?category=agriculture-sustainability2025-05-27T22:00:36Zmonthly0.8https://voxel51.com/whitepapers/your-data-your-advantage2025-08-14T01:28:26Zmonthly0.8https://voxel51.com/2025-08-19T11:49:47Zmonthly1https://voxel51.com/blog/composed-image-retrieval-at-cvpr-20252025-06-05T14:26:00Zmonthly0.8https://voxel51.com/customers/ancera2025-08-22T16:47:20Zmonthly0.8https://voxel51.com/get-started2025-06-06T17:47:36Zmonthly0.9https://voxel51.com/blog/author/antonio-rueda-toicen2025-08-12T00:44:52Zmonthly0.8https://voxel51.com/customers/fyma2025-05-29T14:34:14Zmonthly0.8https://voxel51.com/customers/finegrain2025-05-29T14:24:18Zmonthly0.8https://voxel51.com/events/paris-ai-ml-and-computer-vision-meetup-july-162025-07-10T19:53:02Zmonthly0.8https://voxel51.com/annotation2025-08-19T17:22:38Zmonthly0.9https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes2025-07-21T08:38:31Zmonthly0.8https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows2025-07-21T09:03:27Zmonthly0.8https://voxel51.com/events/ai-ml-and-computer-vision-meetup-en-espanol-august-21-20252025-07-31T23:11:46Zmonthly0.8https://voxel51.com/customers/argosai2025-06-03T20:26:32Zmonthly0.8https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai2025-08-01T21:31:38Zmonthly0.8https://voxel51.com/glossary2025-05-28T07:18:10Zmonthly0.8https://voxel51.com/integrations2025-08-12T00:05:52Zmonthly0.8https://voxel51.com/plugins2025-08-12T00:05:17Zmonthly0.8https://voxel51.com/press2025-07-24T16:47:07Zmonthly0.8https://voxel51.com/webinars2025-06-02T20:52:14Zmonthly0.8https://voxel51.com/whitepapers2025-07-29T08:55:19Zmonthly0.8https://voxel51.com/blog/fiftyone-open-images-collaboration2025-05-20T15:42:44Zmonthly0.8https://voxel51.com/blog/fiftyone-open-source-launch2025-05-20T15:42:44Zmonthly0.8https://voxel51.com/blog/voxel51-physical-distancing-index2025-05-20T15:42:44Zmonthly0.8https://voxel51.com/blog/voxel51-launches-image-to-video-tool2025-05-20T15:42:44Zmonthly0.8https://voxel51.com/blog/voxel51-raises-2-million-to-advance-video-understanding2025-05-20T15:42:44Zmonthly0.8https://voxel51.com/blog/introducing-voxelgpt2025-05-20T15:42:44Zmonthly0.8https://voxel51.com/blog/voxel51-raises-30m-series-b-funding-to-make-visual-ai-a-reality2025-05-20T15:42:44Zmonthly0.8https://voxel51.com/blog/voxel51-launches-fiftyone-open-source-1-0-accelerating-the-creation-of-production-ready-visual-ai-applications2025-05-20T15:42:44Zmonthly0.8https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset2025-05-22T04:24:10Zmonthly0.8https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision2025-05-22T04:24:10Zmonthly0.8https://voxel51.com/blog/forsight-finds-a-centralized-dataset-management-solution-in-fiftyone-teams2025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes2025-05-22T04:24:21Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-20222025-05-22T04:24:10Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-20222025-05-22T04:23:56Zmonthly0.8https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see2025-05-22T04:23:56Zmonthly0.8https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far2025-05-22T04:23:56Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-20222025-05-22T04:24:10Zmonthly0.8https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-20222025-05-22T04:24:10Zmonthly0.8https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools2025-05-22T04:24:10Zmonthly0.8https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-20222025-05-22T04:24:10Zmonthly0.8https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision2025-05-22T04:24:10Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-20222025-05-22T04:24:10Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-20222025-05-22T04:24:10Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-20222025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/computer-vision-meetup-update-november-222025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-20222025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards2025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-20222025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-20222025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/its-our-birthday-voxel51-turns-four2025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-20222025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/hack-on-fiftyone-in-hacktoberfest-20222025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams2025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-20222025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/announcing-our-12-5m-series-a-funding-to-bring-transparency-and-clarity-to-the-worlds-data2025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins2025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-20222025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-20222025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-20222025-05-22T04:03:08Zmonthly0.8https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data2025-05-22T04:23:56Zmonthly0.8https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel512025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-20222025-05-22T04:23:56Zmonthly0.8https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-20222025-05-22T04:23:56Zmonthly0.8https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel512025-05-22T04:23:56Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-20232025-05-22T04:23:56Zmonthly0.8https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset2025-05-22T04:23:40Zmonthly0.8https://voxel51.com/blog/finding-images-with-words2025-05-22T04:23:40Zmonthly0.8https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone2025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/meetup-recap-how-to-build-high-quality-machine-learning-datasets-and-computer-vision-models2025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models2025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/how-to-train-your-dragon-detector2025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-20232025-05-22T04:23:40Zmonthly0.8https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models2025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-20232025-05-22T04:23:40Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-20232025-05-22T04:23:40Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-20232025-05-22T04:23:34Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-20232025-05-22T04:23:34Zmonthly0.8https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-20232025-05-22T04:23:29Zmonthly0.8https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-20232025-05-22T04:23:21Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-20232025-05-22T04:23:04Zmonthly0.8https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox2025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/the-making-of-avatar-the-way-of-water2025-05-22T04:23:40Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-20232025-05-22T04:23:40Zmonthly0.8https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation2025-05-22T04:03:25Zmonthly0.8https://voxel51.com/blog/fiftyone-turns-one2025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation2025-05-22T04:03:21Zmonthly0.8https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone2025-05-22T04:03:25Zmonthly0.8https://voxel51.com/blog/fiftyone-six-months-post-launch2025-05-22T04:03:25Zmonthly0.8https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision2025-05-22T04:03:25Zmonthly0.8https://voxel51.com/blog/people-voxel51-spotlight-on-jimmy-guerrero2025-05-22T04:23:40Zmonthly0.8https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-20232025-08-12T01:03:17Zmonthly0.8https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free2025-05-22T04:23:34Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-20232025-05-22T04:23:34Zmonthly0.8https://voxel51.com/blog/giving-yolov8-a-second-look-part-12025-05-22T04:23:29Zmonthly0.8https://voxel51.com/blog/giving-yolov8-a-second-look-part-22025-05-22T04:23:29Zmonthly0.8https://voxel51.com/blog/giving-yolov8-a-second-look-part-32025-05-22T04:23:34Zmonthly0.8https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap2025-05-22T04:23:34Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-0-192025-05-22T04:23:34Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-20232025-05-22T04:23:34Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-20232025-05-22T04:23:29Zmonthly0.8https://voxel51.com/blog/people-voxel51-spotlight-on-lanny-wang2025-05-22T04:23:29Zmonthly0.8https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset2025-05-22T04:23:29Zmonthly0.8https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision2025-05-22T04:23:29Zmonthly0.8https://voxel51.com/blog/exploring-google-open-images-v72025-05-22T04:23:29Zmonthly0.8https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing2025-05-22T04:23:29Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-20232025-05-22T04:23:21Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-20232025-05-22T04:23:21Zmonthly0.8https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding2025-05-22T04:23:21Zmonthly0.8https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series2025-05-22T04:23:21Zmonthly0.8https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e2025-05-22T04:23:21Zmonthly0.8https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data2025-05-22T04:23:21Zmonthly0.8https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap2025-05-22T04:23:04Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-20232025-05-22T04:23:04Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-20232025-05-22T04:23:04Zmonthly0.8https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen2025-05-22T04:23:04Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-20232025-05-22T04:23:04Zmonthly0.8https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset2025-05-22T04:23:04Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-20232025-05-22T04:22:56Zmonthly0.8https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision2025-05-22T04:22:56Zmonthly0.8https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap2025-05-22T04:22:56Zmonthly0.8https://voxel51.com/blog/generate-movement-from-text-descriptions-with-t2m-gpt2025-05-22T04:22:56Zmonthly0.8https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone2025-05-22T04:22:56Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-20232025-05-22T04:22:56Zmonthly0.8https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip2025-05-22T04:22:56Zmonthly0.8https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone2025-05-22T04:22:56Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-20232025-05-22T04:22:56Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-20232025-05-22T04:22:48Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-20232025-05-22T04:22:48Zmonthly0.8https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision2025-05-22T04:22:48Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-20232025-05-22T04:22:48Zmonthly0.8https://voxel51.com/blog/cvpr-2023-survival-guide2025-05-22T04:22:48Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-20232025-05-22T04:22:48Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-20232025-05-22T04:22:48Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-0-212025-05-22T04:22:48Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-20232025-05-22T04:22:48Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-20232025-05-22T04:22:45Zmonthly0.8https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision2025-05-22T04:22:45Zmonthly0.8https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr2025-05-22T04:22:45Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-20232025-05-22T04:22:45Zmonthly0.8https://voxel51.com/blog/too-many-pixels-so-little-time2025-05-22T04:22:45Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-20232025-05-22T04:22:45Zmonthly0.8https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins2025-05-22T04:22:45Zmonthly0.8https://voxel51.com/blog/visualize-cvpr-2023-datasets-at-cvpr-20232025-05-22T04:22:45Zmonthly0.8https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr2025-05-22T04:22:45Zmonthly0.8https://voxel51.com/blog/conquering-controlnet2025-05-22T04:22:45Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-20232025-05-22T04:22:45Zmonthly0.8https://voxel51.com/blog/the-computer-vision-interface-for-vector-search2025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-20232025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-20232025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-20232025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-20232025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-20232025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/teaching-androids-to-dream-of-sheep2025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-20232025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-20232025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-20232025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/spending-my-first-week-with-fiftyone2025-05-22T04:22:34Zmonthly0.8https://voxel51.com/blog/opencv-ai-competition-20232025-05-22T04:22:34Zmonthly0.8https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-20232025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/celebrating-three-years-of-fiftyone2025-05-22T04:22:39Zmonthly0.8https://voxel51.com/blog/build-your-own-ai-art-gallery2025-05-22T04:22:34Zmonthly0.8https://voxel51.com/blog/how-computer-vision-is-changing-healthcare2025-05-22T04:22:34Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-20232025-05-22T04:22:34Zmonthly0.8https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-20232025-05-22T04:22:34Zmonthly0.8https://voxel51.com/blog/ask-your-images-anything2025-05-22T04:22:34Zmonthly0.8https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-20232025-05-22T04:22:34Zmonthly0.8https://voxel51.com/blog/computer-vision-mastering-drone-data-training2025-05-22T04:22:14Zmonthly0.8https://voxel51.com/blog/build-custom-computer-vision-applications2025-05-22T04:22:34Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-20232025-05-22T04:22:34Zmonthly0.8https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-20232025-05-22T04:22:28Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-20232025-05-22T04:22:28Zmonthly0.8https://voxel51.com/blog/facet-benchmark2025-05-22T04:22:28Zmonthly0.8https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone2025-05-22T04:22:28Zmonthly0.8https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-20232025-05-22T04:22:28Zmonthly0.8https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract2025-05-22T04:22:14Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration2025-05-22T04:22:28Zmonthly0.8https://voxel51.com/blog/nuscenes-dataset-navigating-the-road-ahead2025-05-22T04:22:28Zmonthly0.8https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-20232025-05-22T04:22:14Zmonthly0.8https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide2025-05-22T04:22:14Zmonthly0.8https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone2025-05-22T04:22:14Zmonthly0.8https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-20232025-05-22T04:22:14Zmonthly0.8https://voxel51.com/blog/3-reasons-to-visit-voxel51-at-iccv232025-05-22T04:22:14Zmonthly0.8https://voxel51.com/blog/computer-vision-badger-custom-github-badges2025-05-22T04:22:14Zmonthly0.8https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning2025-05-22T04:22:14Zmonthly0.8https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-22025-05-22T04:22:11Zmonthly0.8https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone2025-05-22T04:22:11Zmonthly0.8https://voxel51.com/blog/computer-vision-video-labels-fiftyone-tips-and-tricks-october-14th-20232025-05-22T04:22:11Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-oct-12-20232025-05-22T04:22:11Zmonthly0.8https://voxel51.com/blog/computer-vision-concept-traversal-plugin-for-fiftyone2025-05-22T04:22:11Zmonthly0.8https://voxel51.com/blog/computer-vision-plugin-for-building-and-managing-plugins2025-05-22T04:21:49Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-20-20232025-05-22T04:22:11Zmonthly0.8https://voxel51.com/blog/computer-vision-sam-for-prediction-kaggle-football-player-segmentation-dataset2025-05-22T04:21:49Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-27-20232025-05-22T04:21:49Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-3-20232025-05-22T04:21:49Zmonthly0.8https://voxel51.com/blog/computer-vision-announcing-updates-to-fiftyone-0-22-2-and-fiftyone-teams-1-4-32025-05-22T04:22:11Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-10-20232025-05-22T04:21:29Zmonthly0.8https://voxel51.com/blog/computer-vision-10-weeks-of-building-fiftyone-plugins2025-05-22T04:21:49Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-30-20232025-05-22T04:21:49Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-nov-2-20232025-05-22T04:21:49Zmonthly0.8https://voxel51.com/blog/computer-vision-fiftyone-0-22-3-and-fiftyone-teams-1-4-42025-05-22T04:21:49Zmonthly0.8https://voxel51.com/blog/tracking-datasets-in-fiftyone2025-05-22T04:21:29Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-24-20232025-05-22T04:21:29Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-17-20232025-05-22T04:21:29Zmonthly0.8https://voxel51.com/blog/fiftyone-plugins-tips-and-tricks-november-22-20232025-05-22T04:21:29Zmonthly0.8https://voxel51.com/blog/computer-vision-generating-videos-from-images-with-stable-video-diffusion-and-fiftyone2025-05-22T04:21:29Zmonthly0.8https://voxel51.com/blog/computer-vision-elevate-your-github-readme-game2025-05-22T04:21:29Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-0-23-and-fiftyone-teams-1-52025-05-22T04:21:29Zmonthly0.8https://voxel51.com/blog/neurips-2023-and-the-state-of-ai-research2025-05-22T04:21:29Zmonthly0.8https://voxel51.com/blog/understanding-llava-large-language-and-vision-assistant2025-05-22T04:21:25Zmonthly0.8https://voxel51.com/blog/announcing-the-voxel51-v7-partnership2025-05-22T04:21:25Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-dec-7-20232025-05-22T04:21:29Zmonthly0.8https://voxel51.com/blog/neurips-2023-survival-guide2025-05-22T04:21:29Zmonthly0.8https://voxel51.com/blog/computer-vision-and-fiftyone-community-year-in-review-20232025-05-22T04:21:25Zmonthly0.8https://voxel51.com/blog/why-2023-was-the-most-exciting-year-in-computer-vision-history-so-far2025-05-22T04:21:25Zmonthly0.8https://voxel51.com/blog/how-to-build-a-semantic-search-engine-for-emojis2025-05-22T04:21:25Zmonthly0.8https://voxel51.com/blog/how-computer-vision-is-changing-sports2025-05-22T04:21:25Zmonthly0.8https://voxel51.com/blog/comparing-vqa-and-action-recognition2025-05-22T04:21:25Zmonthly0.8https://voxel51.com/blog/how-to-estimate-depth-from-a-single-image2025-05-22T04:20:47Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-jan-25-20242025-05-22T04:21:19Zmonthly0.8https://voxel51.com/blog/how-computer-vision-is-changing-retail2025-08-12T01:11:43Zmonthly0.8https://voxel51.com/blog/how-to-visualize-your-data-with-dimension-reduction-techniques2025-05-22T04:21:19Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-16-20242025-05-22T04:21:19Zmonthly0.8https://voxel51.com/blog/exploring-gradcam-and-more-with-fiftyone2025-05-22T04:21:19Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-23-20242025-05-22T04:21:19Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-feb-15-20242025-05-22T04:21:19Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-march-1-20242025-05-22T04:21:19Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-march-8-20242025-05-22T04:21:11Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-0-23-5-and-fiftyone-teams-1-5-62025-05-22T04:21:19Zmonthly0.8https://voxel51.com/blog/finding-outliers-in-your-vision-datasets2025-05-22T04:21:11Zmonthly0.8https://voxel51.com/blog/data-augmentation-is-still-data-curation2025-08-12T01:15:51Zmonthly0.8https://voxel51.com/blog/how-computer-vision-is-changing-security2025-05-22T04:21:11Zmonthly0.8https://voxel51.com/blog/a-history-of-clip-model-training-data-advances2025-05-22T04:21:11Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-0-23-6-and-fiftyone-teams-1-5-72025-05-22T04:21:11Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-march-15-20242025-05-22T04:21:11Zmonthly0.8https://voxel51.com/blog/streamline-computer-vision-workflows-with-hugging-face-transformers-and-fiftyone2025-05-22T04:21:11Zmonthly0.8https://voxel51.com/blog/how-to-easily-cluster-your-computer-vision-datasets2025-05-22T04:21:11Zmonthly0.8https://voxel51.com/blog/efficiently-managing-and-querying-visual-data-with-mongodb-atlas-vector-search-and-fiftyone2025-05-22T04:21:11Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-march-22-20242025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/fiftyone-to-integrate-with-nvidia-omniverse-simulation-services2025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-march-21-20242025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-march-29-20242025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/fiftyone-wins-2024-artificial-intelligence-excellence-award2025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/finding-the-optimal-confidence-threshold2025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-5-20242025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-0-23-7-and-fiftyone-teams-1-5-82025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams2025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/iso-27001-certified2025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-march-29-20242025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/how-to-cluster-images2025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-april-12-20242025-05-22T04:20:59Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-12-20242025-05-22T04:20:47Zmonthly0.8https://voxel51.com/blog/how-computer-vision-is-transforming-robotics2025-08-12T01:08:48Zmonthly0.8https://voxel51.com/blog/cvpr-2024-survival-guide-five-vision-language-papers-you-dont-want-to-miss2025-05-22T04:20:47Zmonthly0.8https://voxel51.com/blog/how-to-detect-small-objects2025-05-22T04:20:47Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-april-18-20242025-05-22T04:20:47Zmonthly0.8https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets2025-05-22T04:20:47Zmonthly0.8https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks2025-05-22T04:20:40Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-teams-1-62025-05-22T04:20:40Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-april-26-20242025-05-22T04:20:47Zmonthly0.8https://voxel51.com/blog/anomaly-detection-with-fiftyone-and-anomalib2025-05-22T04:20:40Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-may-2-20242025-05-22T04:20:40Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-may-10-20242025-05-22T04:20:40Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-may-8-20242025-05-22T04:20:40Zmonthly0.8https://voxel51.com/blog/secury360-strengthens-security-with-fiftyone-teams2025-05-22T04:20:40Zmonthly0.8https://voxel51.com/blog/announcing-series-b-led-by-bessemer-venture-partners2025-05-30T00:34:40Zmonthly0.8https://voxel51.com/blog/voxel51-and-bessemer-venture-partners-collaborate-to-make-visual-ai-a-reality2025-05-22T04:20:40Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-may-24-20242025-05-22T04:20:40Zmonthly0.8https://voxel51.com/blog/how-shap-e-changed-how-we-think-about-diffusion-models2025-05-22T04:20:35Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-0-24-with-3d-meshes-and-custom-workspaces2025-05-22T04:20:35Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-may-30-20242025-05-22T04:20:35Zmonthly0.8https://voxel51.com/blog/voxel51-at-cvpr-20242025-05-22T04:20:35Zmonthly0.8https://voxel51.com/blog/5-papers-on-my-cvpr-2024-must-see-list2025-05-22T04:20:35Zmonthly0.8https://voxel51.com/blog/how-i-built-an-in-cabin-perception-dataset2025-05-22T04:20:35Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-june-21-20242025-05-22T04:20:35Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-june-27-20242025-05-22T04:20:35Zmonthly0.8https://voxel51.com/blog/segment-anything-in-a-ct-scan-with-nvidia-vista-3d2025-05-22T04:20:35Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-19-20242025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-july-3-20242025-05-22T04:20:35Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-20242025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-26-20242025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-august-02-20242025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/announcing-the-data-centric-ai-competition-revolutionizing-object-detection-through-smart-data-curation2025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-august-9-20242025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/what-is-visual-ai-going-beyond-computer-vision2025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-august-8-20242025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/four-years-of-open-source-fiftyone2025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-august-15-20242025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-august-16-20242025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/sam-2-is-now-available-in-fiftyone2025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-0-252025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-august-23-20242025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-august-30-20242025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-august-29-20242025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-september-12-20242025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-september-13-20242025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/recapping-the-visual-ai-in-healthcare-meetup-september-19-20242025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-september-20-20242025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/segments-ai-plugin-for-fiftyone2025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-september-26-20242025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/the-power-of-open-source-ai-how-fiftyone-drives-the-future-of-visual-ai2025-05-22T04:20:26Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-1-02025-05-22T04:20:22Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-october-4-20242025-05-22T04:20:22Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-october-10-20242025-05-22T04:20:22Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-october-11-20242025-05-22T04:20:22Zmonthly0.8https://voxel51.com/blog/cotracker3-a-point-tracker-using-real-videos2025-05-22T04:20:22Zmonthly0.8https://voxel51.com/blog/cotracker3-enhanced-point-tracking-with-less-data2025-05-22T04:20:22Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-october-24-20242025-05-22T04:20:22Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-november-1-20242025-05-22T04:20:22Zmonthly0.8https://voxel51.com/blog/data-quality-the-hidden-driver-of-ai-success2025-05-22T04:20:22Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-november-14-20242025-08-04T16:43:20Zmonthly0.8https://voxel51.com/blog/recapping-eccv-2024-redux-day-12025-05-22T04:20:22Zmonthly0.8https://voxel51.com/blog/recapping-eccv-2024-redux-day-32025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/recapping-eccv-2024-redux-day-42025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/the-neurlps-2024-preshow-naturalbench-evaluating-vision-language-models-on-natural-adversarial-samples2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/the-neurlps-2024-preshow-a-textbook-remedy-for-domain-shifts-knowledge-priors-for-medical-image-analysis2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/the-neurlps-2024-preshow-a-label-is-worth-a-thousand-images-in-dataset-distillation2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/the-neurlps-2024-preshow-what-matters-when-building-vision-language-models2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/the-neurips-2024-preshow-zero-shot-learning-a-misnomer2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/the-neurips-2024-preshow-creating-spiqa-addressing-the-limitations-of-existing-datasets-for-scientific-vqa2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/journey-into-visual-ai-exploring-fiftyone-together-part-i-introduction2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-teams-2-22025-05-22T04:20:15Zmonthly0.8https://voxel51.com/blog/the-neurips-2024-preshow-are-we-measuring-what-we-think-we-are-the-perils-of-contaminated-benchmark-datasets2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/the-neurips-2024-preshow-a-data-centric-look-at-curation-strategies-for-image-classification2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/the-neurips-2024-preshow-using-knowledge-graphs-to-diagnose-and-debias-visual-datasets2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/the-neurips-2024-preshow-data-quality-over-quantity-why-real-images-still-reign-supreme-for-vision-model-training2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/the-neurips-2024-preshow-more-than-meets-the-eye-how-transformations-reveal-the-hidden-biases-shaping-our-datasets2025-05-22T04:20:18Zmonthly0.8https://voxel51.com/blog/five-must-read-data-centric-ai-papers-from-neurips-20242025-05-22T04:20:15Zmonthly0.8https://voxel51.com/blog/on-leaky-datasets-and-a-clever-horse2025-05-22T04:20:15Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-december-12-20242025-05-22T04:20:15Zmonthly0.8https://voxel51.com/blog/what-ai-means-for-science-in-20252025-05-22T04:20:15Zmonthly0.8https://voxel51.com/blog/why-2024-was-the-best-year-for-visual-ai-so-far2025-05-22T04:20:15Zmonthly0.8https://voxel51.com/blog/journey-into-visual-ai-exploring-fiftyone-together-part-ii-getting-started2025-05-22T04:20:15Zmonthly0.8https://voxel51.com/blog/journey-into-visual-ai-exploring-fiftyone-together-part-iii-preparing-a-computer-vision-challenge2025-05-22T04:20:15Zmonthly0.8https://voxel51.com/blog/bias-in-data-what-embeddings-reveal-about-real-vs-synthetic-data-distribution2025-05-22T04:20:15Zmonthly0.8https://voxel51.com/blog/how-to-make-the-best-self-driving-dataset2025-05-22T04:20:15Zmonthly0.8https://voxel51.com/blog/the-frontier-of-visual-ai-in-medical-imaging2025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-24-20232025-05-22T04:23:04Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-0-202025-05-22T04:23:21Zmonthly0.8https://voxel51.com/blog/announcing-fiftyone-teams-1-22025-05-22T04:23:04Zmonthly0.8https://voxel51.com/blog/recapping-the-computer-vision-meetup-sept-14-20232025-05-22T04:22:28Zmonthly0.8https://voxel51.com/blog/caffe-computer-vision-glacial-mass-modeling2025-05-22T04:22:28Zmonthly0.8https://voxel51.com/blog/heatmaps-fiftyone-tips-and-tricks-october-6th-20232025-05-22T04:22:11Zmonthly0.8https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-oct-5-20232025-05-22T04:22:11Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-community-update-october-20232025-05-22T04:22:11Zmonthly0.8https://voxel51.com/blog/happy-5th-birthday-voxel512025-05-22T04:22:11Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-community-update-november-20232025-05-22T04:21:49Zmonthly0.8https://voxel51.com/blog/fiftyone-0-23-3-and-fiftyone-teams-1-5-42025-05-22T04:21:25Zmonthly0.8https://voxel51.com/blog/fiftyone-computer-vision-community-update-february-20242025-05-22T04:21:19Zmonthly0.8https://voxel51.com/blog/solving-the-ai-blindspot-using-data-to-drive-models-in-automotive2025-05-22T04:20:31Zmonthly0.8https://voxel51.com/blog/voxel51-filtered-views-newsletter-january-17-20252025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/how-to-tame-your-data-dragon2025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/journey-into-visual-ai-exploring-fiftyone-together-part-iv-model-evaluation2025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/elderly-action-recognition-no-one-should-age-alone-ais-promise-for-the-next-generation-of-elders2025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/understanding-dataset-difficulty-with-class-wise-autoencoders2025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/voxel51-and-ori-partner-to-accelerate-visual-ai-innovation-in-u-s-government2025-05-20T15:42:44Zmonthly0.8https://voxel51.com/blog/visual-understanding-with-aimv22025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/beyond-the-microscope-diving-into-bioscan-5m-a-new-dataset-for-insect-biodiversity-research2025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/aimv2-outperforms-clip-on-synthetic-dataset-imagenet-d2025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/webuot-1m-a-dataset-for-underwater-object-tracking2025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/imagenet-d-new-synthetic-test-set-designed-to-rigorously-evaluate-the-robustness-of-neural-networks2025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/can-vlms-hear-what-they-see2025-05-22T04:20:04Zmonthly0.8https://voxel51.com/blog/gaussian-splatting-from-research-to-reality2025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/supercharge-your-visual-ai-workflow-fiftyone-new-plugin-for-janus-pro2025-05-22T04:20:08Zmonthly0.8https://voxel51.com/blog/memes-are-the-vlm-benchmark-we-deserve2025-05-22T04:20:04Zmonthly0.8https://voxel51.com/blog/this-visual-illusions-benchmark-makes-me-question-the-power-of-vlms2025-05-22T04:20:04Zmonthly0.8https://voxel51.com/blog/introducing-fiftyone-enterprise-visual-ai-workflows2025-05-30T06:43:19Zmonthly0.8https://voxel51.com/blog/new-data-and-model-workflows-from-voxel51-accelerate-visual-ai-development-for-enterprises2025-05-20T15:42:44Zmonthly0.8https://voxel51.com/blog/streamline-visual-data-discovery-with-fiftyone-data-lens2025-05-22T04:20:04Zmonthly0.8https://voxel51.com/blog/visualizing-model-certainty-in-the-unknown2025-05-22T04:20:04Zmonthly0.8https://voxel51.com/blog/build-better-visual-ai-datasets-with-the-fiftyone-data-quality-workflow2025-05-22T04:20:04Zmonthly0.8https://voxel51.com/blog/unified-model-insights-with-fiftyone-model-evaluation-workflows2025-05-22T04:20:04Zmonthly0.8https://voxel51.com/blog/computer-vision-for-earth-observation-from-manual-digitizing-to-ai-powered-analysis2025-05-22T04:20:04Zmonthly0.8https://voxel51.com/events/visual-ai-in-healthcare-june-26-20252025-08-03T20:00:32Zmonthly0.8https://voxel51.com/events/best-of-cvpr-july-10-20252025-08-04T13:10:56Zmonthly0.8https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-jan-22-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/visual-ai-for-geospatial-jan-29-20252025-08-05T09:24:45Zmonthly0.8https://voxel51.com/events/visual-ai-for-geospatial-jan-31-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/visual-ai-hackathon-jan-31-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/berlin-ai-ml-computer-vision-meetup-feb-7-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/elderly-action-recognition-challenge-wacv-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/best-of-neurips-feb-6-20252025-08-05T09:18:34Zmonthly0.8https://voxel51.com/events/visual-ai-hackathon-march-9-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/munich-ai-ml-computer-vision-meetup-feb-6-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/visual-ai-hackathon-feb-15-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/stuttgart-ai-ml-computer-vision-meetup-feb-5-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/boston-ai-ml-computer-vision-meetup-feb-28-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/best-of-neurips-feb-4-20252025-08-05T09:19:39Zmonthly0.8https://voxel51.com/events/dusseldorf-ai-ml-computer-vision-meetup-feb-4-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/getting-started-with-fiftyone-workshop-feb-19-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/ai-machine-learning-computer-vision-meetup-feb-20-20252025-08-05T09:17:10Zmonthly0.8https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-feb-26-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/visual-ai-hackathon-march-21-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/visual-ai-hackathon-march-22-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/visual-ai-in-agriculture-march-262025-08-04T15:29:31Zmonthly0.8https://voxel51.com/events/visual-ai-hackathon-march-15-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/getting-started-with-fiftyone-workshop-march-12-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-workshop-march-19-20252025-08-04T16:17:49Zmonthly0.8https://voxel51.com/events/visual-ai-hackathon-april-4-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/building-visual-ai-in-the-enterprise-workshop-march-27-20252025-08-06T23:36:44Zmonthly0.8https://voxel51.com/events/world-agri-tech2025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/foundations-of-computer-vision-workshop-march-4-20252025-08-04T14:27:52Zmonthly0.8https://voxel51.com/events/neural-networks-fundamentals-multilayer-perceptrons-for-regression-workshop-march-11-20252025-08-04T14:28:54Zmonthly0.8https://voxel51.com/events/nvidia-gtc-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/training-evaluation-of-classification-models-workshop-march-18-20252025-08-04T14:39:59Zmonthly0.8https://voxel51.com/events/tech-ad-europe-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/raleigh-ai-machine-learning-and-computer-vision-meetup-april-17-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/vand-3-0-cvpr-2025-workshop2025-05-26T10:41:40Zmonthly0.8https://voxel51.com/events/ai-machine-learning-computer-vision-meetup-april-24-20252025-08-04T14:52:53Zmonthly0.8https://voxel51.com/events/vand-3-0-challenge-at-cvpr-20252025-06-06T13:09:45Zmonthly0.8https://voxel51.com/events/ai-machine-learning-computer-vision-meetup-march-20-20252025-08-04T16:15:51Zmonthly0.8https://voxel51.com/events/convolutional-neural-networks-lenet5-workshop-march-25-20252025-08-04T14:41:03Zmonthly0.8https://voxel51.com/events/training-techniques-for-convolutional-networks-workshop-april-1-20252025-08-04T14:42:24Zmonthly0.8https://voxel51.com/events/multi-label-classification-with-binary-cross-entropy-amazon-satellite-images-workshop-april-8-20252025-08-04T14:43:40Zmonthly0.8https://voxel51.com/events/interpretability-in-computer-vision-cam-grad-cam-workshop-april-15-20252025-08-04T14:43:35Zmonthly0.8https://voxel51.com/events/convolutional-neural-networks-advanced-upsampling-u-net-for-semantic-segmentation-workshop-april-22-20252025-08-04T14:43:57Zmonthly0.8https://voxel51.com/events/model-optimization-data-augmentation-regularization-workshop-april-29-20252025-08-04T14:44:22Zmonthly0.8https://voxel51.com/events/image-embeddings-zero-shot-classification-with-clip-workshop-may-6-20252025-08-04T14:37:12Zmonthly0.8https://voxel51.com/events/object-detection-instance-segmentation-yolo-in-practice-workshop-may-13-20252025-08-04T14:46:18Zmonthly0.8https://voxel51.com/events/image-generation-diffusion-models-u-net-workshop-may-20-20252025-08-04T14:21:10Zmonthly0.8https://voxel51.com/events/boston-ai-ml-and-computer-vision-meetup-april-18-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/deep-learning-fundamentals-with-pytorch-and-fiftyone-workshop-april-5-6-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/new-york-ai-ml-and-computer-vision-meetup-april-3-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/computer-vision-developer-hour-march-182025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/computer-vision-developer-hour-march-252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/computer-vision-developer-hour-april-12025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/computer-vision-developer-hour-april-1-22025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/computer-vision-developer-hour-april-152025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/computer-vision-developer-hour-april-222025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/computer-vision-developer-hour-april-292025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/chicago-ai-ml-and-computer-vision-meetup-april-10-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/berlin-ai-ml-computer-vision-meetup-april-25-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-workshop-april-23-20252025-08-04T14:53:39Zmonthly0.8https://voxel51.com/events/getting-started-with-fiftyone-workshop-april-16-20252025-08-06T23:34:55Zmonthly0.8https://voxel51.com/events/building-visual-ai-in-the-enterprise-workshop-april-30-20252025-08-06T23:36:49Zmonthly0.8https://voxel51.com/events/ai-ml-and-computer-vision-meetup-may-22-20252025-08-04T14:19:27Zmonthly0.8https://voxel51.com/events/munich-ai-ml-and-computer-vision-meetup-april-24-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/sds-2025-workshop2025-05-27T08:07:50Zmonthly0.8https://voxel51.com/events/cologne-ai-ml-and-computer-vision-meetup-april-22-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/computer-vision-for-autonomous-driving-april-26-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/stuttgart-ai-ml-and-computer-vision-meetup-april-23-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/detecting-the-unexpected-practical-approaches-to-anomaly-detection-in-visual-data-workshop2025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/best-of-wacv-may-29-20252025-08-04T14:13:23Zmonthly0.8https://voxel51.com/events/best-of-wacv-may-30-20252025-08-04T14:12:16Zmonthly0.8https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-workshop-may-21-20252025-08-04T14:20:07Zmonthly0.8https://voxel51.com/events/getting-started-with-fiftyone-workshop-may-14-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/ai-ml-and-computer-vision-meetup-en-espanol-june-20-20252025-08-04T13:21:33Zmonthly0.8https://voxel51.com/events/ann-arbor-ai-ml-and-computer-vision-meetup-may-14-20252025-05-16T21:05:55Zmonthly0.8https://voxel51.com/events/visual-ai-in-healthcare-june-25-20252025-08-04T13:17:22Zmonthly0.8https://voxel51.com/events/midl-2025-workshop2025-06-05T09:37:59Zmonthly0.8https://voxel51.com/events/mosaic-ai-fiftyone-scaling-physical-ai-for-mobility-and-autonomous-use-cases-june-17-20252025-08-04T14:10:09Zmonthly0.8https://voxel51.com/events/computer-vision-developer-hour-may-132025-05-16T21:20:00Zmonthly0.8https://voxel51.com/blog/author/mt\_admin2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/mt\_pierce2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/stgvoxel512025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/travis2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/kim2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/jason2025-08-12T00:21:12Zmonthly0.8https://voxel51.com/blog/author/eric2025-07-15T07:07:11Zmonthly0.8https://voxel51.com/blog/author/benjamin2025-07-15T07:04:54Zmonthly0.8https://voxel51.com/blog/author/dave2025-07-15T07:06:26Zmonthly0.8https://voxel51.com/blog/author/tarmily2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/vini2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/mt\_andrew2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/wpengine2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/yuheng-li2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/chien-vu2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/allen2025-07-15T07:04:24Zmonthly0.8https://voxel51.com/blog/author/leila2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/ritchie2025-07-15T07:14:31Zmonthly0.8https://voxel51.com/blog/author/jon2025-07-15T07:09:57Zmonthly0.8https://voxel51.com/blog/author/robert2025-08-12T00:21:16Zmonthly0.8https://voxel51.com/blog/author/prerna2025-07-15T07:14:03Zmonthly0.8https://voxel51.com/blog/author/markus2025-07-15T07:12:39Zmonthly0.8https://voxel51.com/blog/author/mt\_cathy2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/jimmy2025-08-12T00:25:15Zmonthly0.8https://voxel51.com/blog/author/jacob2025-07-15T07:08:12Zmonthly0.8https://voxel51.com/blog/author/dillon2025-07-15T07:06:48Zmonthly0.8https://voxel51.com/blog/author/dan2025-08-12T00:23:17Zmonthly0.8https://voxel51.com/blog/author/harpreet2025-08-12T00:21:19Zmonthly0.8https://voxel51.com/blog/author/paularamos2025-08-12T00:21:22Zmonthly0.8https://voxel51.com/blog/author/vallyn2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/brent2025-08-12T00:35:27Zmonthly0.8https://voxel51.com/blog/author/manushree2025-08-12T00:41:05Zmonthly0.8https://voxel51.com/blog/author/jacob-sela2025-08-12T00:41:33Zmonthly0.8https://voxel51.com/blog/author/mike2025-07-15T07:13:10Zmonthly0.8https://voxel51.com/blog/author/monica2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/sliday2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/amanda-seo-consultant2025-07-15T07:04:19Zmonthly0.8https://voxel51.com/blog/author/sofia-sliday2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/steve2025-07-15T07:16:24Zmonthly0.8https://voxel51.com/blog/author/oscarsidebo2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/oskarengstrom2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/author/kirtivoxel51-com2025-07-15T07:10:46Zmonthly0.8https://voxel51.com/blog/author/nickvoxel51-com2025-08-12T00:25:19Zmonthly0.8https://voxel51.com/blog/author/brianvoxel51-com2025-08-12T00:27:25Zmonthly0.8https://voxel51.com/blog/author/deanlee808gmail-com2025-05-15T14:21:35Zmonthly0.8https://voxel51.com/blog/category/uncategorized2025-05-19T05:46:49Zmonthly0.8https://voxel51.com/blog/category/tutorials2025-05-19T05:46:49Zmonthly0.8https://voxel51.com/blog/category/vector-search2025-05-19T05:46:49Zmonthly0.8https://voxel51.com/blog/category/industry-solutions2025-05-19T05:46:49Zmonthly0.8https://voxel51.com/blog/category/plugins2025-05-19T05:46:49Zmonthly0.8https://voxel51.com/blog/category/integrations2025-05-19T05:46:49Zmonthly0.8https://voxel51.com/blog/category/mlvoxel512025-05-19T05:46:49Zmonthly0.8https://voxel51.com/blog/category/datasets2025-05-19T05:46:49Zmonthly0.8https://voxel51.com/blog/category/event-recaps2025-05-19T05:46:49Zmonthly0.8https://voxel51.com/blog/category/product-news2025-05-20T23:39:17Zmonthly0.8https://voxel51.com/blog/category/tips-tricks2025-05-19T07:02:39Zmonthly0.8https://voxel51.com/blog/category/computer-vision2025-05-19T05:46:49Zmonthly0.8https://voxel51.com/blog/tag/multi-modal-ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/transformers2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/year-in-review2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/carnegie-mellon2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/computer-vision-meetup2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/data-annotation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/qdrant2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/similarity-learning2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/wearable-vision-sensors2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/filtering2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dataset-curation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dataset-improvement2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/integrations2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/aggregations2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/pandas2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/pandas-style-queries2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/interactive-plots2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/autonomous-vehicles2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/end2end-learning2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/retail-use-case2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/supply-chain-use-case2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/synthetic-data2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/synthetic-data-generator2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/iou2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mistakenness2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/viewfield2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/exporting2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/community-rewards2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/anchor-boxes2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/embeddings2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/metadata2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-brain2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/persisting-datasets2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/video-datasets2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/voxel51-birthday2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/voxel51-milestone2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mongodb2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/pytorch2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/sorting2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/hacktoberfest2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/open-source2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-0-172025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/custom-datasets2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/tiff-images2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/funding2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/series-a2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/custom-plugins2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/importing2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/oss-community2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/success-story2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/labels2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/bdd100k2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/berkeley-deep-drive2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/clip2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/natural-language-processing2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/nlp2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/openai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/pinecone2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/classification2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/nearest-neighbors2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/data-centric-machine-learning2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/data-centric-ml2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mlops2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mlops-meetup2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/kinetics2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/video-classification-models2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/clearml2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/deep-learning2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/experiment-tracking2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/object-detection2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/computer-vision2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/activitynet2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/model-evaluation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/view-stages2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/viewstage2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/jupyter-notebook2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/evaluation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/adding-data2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/merging-data2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-app2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/tensorflow2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/annotation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cifar-102025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/lightning-flash2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mnist2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/coco2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/labelbox2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mobilenet2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/model-zoo2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/umap2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/avatar2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/frame-interpolation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/kitti-dataset2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/performance-capture2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/sensor-fusion2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/stereoscopic-vision2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/hugging-face2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/hyperparameter-scheduling2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/hyperparameters2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vision-transformers2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vit2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/anniversary2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/image-dataset2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/images2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/open-images2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/open-images-v62025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/colab-notebook2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/notebooks2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/company-culture2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/developer-relations2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/hiring2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/open-positions2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/open-roles2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/people-at-voxel512025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/agriculture2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/agriculture-use-case2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/industry-spotlight2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/use-case2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/computer-vision-news-recap2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cookiecutter2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/docker2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/github-actions2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/poetry2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/geojson2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/labeling-mistakes2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/instance-segmentation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ms-coco2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ultralytics2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/yolo2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/yolov82025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mean-average-precision2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/albumentations2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/diffusion-models2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fine-tune-models2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/gans2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/anomalib2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/asr2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/edge-ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/openai-whisper-model2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/openvino2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/speech-recognition2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/whisper2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/whisper-model2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-0-192025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/on-disk-segmentations2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/saved-views2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/spaces2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ui-filtering2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/community-update2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/action-recognition2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ucf1012025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/heatmaps2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/segmentations2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/google2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/keypoints2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/open-images-v72025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/point-labels2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/anomaly-detection2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/bin-picking2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cognex2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/datalogic2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/defect-detection2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/depalletizing2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/industrial-automation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/instrumental-ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/machine-tending2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/machine-vision2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/manufacturing2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/matroid2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mech-mind-robotics2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/palletizing2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/pickit-3d2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/predictive-maintenance2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/preml-gmbh2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/prophesee2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/protex-ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/rios-intelligent-machines2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/stemmer-imaging2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/quickstart-dataset2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/quickstart-video-dataset2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/image-restoration2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/low-light-image-enhancement2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/models-in-production2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/recycling-max-pooling2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/rmp2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cityscapes2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/semantic-urban-scene-understanding2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/events2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/getting-started-with-fiftyone2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/training2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/workshop2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/workshop-series2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/3d-point-cloud2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/birds-eye-view2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dbscan2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/point-cloud2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/point-cloud-synthesis2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/point-e2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/quickstart-groups-dataset2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/natural-language-search2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/semantic-search2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/similarity-search2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vector-database2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/regex2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/3d2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-0-202025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/nlp-search2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vector-search-database2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vector-search-engines2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-teams-1-22025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/getting-started-with-fiftyone-workshop2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/computer-vision-events2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/video-data-analysis2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/gligen2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/grounded-language-to-image-generation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/text-to-image2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/text-to-image-diffusion-models2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/image-matching-challenge-20232025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/kaggle2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/kaggle-competition2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/datasets2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/filtered-views2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/machine-learning2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/newsletter2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cotracker32025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/point-tracking2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/gpt2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/human-motion-synthesis2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mocap2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/motion-capture2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/motion-estimation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/motion-prediction2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/t2m-gpt2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vq-vae2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/hyperparameter-sweep2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/model-selection2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/model-training2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/weights-biases2025-05-22T04:26:45Zmonthly0.8https://voxel51.com/blog/tag/attention2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/amazon-dataset2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/armbench2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dataset-visualization2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/object-segmentation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/pick-and-place-robots2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/robotics2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/robots2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/deci-ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/supergradient2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/yolo-nas2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/model-predictions2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/uniqueness2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/voc-annotation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/voc-label2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cvf2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cvpr2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ieee2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/midjourney2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/nerf2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dataset-views2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/object-patches2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/sidebar2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dreambooth2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/f2-nerf2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/imagebind2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mask-dino2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mobilenerf2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/sadtalker2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/soldier2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/videofusion2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/opendoor2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/real-estate2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/wildlife-ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/color-schemes2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dynamic-groups2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-0-212025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/operators2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/plugins2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/langchain2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/voxelgpt2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cvpr-20232025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/tradeshows2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/sama-coco2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/yolov52025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai-stack2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/data-centric-ai-tooling2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/try-fiftyone-ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/compute\_hardness2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/custom-color-schemes2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/webinar2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/arkittrack2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/geonet2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/imagenet-e2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/jrdb-pose2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/llcm2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mobile-hdr2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/mvimgnet2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/spring2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/synsl-120k2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/caption2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/controlnet2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/google-conceptual-captions2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/multimodal2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/lancedb2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/milvus2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/meetup2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dreamsim2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/facial-video-representation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/human-perception2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/marlin2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/nights-dataset2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vector-search2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/lpips2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/nights2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/perceptual-metrics2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/similarity2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/keypoint-skeletons2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fields2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/sample-fields2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/medical-imaging2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/neural-congealing2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/radiotherapy2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/data-curation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/competition2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/opencv2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-community2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai-art2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dalle22025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/genai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/generative-ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/replicate2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/stable-diffusion2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vqgan2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/artificial-intelligence2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/computer-aided-diagnosis2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/disease-detection2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/healthcare2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/indiustry-spotlight2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/medicine2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/segment-anything2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/surgical-guidance2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cli2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/blip2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/visual-question-answering2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vqa2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/groups2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/drones2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/javascript2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/material-ui2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/react2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/youtube2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/group-datasets2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai-machine-learning-data-science-meetup2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai-ml-ds-meetup2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/meetups2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/benchmark2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/bias2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/disparity2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fairness2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/meta2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/sa1b2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/deduplication2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/polylines2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/iccv2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/lidar2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/iccv232025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/badger2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/custom-badges2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/python-library2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/labeling2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/pose-estimation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/tag-12025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/asl2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/zilliz2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/tips-tricks2025-05-22T04:27:40Zmonthly0.8https://voxel51.com/blog/tag/video-object-tracking2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/neurips2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/partnership2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/v72025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai-referee2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fitness2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/industry-use-cases2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/injury-prevention2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/rehabilitation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/sports2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/sports-analytics2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/sports-fan-enhancement2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/autonomous-checkout2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/customer-behavior-analysis2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/inventory-management2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/product-recommendations2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/recommender-engines2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/retail2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/retail-industry2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/supply-chain2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/virtual-try-ons2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dimensionality-reduction2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/pca2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/resnet502025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/tsne2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/visualization2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dynamic-attributes2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/facial-recognition2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/intruder-detection2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/safety-security2025-05-22T04:27:51Zmonthly0.8https://voxel51.com/blog/tag/safety-equipment-recognition2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/search-and-rescue2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/security-checkpoint-screening2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/security-industry2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/nvidia2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/nvidia-omniverse2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai-excellence-award2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/awards2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/detector-model2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/rios2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/iso-270012025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/iso-27001-certification2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/adt-commercial2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dataset-management2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/everon2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/autonomous-navigation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/environmental-monitoring2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/object-recognition2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/quality-control2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/robotics-industry2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/surveillance2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/visual-inspection2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/text2image2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/llms2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/series-b2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/bessemer-venture-partners2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/bvp2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/3d-object-generation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/shap-e2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/3d-geometries2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/3d-meshes2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/custom-workspaces2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-0-242025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/llama22025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cvpr-20242025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/adas2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/automotive2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/driving-use-case2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/in-cabin-perception2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/zero-shot2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/medical-imagery2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/hugging-face-competition2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/visual-ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/what-is-visual-ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/rag-models2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/sam-22025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/custom-dashboards2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/elasticsearch2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-0-252025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/python-panels2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/data-annotation-workflows2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/segments-ai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-1-02025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/data-quality2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/high-quality-data2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/eccv2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/naturalbench2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/neurlps2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vlm2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vqa-benchmarks2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vqav22025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/image-analysi2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/knobo2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/medical2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/medical-ai-model2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/medical-education2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/neurlps-20242025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dataset-distillation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/soft-labels2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/spiqa2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/visualai2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/builtin-compute2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/data-lens2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/panels2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/query-performance2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/leaky-splits2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/leaky-splits-analysis2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai-in-chemistry2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai-in-natural-sciences2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai-in-physics2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai-in-science2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cifake2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/data-bias2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/embeddings-comparison2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/synthetic-vs-real-data2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/health2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/coreset-selection2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/zcore2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/zero-shot-coreset-selection2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/zero-shot-data-reduction2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/class-wise-autoencoders2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ml-research2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/reconstruction-error-ratios2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/rers2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/janus-pro2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/gaussian-splatting2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/gaussian-splats2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/3d-reconstruction2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/imagenet2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/imagenet-d2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/aimv22025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/zero-shot-classification2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/bioscan-5m2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/bioclip2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/barcodebert2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/geolocation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/webuot-1m2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/video-embeddings2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/hiera2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/text-embeddings2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/jina-embeddings-v32025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/visual-spectrogram-classification2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/vsc2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/esc-502025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/esc-50-dataset2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/music2latent2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/clap2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/spectrograms2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/moondream22025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/meme-understanding2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ocr2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/contextual-caption-generation2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/siglip-22025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-enterprise2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/model-certainty-visualization2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/dataset-zoo2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/families-in-the-wild2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiw2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-0-182025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/webinar-recap2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/case-study2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/fiftyone-teams2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/product-release2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/bounding-boxes2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/exporting-with-splits2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/faq2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/large-datasets2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/cvat2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/detections2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/grouped-datasets2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/label-studio2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/point-clouds2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/chatgpt2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/ai-generated-artwork2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/computer-vision-applications2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/computer-vision-research2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/computer-vision-trends2025-05-19T05:50:26Zmonthly0.8https://voxel51.com/blog/tag/data-centric-computer-vision2025-05-19T05:50:26Zmonthly0.8 https://voxel51.com/curation 2025-08-19T11:10:32Z monthly 0.9 ... https://voxel51.com/industries/manufacturing 2025-08-08T16:56:14Z monthly 0.9 ... https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections 2025-07-04T13:05:27Z monthly 0.8 ... https://voxel51.com/customers/rios 2025-08-22T16:52:10Z monthly 0.8 ... https://voxel51.com/customers/metu 2025-05-29T15:00:46Z monthly 0.8 ... https://voxel51.com/customers/updata 2025-05-29T14:16:20Z monthly 0.8 ... https://voxel51.com/events/visual-ai-in-healthcare-june-27-2025 2025-08-04T13:15:38Z monthly 0.8 ... https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately 2025-07-21T07:52:58Z monthly 0.8 ... https://voxel51.com/whitepapers/the-best-data-centric-computer-vision-tools-for-the-enterprise 2025-08-14T01:30:08Z monthly 0.8 ... https://voxel51.com/customers?category=av-physical-ai 2025-05-31T18:26:02Z monthly 0.8 ... https://voxel51.com/events/visual-ai-in-manufacturing-september-10-2025 2025-08-05T07:16:14Z monthly 0.8 ... https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning 2025-07-21T08:31:08Z monthly 0.8 ... https://voxel51.com/customers/ibm 2025-08-22T16:41:36Z monthly 0.8 ... https://voxel51.com/customers/binit 2025-05-29T14:01:32Z monthly 0.8 ... https://voxel51.com/point-cloud 2025-07-10T17:30:51Z monthly 0.9 ... https://voxel51.com/events/from-research-to-reality-building-gui-agents-that-actually-work-august-22-2025 2025-07-16T19:03:54Z monthly 0.8 ... https://voxel51.com/fiftyone 2025-07-18T17:45:52Z monthly 0.9 ... https://voxel51.com/industries/sports 2025-06-06T08:02:32Z monthly 0.9 ... https://voxel51.com/customers/fast-code-ai 2025-05-29T14:55:38Z monthly 0.8 ... https://voxel51.com/customers?category=retail-consumer 2025-05-31T18:26:06Z monthly 0.8 ... https://voxel51.com/careers 2025-06-02T18:59:39Z monthly 0.9 ... https://voxel51.com/events/boston-ai-ml-and-computer-vision-meetup-workshop-september-25-2025 2025-08-04T18:02:54Z monthly 0.8 ... https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub 2025-07-18T16:03:34Z monthly 0.8 ... https://voxel51.com/about 2025-06-18T01:42:24Z monthly 0.9 ... https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1 2025-06-09T14:59:15Z monthly 0.8 ... https://voxel51.com/events/tokyo-ai-ml-and-computer-vision-meetup-july-31 2025-07-28T23:34:58Z monthly 0.8 ... https://voxel51.com/industries/retail 2025-06-06T08:01:46Z monthly 0.9 ... https://voxel51.com/events/brussels-ai-ml-and-computer-vision-meetup-july-15 2025-06-24T22:49:48Z monthly 0.8 ... https://voxel51.com/events/ai-ml-and-computer-vision-meetup-june-19-2025 2025-08-04T14:09:33Z monthly 0.8 ... https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier 2025-06-30T22:22:30Z monthly 0.8 ... https://voxel51.com/industries/agriculture 2025-06-02T06:46:31Z monthly 0.9 ... https://voxel51.com/blog/cvpr-2025 2025-06-04T23:30:51Z monthly 0.8 ... https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation 2025-07-29T21:04:11Z monthly 0.8 ... https://voxel51.com/whitepapers/why-vision-ai-models-fails 2025-08-14T01:29:44Z monthly 0.8 ... https://voxel51.com/customers/aquabyte 2025-05-29T14:06:07Z monthly 0.8 ... https://voxel51.com/industries/robotics 2025-06-06T08:01:59Z monthly 0.9 ... https://voxel51.com/blog/how-we-built-annotation-savings-estimator 2025-06-18T19:07:22Z monthly 0.8 ... https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities 2025-08-12T01:05:34Z monthly 0.8 ... https://voxel51.com/research 2025-08-05T06:34:56Z monthly 0.9 ... https://voxel51.com/industries/autonomous-vehicles-systems 2025-06-06T07:50:19Z monthly 0.9 ... https://voxel51.com/blog/comprehensive-guide-point-cloud-data 2025-07-25T15:32:00Z monthly 0.8 ... https://voxel51.com/events/scaling-computer-vision-ai-in-the-enterprise-28-august-2025 2025-08-18T17:53:30Z monthly 0.8 ... https://voxel51.com/data-centric-visual-ai 2025-05-29T10:25:17Z monthly 0.9 ... https://voxel51.com/blog/nvidia-ai-podcast-adas-with-porsche 2025-07-30T16:57:42Z monthly 0.8 ... https://voxel51.com/customers/secury360 2025-08-12T00:58:47Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies 2025-07-17T03:35:36Z monthly 0.8 ... https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone 2025-08-11T15:00:10Z monthly 0.8 ... https://voxel51.com/events/ai-ml-and-computer-vision-meetup-july-17-2025 2025-08-04T12:55:04Z monthly 0.8 ... https://voxel51.com/blog/tag/visual-agents 2025-06-30T22:21:31Z monthly 0.8 ... https://voxel51.com/blog/tag/auto-labeling 2025-06-03T09:21:49Z monthly 0.8 ... https://voxel51.com/events/from-research-to-reality-building-gui-agents-that-actually-work-august-15-2025 2025-07-16T19:04:14Z monthly 0.8 ... https://voxel51.com/events/getting-started-with-fiftyone-for-healthcare-use-cases-july-23-2025 2025-08-04T12:53:48Z monthly 0.8 ... https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation 2025-07-21T08:55:03Z monthly 0.8 ... https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape 2025-07-22T14:38:58Z monthly 0.8 ... https://voxel51.com/events/building-visual-ai-in-the-enterprise-workshop-june-4-2025 2025-08-06T23:36:35Z monthly 0.8 ... https://voxel51.com/events/getting-started-with-fiftyone-workshop-june-18-2025 2025-08-04T14:08:17Z monthly 0.8 ... https://voxel51.com/blog/a-guide-to-ai-image-segmentation 2025-07-21T08:26:37Z monthly 0.8 ... https://voxel51.com/blog/category/press 2025-05-20T15:41:34Z monthly 0.8 ... https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services 2025-05-22T01:53:59Z monthly 0.8 ... https://voxel51.com/customers/aidence 2025-05-29T14:39:53Z monthly 0.8 ... https://voxel51.com/blog/smarter-automotive-datasets-selection 2025-07-24T16:51:40Z monthly 0.8 ... https://voxel51.com/customers/kitro 2025-05-29T14:10:02Z monthly 0.8 ... https://voxel51.com/pricing 2025-08-18T16:59:51Z monthly 0.9 ... https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail 2025-07-21T07:32:05Z monthly 0.8 ... https://voxel51.com/customers/indian-institute-of-it-and-management 2025-05-29T14:14:16Z monthly 0.8 ... https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3 2025-06-09T15:07:31Z monthly 0.8 ... https://voxel51.com/events/raleigh-ai-ml-and-computer-vision-meetup-august-20-2025 2025-07-23T14:36:23Z monthly 0.8 ... https://voxel51.com/events/best-of-cvpr-july-9-2025 2025-08-04T13:12:38Z monthly 0.8 ... https://voxel51.com/customers/taranis 2025-05-29T14:52:30Z monthly 0.8 ... https://voxel51.com/blog/how-automated-data-labeling-enhances-computer-vision-efficiency-and-accuracy 2025-07-21T09:10:23Z monthly 0.8 ... https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs 2025-06-05T16:44:08Z monthly 0.8 ... https://voxel51.com/customers/forsight 2025-08-22T16:44:18Z monthly 0.8 ... https://voxel51.com/customers?category=dev-tools 2025-05-27T22:05:37Z monthly 0.8 ... https://voxel51.com/events/amsterdam-ai-ml-and-computer-vision-meetup-july-14 2025-06-25T13:44:05Z monthly 0.8 ... https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai 2025-06-12T16:07:55Z monthly 0.8 ... https://voxel51.com/events/visual-ai-in-manufacturing-how-multimodal-data-powers-adaptive-process-control 2025-08-04T23:47:14Z monthly 0.8 ... https://voxel51.com/customers/rif-robotics 2025-05-29T13:55:24Z monthly 0.8 ... https://voxel51.com/events/ai-ml-and-computer-vision-meetup-en-espanol-october-23-2025 2025-08-18T17:09:13Z monthly 0.8 ... https://voxel51.com/whitepapers/visual-ai-for-defect-detection-in-manufacturing 2025-08-14T01:29:12Z monthly 0.8 ... https://voxel51.com/customers/adt-commercial 2025-08-22T16:39:05Z monthly 0.8 ... https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness 2025-07-21T08:51:21Z monthly 0.8 ... https://voxel51.com/blog/what-makes-good-data-a-view-from-the-front-lines-of-ai 2025-07-17T16:19:10Z monthly 0.8 ... https://voxel51.com/events/from-research-to-reality-building-gui-agents-that-actually-work-august-29-2025 2025-07-16T19:03:37Z monthly 0.8 ... https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation 2025-06-27T19:15:55Z monthly 0.8 ... https://voxel51.com/blog/the-complete-guide-to-auto-labeling 2025-07-21T07:54:15Z monthly 0.8 ... https://voxel51.com/customers/ai-fish 2025-05-29T14:03:48Z monthly 0.8 ... https://voxel51.com/voxelgpt 2025-05-29T06:16:23Z monthly 0.9 ... https://voxel51.com/events/advanced-car-damage-detection-with-fiftyone-and-the-cardd-dataset-july-12 2025-06-29T20:12:14Z monthly 0.8 ... https://voxel51.com/events/valencia-ai-ml-and-computer-vision-meetup-september-25-2025 2025-08-05T16:40:35Z monthly 0.8 ... https://voxel51.com/events/verified-auto-labeling-smarter-annotation-at-scale-june-24-2025 2025-08-04T13:17:50Z monthly 0.8 ... https://voxel51.com/blog/vggt-is-a-pure-neural-approach-to-3d-vision 2025-06-26T06:12:03Z monthly 0.8 ... https://voxel51.com/customers/allstate 2025-08-22T16:42:28Z monthly 0.8 ... https://voxel51.com/blog/tag/video 2025-06-03T09:56:20Z monthly 0.8 ... https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision 2025-07-21T09:07:06Z monthly 0.8 ... https://voxel51.com/events/visual-ai-in-manufacturing-and-robotics-september-11-2025 2025-08-11T17:35:39Z monthly 0.8 ... https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks 2025-08-20T14:22:52Z monthly 0.8 ... https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using 2025-06-23T07:57:21Z monthly 0.8 ... https://voxel51.com/events/boston-ai-ml-and-computer-vision-meetup-june-26-2025 2025-06-06T12:47:49Z monthly 0.8 ... https://voxel51.com/events/virtual-how-porsche-uses-auto-labeling-to-supercharge-av-development-4-september-2025 2025-08-18T17:53:01Z monthly 0.8 ... https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai 2025-07-03T20:59:41Z monthly 0.8 ... https://voxel51.com/events/ai-ml-and-computer-vision-meetup-aug-28-2025 2025-08-14T08:16:54Z monthly 0.8 ... https://voxel51.com/events/visual-ai-in-manufacturing-and-robotics-september-12-2025 2025-08-21T19:15:12Z monthly 0.8 ... https://voxel51.com/events/preventing-critical-misses-in-defect-detection-a-data-centric-approach-with-mongodb-and-voxel51 2025-08-18T17:53:15Z monthly 0.8 ... https://voxel51.com/blog/tag/cost-estimation 2025-06-03T09:42:20Z monthly 0.8 ... https://voxel51.com/industries/security 2025-06-06T08:02:12Z monthly 0.9 ... https://voxel51.com/blog/deploy-computer-vision-in-manufacturing 2025-08-06T19:40:53Z monthly 0.8 ... https://voxel51.com/customers?category=security 2025-05-27T21:57:12Z monthly 0.8 ... https://voxel51.com/industries/defense 2025-06-06T07:54:50Z monthly 0.9 ... https://voxel51.com/events/getting-started-with-fiftyone-for-manufacturing-use-cases-sept-30-2025 2025-08-04T18:00:00Z monthly 0.8 ... https://voxel51.com/blog/uncommon-objects-in-3d 2025-06-18T21:17:05Z monthly 0.8 ... https://voxel51.com/events/madrid-ai-ml-and-computer-vision-meetup-september-26-2025 2025-08-05T16:08:13Z monthly 0.8 ... https://voxel51.com/blog/opening-remarks-from-cvpr-2025 2025-06-20T21:30:21Z monthly 0.8 ... https://voxel51.com/customers/raytheon-technologies 2025-08-22T16:51:21Z monthly 0.8 ... https://voxel51.com/events/exposing-your-datas-blind-spots-scenario-mining-for-safer-av 2025-08-05T14:23:15Z monthly 0.8 ... https://voxel51.com/customers/vivint 2025-05-29T14:42:59Z monthly 0.8 ... https://voxel51.com/events/workshop-train-a-medical-ai-model-in-one-day-july-25-2025 2025-07-14T16:26:06Z monthly 0.8 ... https://voxel51.com/events/women-in-ai-july-24 2025-08-04T12:51:07Z monthly 0.8 ... https://voxel51.com/sales 2025-05-22T10:22:48Z monthly 0.9 ... https://voxel51.com/blog/tag/agi 2025-06-17T15:49:05Z monthly 0.8 ... https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition 2025-07-21T09:21:02Z monthly 0.8 ... https://voxel51.com/blog 2025-07-30T16:41:44Z monthly 0.8 ... https://voxel51.com/blog/category/learn 2025-07-21T07:23:09Z monthly 0.8 ... https://voxel51.com/community 2025-07-17T16:47:53Z monthly 0.9 ... https://voxel51.com/events/women-in-ai-october-2-2025 2025-07-31T23:06:57Z monthly 0.8 ... https://voxel51.com/industries/healthcare 2025-06-24T20:37:20Z monthly 0.9 ... https://voxel51.com/customers/aisprid 2025-05-30T17:45:02Z monthly 0.8 ... https://voxel51.com/customers/g42 2025-05-29T14:21:20Z monthly 0.8 ... https://voxel51.com/events/food-waste-estimation-hackathon-computer-vision-for-sustainability-1-august-2025 2025-07-31T15:27:43Z monthly 0.8 ... https://voxel51.com/careers 2025-06-02T06:57:52Z monthly 0.8 ... https://voxel51.com/customers 2025-06-26T22:19:00Z monthly 0.8 ... https://voxel51.com/events/understanding-visual-agents-august-7-2025 2025-08-12T22:40:02Z monthly 0.8 ... https://voxel51.com/customers/protex-ai 2025-07-15T16:51:21Z monthly 0.8 ... https://voxel51.com/customers/qinecsa 2025-05-29T14:18:36Z monthly 0.8 ... https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data 2025-07-21T09:05:13Z monthly 0.8 ... https://voxel51.com/link-catcher 2025-05-12T20:01:57Z monthly 0.9 ... https://voxel51.com/evaluation 2025-08-13T19:41:15Z monthly 0.9 ... https://voxel51.com/customers/smart-eye 2025-08-22T16:38:01Z monthly 0.8 ... https://voxel51.com/blog/tag/verified-auto-labeling 2025-06-03T09:41:50Z monthly 0.8 ... https://voxel51.com/customers/safelyyou 2025-08-22T16:54:17Z monthly 0.8 ... https://voxel51.com/events/best-of-cvpr-july-11-2025 2025-08-04T12:57:22Z monthly 0.8 ... https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2 2025-06-09T15:06:31Z monthly 0.8 ... https://voxel51.com/customers/wildlife-ai 2025-05-29T13:57:59Z monthly 0.8 ... https://voxel51.com/customers/berkshire-grey 2025-08-22T16:50:24Z monthly 0.8 ... https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models 2025-07-21T09:14:50Z monthly 0.8 ... https://voxel51.com/customers/lancedb 2025-05-29T14:57:24Z monthly 0.8 ... https://voxel51.com/blog/tag/annotation-savings 2025-06-03T09:42:06Z monthly 0.8 ... https://voxel51.com/blog/visual-agents-at-cvpr-2025 2025-06-02T20:19:41Z monthly 0.8 ... https://voxel51.com/customers?category=aerospace-and-defense 2025-05-27T21:43:27Z monthly 0.8 ... https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape 2025-06-26T04:29:55Z monthly 0.8 ... https://voxel51.com/customers?category=health-and-medicine 2025-05-27T06:31:54Z monthly 0.8 ... https://voxel51.com/customers/seafar 2025-05-29T14:30:29Z monthly 0.8 ... https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic 2025-06-17T15:55:36Z monthly 0.8 ... https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance 2025-08-19T17:19:59Z monthly 0.8 ... https://voxel51.com/whitepapers/auto-labeling-data-for-object-detection 2025-08-14T01:30:27Z monthly 0.8 ... https://voxel51.com/industries/aviation 2025-06-06T07:50:30Z monthly 0.9 ... https://voxel51.com/events 2025-05-20T12:29:44Z monthly 0.8 ... https://voxel51.com/customers?category=agriculture-sustainability 2025-05-27T22:00:36Z monthly 0.8 ... https://voxel51.com/whitepapers/your-data-your-advantage 2025-08-14T01:28:26Z monthly 0.8 ... https://voxel51.com/ 2025-08-19T11:49:47Z monthly 1 ... https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025 2025-06-05T14:26:00Z monthly 0.8 ... https://voxel51.com/customers/ancera 2025-08-22T16:47:20Z monthly 0.8 ... https://voxel51.com/get-started 2025-06-06T17:47:36Z monthly 0.9 ... https://voxel51.com/blog/author/antonio-rueda-toicen 2025-08-12T00:44:52Z monthly 0.8 ... https://voxel51.com/customers/fyma 2025-05-29T14:34:14Z monthly 0.8 ... https://voxel51.com/customers/finegrain 2025-05-29T14:24:18Z monthly 0.8 ... https://voxel51.com/events/paris-ai-ml-and-computer-vision-meetup-july-16 2025-07-10T19:53:02Z monthly 0.8 ... https://voxel51.com/annotation 2025-08-19T17:22:38Z monthly 0.9 ... https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes 2025-07-21T08:38:31Z monthly 0.8 ... https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows 2025-07-21T09:03:27Z monthly 0.8 ... https://voxel51.com/events/ai-ml-and-computer-vision-meetup-en-espanol-august-21-2025 2025-07-31T23:11:46Z monthly 0.8 ... https://voxel51.com/customers/argosai 2025-06-03T20:26:32Z monthly 0.8 ... https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai 2025-08-01T21:31:38Z monthly 0.8 ... https://voxel51.com/glossary 2025-05-28T07:18:10Z monthly 0.8 ... https://voxel51.com/integrations 2025-08-12T00:05:52Z monthly 0.8 ... https://voxel51.com/plugins 2025-08-12T00:05:17Z monthly 0.8 ... https://voxel51.com/press 2025-07-24T16:47:07Z monthly 0.8 ... https://voxel51.com/webinars 2025-06-02T20:52:14Z monthly 0.8 ... https://voxel51.com/whitepapers 2025-07-29T08:55:19Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-open-images-collaboration 2025-05-20T15:42:44Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-open-source-launch 2025-05-20T15:42:44Z monthly 0.8 ... https://voxel51.com/blog/voxel51-physical-distancing-index 2025-05-20T15:42:44Z monthly 0.8 ... https://voxel51.com/blog/voxel51-launches-image-to-video-tool 2025-05-20T15:42:44Z monthly 0.8 ... https://voxel51.com/blog/voxel51-raises-2-million-to-advance-video-understanding 2025-05-20T15:42:44Z monthly 0.8 ... https://voxel51.com/blog/introducing-voxelgpt 2025-05-20T15:42:44Z monthly 0.8 ... https://voxel51.com/blog/voxel51-raises-30m-series-b-funding-to-make-visual-ai-a-reality 2025-05-20T15:42:44Z monthly 0.8 ... https://voxel51.com/blog/voxel51-launches-fiftyone-open-source-1-0-accelerating-the-creation-of-production-ready-visual-ai-applications 2025-05-20T15:42:44Z monthly 0.8 ... https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset 2025-05-22T04:24:10Z monthly 0.8 ... https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision 2025-05-22T04:24:10Z monthly 0.8 ... https://voxel51.com/blog/forsight-finds-a-centralized-dataset-management-solution-in-fiftyone-teams 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes 2025-05-22T04:24:21Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022 2025-05-22T04:24:10Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022 2025-05-22T04:23:56Z monthly 0.8 ... https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see 2025-05-22T04:23:56Z monthly 0.8 ... https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far 2025-05-22T04:23:56Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022 2025-05-22T04:24:10Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022 2025-05-22T04:24:10Z monthly 0.8 ... https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools 2025-05-22T04:24:10Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022 2025-05-22T04:24:10Z monthly 0.8 ... https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision 2025-05-22T04:24:10Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022 2025-05-22T04:24:10Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022 2025-05-22T04:24:10Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-meetup-update-november-22 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/its-our-birthday-voxel51-turns-four 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/hack-on-fiftyone-in-hacktoberfest-2022 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/announcing-our-12-5m-series-a-funding-to-bring-transparency-and-clarity-to-the-worlds-data 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022 2025-05-22T04:03:08Z monthly 0.8 ... https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data 2025-05-22T04:23:56Z monthly 0.8 ... https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022 2025-05-22T04:23:56Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022 2025-05-22T04:23:56Z monthly 0.8 ... https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51 2025-05-22T04:23:56Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023 2025-05-22T04:23:56Z monthly 0.8 ... https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset 2025-05-22T04:23:40Z monthly 0.8 ... https://voxel51.com/blog/finding-images-with-words 2025-05-22T04:23:40Z monthly 0.8 ... https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/meetup-recap-how-to-build-high-quality-machine-learning-datasets-and-computer-vision-models 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/how-to-train-your-dragon-detector 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023 2025-05-22T04:23:40Z monthly 0.8 ... https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023 2025-05-22T04:23:40Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023 2025-05-22T04:23:40Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023 2025-05-22T04:23:34Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023 2025-05-22T04:23:34Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023 2025-05-22T04:23:29Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023 2025-05-22T04:23:21Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023 2025-05-22T04:23:04Z monthly 0.8 ... https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/the-making-of-avatar-the-way-of-water 2025-05-22T04:23:40Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023 2025-05-22T04:23:40Z monthly 0.8 ... https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation 2025-05-22T04:03:25Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-turns-one 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation 2025-05-22T04:03:21Z monthly 0.8 ... https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone 2025-05-22T04:03:25Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-six-months-post-launch 2025-05-22T04:03:25Z monthly 0.8 ... https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision 2025-05-22T04:03:25Z monthly 0.8 ... https://voxel51.com/blog/people-voxel51-spotlight-on-jimmy-guerrero 2025-05-22T04:23:40Z monthly 0.8 ... https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023 2025-08-12T01:03:17Z monthly 0.8 ... https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free 2025-05-22T04:23:34Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023 2025-05-22T04:23:34Z monthly 0.8 ... https://voxel51.com/blog/giving-yolov8-a-second-look-part-1 2025-05-22T04:23:29Z monthly 0.8 ... https://voxel51.com/blog/giving-yolov8-a-second-look-part-2 2025-05-22T04:23:29Z monthly 0.8 ... https://voxel51.com/blog/giving-yolov8-a-second-look-part-3 2025-05-22T04:23:34Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap 2025-05-22T04:23:34Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-0-19 2025-05-22T04:23:34Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023 2025-05-22T04:23:34Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023 2025-05-22T04:23:29Z monthly 0.8 ... https://voxel51.com/blog/people-voxel51-spotlight-on-lanny-wang 2025-05-22T04:23:29Z monthly 0.8 ... https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset 2025-05-22T04:23:29Z monthly 0.8 ... https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision 2025-05-22T04:23:29Z monthly 0.8 ... https://voxel51.com/blog/exploring-google-open-images-v7 2025-05-22T04:23:29Z monthly 0.8 ... https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing 2025-05-22T04:23:29Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023 2025-05-22T04:23:21Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023 2025-05-22T04:23:21Z monthly 0.8 ... https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding 2025-05-22T04:23:21Z monthly 0.8 ... https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series 2025-05-22T04:23:21Z monthly 0.8 ... https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e 2025-05-22T04:23:21Z monthly 0.8 ... https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data 2025-05-22T04:23:21Z monthly 0.8 ... https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap 2025-05-22T04:23:04Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023 2025-05-22T04:23:04Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023 2025-05-22T04:23:04Z monthly 0.8 ... https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen 2025-05-22T04:23:04Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023 2025-05-22T04:23:04Z monthly 0.8 ... https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset 2025-05-22T04:23:04Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023 2025-05-22T04:22:56Z monthly 0.8 ... https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision 2025-05-22T04:22:56Z monthly 0.8 ... https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap 2025-05-22T04:22:56Z monthly 0.8 ... https://voxel51.com/blog/generate-movement-from-text-descriptions-with-t2m-gpt 2025-05-22T04:22:56Z monthly 0.8 ... https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone 2025-05-22T04:22:56Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023 2025-05-22T04:22:56Z monthly 0.8 ... https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip 2025-05-22T04:22:56Z monthly 0.8 ... https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone 2025-05-22T04:22:56Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023 2025-05-22T04:22:56Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023 2025-05-22T04:22:48Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023 2025-05-22T04:22:48Z monthly 0.8 ... https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision 2025-05-22T04:22:48Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023 2025-05-22T04:22:48Z monthly 0.8 ... https://voxel51.com/blog/cvpr-2023-survival-guide 2025-05-22T04:22:48Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023 2025-05-22T04:22:48Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023 2025-05-22T04:22:48Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-0-21 2025-05-22T04:22:48Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023 2025-05-22T04:22:48Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023 2025-05-22T04:22:45Z monthly 0.8 ... https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision 2025-05-22T04:22:45Z monthly 0.8 ... https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr 2025-05-22T04:22:45Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023 2025-05-22T04:22:45Z monthly 0.8 ... https://voxel51.com/blog/too-many-pixels-so-little-time 2025-05-22T04:22:45Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023 2025-05-22T04:22:45Z monthly 0.8 ... https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins 2025-05-22T04:22:45Z monthly 0.8 ... https://voxel51.com/blog/visualize-cvpr-2023-datasets-at-cvpr-2023 2025-05-22T04:22:45Z monthly 0.8 ... https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr 2025-05-22T04:22:45Z monthly 0.8 ... https://voxel51.com/blog/conquering-controlnet 2025-05-22T04:22:45Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023 2025-05-22T04:22:45Z monthly 0.8 ... https://voxel51.com/blog/the-computer-vision-interface-for-vector-search 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/teaching-androids-to-dream-of-sheep 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/spending-my-first-week-with-fiftyone 2025-05-22T04:22:34Z monthly 0.8 ... https://voxel51.com/blog/opencv-ai-competition-2023 2025-05-22T04:22:34Z monthly 0.8 ... https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/celebrating-three-years-of-fiftyone 2025-05-22T04:22:39Z monthly 0.8 ... https://voxel51.com/blog/build-your-own-ai-art-gallery 2025-05-22T04:22:34Z monthly 0.8 ... https://voxel51.com/blog/how-computer-vision-is-changing-healthcare 2025-05-22T04:22:34Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023 2025-05-22T04:22:34Z monthly 0.8 ... https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023 2025-05-22T04:22:34Z monthly 0.8 ... https://voxel51.com/blog/ask-your-images-anything 2025-05-22T04:22:34Z monthly 0.8 ... https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023 2025-05-22T04:22:34Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-mastering-drone-data-training 2025-05-22T04:22:14Z monthly 0.8 ... https://voxel51.com/blog/build-custom-computer-vision-applications 2025-05-22T04:22:34Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023 2025-05-22T04:22:34Z monthly 0.8 ... https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023 2025-05-22T04:22:28Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023 2025-05-22T04:22:28Z monthly 0.8 ... https://voxel51.com/blog/facet-benchmark 2025-05-22T04:22:28Z monthly 0.8 ... https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone 2025-05-22T04:22:28Z monthly 0.8 ... https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023 2025-05-22T04:22:28Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract 2025-05-22T04:22:14Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration 2025-05-22T04:22:28Z monthly 0.8 ... https://voxel51.com/blog/nuscenes-dataset-navigating-the-road-ahead 2025-05-22T04:22:28Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023 2025-05-22T04:22:14Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide 2025-05-22T04:22:14Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone 2025-05-22T04:22:14Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023 2025-05-22T04:22:14Z monthly 0.8 ... https://voxel51.com/blog/3-reasons-to-visit-voxel51-at-iccv23 2025-05-22T04:22:14Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-badger-custom-github-badges 2025-05-22T04:22:14Z monthly 0.8 ... https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning 2025-05-22T04:22:14Z monthly 0.8 ... https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-2 2025-05-22T04:22:11Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone 2025-05-22T04:22:11Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-video-labels-fiftyone-tips-and-tricks-october-14th-2023 2025-05-22T04:22:11Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-oct-12-2023 2025-05-22T04:22:11Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-concept-traversal-plugin-for-fiftyone 2025-05-22T04:22:11Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-plugin-for-building-and-managing-plugins 2025-05-22T04:21:49Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-20-2023 2025-05-22T04:22:11Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-sam-for-prediction-kaggle-football-player-segmentation-dataset 2025-05-22T04:21:49Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-27-2023 2025-05-22T04:21:49Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-3-2023 2025-05-22T04:21:49Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-announcing-updates-to-fiftyone-0-22-2-and-fiftyone-teams-1-4-3 2025-05-22T04:22:11Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-10-2023 2025-05-22T04:21:29Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-10-weeks-of-building-fiftyone-plugins 2025-05-22T04:21:49Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-30-2023 2025-05-22T04:21:49Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-nov-2-2023 2025-05-22T04:21:49Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-fiftyone-0-22-3-and-fiftyone-teams-1-4-4 2025-05-22T04:21:49Z monthly 0.8 ... https://voxel51.com/blog/tracking-datasets-in-fiftyone 2025-05-22T04:21:29Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-24-2023 2025-05-22T04:21:29Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-17-2023 2025-05-22T04:21:29Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-plugins-tips-and-tricks-november-22-2023 2025-05-22T04:21:29Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-generating-videos-from-images-with-stable-video-diffusion-and-fiftyone 2025-05-22T04:21:29Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-elevate-your-github-readme-game 2025-05-22T04:21:29Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-0-23-and-fiftyone-teams-1-5 2025-05-22T04:21:29Z monthly 0.8 ... https://voxel51.com/blog/neurips-2023-and-the-state-of-ai-research 2025-05-22T04:21:29Z monthly 0.8 ... https://voxel51.com/blog/understanding-llava-large-language-and-vision-assistant 2025-05-22T04:21:25Z monthly 0.8 ... https://voxel51.com/blog/announcing-the-voxel51-v7-partnership 2025-05-22T04:21:25Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-dec-7-2023 2025-05-22T04:21:29Z monthly 0.8 ... https://voxel51.com/blog/neurips-2023-survival-guide 2025-05-22T04:21:29Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-and-fiftyone-community-year-in-review-2023 2025-05-22T04:21:25Z monthly 0.8 ... https://voxel51.com/blog/why-2023-was-the-most-exciting-year-in-computer-vision-history-so-far 2025-05-22T04:21:25Z monthly 0.8 ... https://voxel51.com/blog/how-to-build-a-semantic-search-engine-for-emojis 2025-05-22T04:21:25Z monthly 0.8 ... https://voxel51.com/blog/how-computer-vision-is-changing-sports 2025-05-22T04:21:25Z monthly 0.8 ... https://voxel51.com/blog/comparing-vqa-and-action-recognition 2025-05-22T04:21:25Z monthly 0.8 ... https://voxel51.com/blog/how-to-estimate-depth-from-a-single-image 2025-05-22T04:20:47Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-jan-25-2024 2025-05-22T04:21:19Z monthly 0.8 ... https://voxel51.com/blog/how-computer-vision-is-changing-retail 2025-08-12T01:11:43Z monthly 0.8 ... https://voxel51.com/blog/how-to-visualize-your-data-with-dimension-reduction-techniques 2025-05-22T04:21:19Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-16-2024 2025-05-22T04:21:19Z monthly 0.8 ... https://voxel51.com/blog/exploring-gradcam-and-more-with-fiftyone 2025-05-22T04:21:19Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-23-2024 2025-05-22T04:21:19Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-feb-15-2024 2025-05-22T04:21:19Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-march-1-2024 2025-05-22T04:21:19Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-march-8-2024 2025-05-22T04:21:11Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-0-23-5-and-fiftyone-teams-1-5-6 2025-05-22T04:21:19Z monthly 0.8 ... https://voxel51.com/blog/finding-outliers-in-your-vision-datasets 2025-05-22T04:21:11Z monthly 0.8 ... https://voxel51.com/blog/data-augmentation-is-still-data-curation 2025-08-12T01:15:51Z monthly 0.8 ... https://voxel51.com/blog/how-computer-vision-is-changing-security 2025-05-22T04:21:11Z monthly 0.8 ... https://voxel51.com/blog/a-history-of-clip-model-training-data-advances 2025-05-22T04:21:11Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-0-23-6-and-fiftyone-teams-1-5-7 2025-05-22T04:21:11Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-march-15-2024 2025-05-22T04:21:11Z monthly 0.8 ... https://voxel51.com/blog/streamline-computer-vision-workflows-with-hugging-face-transformers-and-fiftyone 2025-05-22T04:21:11Z monthly 0.8 ... https://voxel51.com/blog/how-to-easily-cluster-your-computer-vision-datasets 2025-05-22T04:21:11Z monthly 0.8 ... https://voxel51.com/blog/efficiently-managing-and-querying-visual-data-with-mongodb-atlas-vector-search-and-fiftyone 2025-05-22T04:21:11Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-march-22-2024 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-to-integrate-with-nvidia-omniverse-simulation-services 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-march-21-2024 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-march-29-2024 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-wins-2024-artificial-intelligence-excellence-award 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/finding-the-optimal-confidence-threshold 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-5-2024 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-0-23-7-and-fiftyone-teams-1-5-8 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/iso-27001-certified 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-march-29-2024 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/how-to-cluster-images 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-april-12-2024 2025-05-22T04:20:59Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-12-2024 2025-05-22T04:20:47Z monthly 0.8 ... https://voxel51.com/blog/how-computer-vision-is-transforming-robotics 2025-08-12T01:08:48Z monthly 0.8 ... https://voxel51.com/blog/cvpr-2024-survival-guide-five-vision-language-papers-you-dont-want-to-miss 2025-05-22T04:20:47Z monthly 0.8 ... https://voxel51.com/blog/how-to-detect-small-objects 2025-05-22T04:20:47Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-april-18-2024 2025-05-22T04:20:47Z monthly 0.8 ... https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets 2025-05-22T04:20:47Z monthly 0.8 ... https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-2-benchmarks 2025-05-22T04:20:40Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-teams-1-6 2025-05-22T04:20:40Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-april-26-2024 2025-05-22T04:20:47Z monthly 0.8 ... https://voxel51.com/blog/anomaly-detection-with-fiftyone-and-anomalib 2025-05-22T04:20:40Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-may-2-2024 2025-05-22T04:20:40Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-may-10-2024 2025-05-22T04:20:40Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-may-8-2024 2025-05-22T04:20:40Z monthly 0.8 ... https://voxel51.com/blog/secury360-strengthens-security-with-fiftyone-teams 2025-05-22T04:20:40Z monthly 0.8 ... https://voxel51.com/blog/announcing-series-b-led-by-bessemer-venture-partners 2025-05-30T00:34:40Z monthly 0.8 ... https://voxel51.com/blog/voxel51-and-bessemer-venture-partners-collaborate-to-make-visual-ai-a-reality 2025-05-22T04:20:40Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-may-24-2024 2025-05-22T04:20:40Z monthly 0.8 ... https://voxel51.com/blog/how-shap-e-changed-how-we-think-about-diffusion-models 2025-05-22T04:20:35Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-0-24-with-3d-meshes-and-custom-workspaces 2025-05-22T04:20:35Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-may-30-2024 2025-05-22T04:20:35Z monthly 0.8 ... https://voxel51.com/blog/voxel51-at-cvpr-2024 2025-05-22T04:20:35Z monthly 0.8 ... https://voxel51.com/blog/5-papers-on-my-cvpr-2024-must-see-list 2025-05-22T04:20:35Z monthly 0.8 ... https://voxel51.com/blog/how-i-built-an-in-cabin-perception-dataset 2025-05-22T04:20:35Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-june-21-2024 2025-05-22T04:20:35Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-june-27-2024 2025-05-22T04:20:35Z monthly 0.8 ... https://voxel51.com/blog/segment-anything-in-a-ct-scan-with-nvidia-vista-3d 2025-05-22T04:20:35Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-19-2024 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-july-3-2024 2025-05-22T04:20:35Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-26-2024 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-august-02-2024 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/announcing-the-data-centric-ai-competition-revolutionizing-object-detection-through-smart-data-curation 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-august-9-2024 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/what-is-visual-ai-going-beyond-computer-vision 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-august-8-2024 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/four-years-of-open-source-fiftyone 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-august-15-2024 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-august-16-2024 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/sam-2-is-now-available-in-fiftyone 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-0-25 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-august-23-2024 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-august-30-2024 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-august-29-2024 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-september-12-2024 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-september-13-2024 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-visual-ai-in-healthcare-meetup-september-19-2024 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-september-20-2024 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/segments-ai-plugin-for-fiftyone 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-september-26-2024 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/the-power-of-open-source-ai-how-fiftyone-drives-the-future-of-visual-ai 2025-05-22T04:20:26Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-1-0 2025-05-22T04:20:22Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-october-4-2024 2025-05-22T04:20:22Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-october-10-2024 2025-05-22T04:20:22Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-october-11-2024 2025-05-22T04:20:22Z monthly 0.8 ... https://voxel51.com/blog/cotracker3-a-point-tracker-using-real-videos 2025-05-22T04:20:22Z monthly 0.8 ... https://voxel51.com/blog/cotracker3-enhanced-point-tracking-with-less-data 2025-05-22T04:20:22Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-october-24-2024 2025-05-22T04:20:22Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-november-1-2024 2025-05-22T04:20:22Z monthly 0.8 ... https://voxel51.com/blog/data-quality-the-hidden-driver-of-ai-success 2025-05-22T04:20:22Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-november-14-2024 2025-08-04T16:43:20Z monthly 0.8 ... https://voxel51.com/blog/recapping-eccv-2024-redux-day-1 2025-05-22T04:20:22Z monthly 0.8 ... https://voxel51.com/blog/recapping-eccv-2024-redux-day-3 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/recapping-eccv-2024-redux-day-4 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/the-neurlps-2024-preshow-naturalbench-evaluating-vision-language-models-on-natural-adversarial-samples 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/the-neurlps-2024-preshow-a-textbook-remedy-for-domain-shifts-knowledge-priors-for-medical-image-analysis 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/the-neurlps-2024-preshow-a-label-is-worth-a-thousand-images-in-dataset-distillation 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/the-neurlps-2024-preshow-what-matters-when-building-vision-language-models 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/the-neurips-2024-preshow-zero-shot-learning-a-misnomer 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/the-neurips-2024-preshow-creating-spiqa-addressing-the-limitations-of-existing-datasets-for-scientific-vqa 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/journey-into-visual-ai-exploring-fiftyone-together-part-i-introduction 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-teams-2-2 2025-05-22T04:20:15Z monthly 0.8 ... https://voxel51.com/blog/the-neurips-2024-preshow-are-we-measuring-what-we-think-we-are-the-perils-of-contaminated-benchmark-datasets 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/the-neurips-2024-preshow-a-data-centric-look-at-curation-strategies-for-image-classification 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/the-neurips-2024-preshow-using-knowledge-graphs-to-diagnose-and-debias-visual-datasets 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/the-neurips-2024-preshow-data-quality-over-quantity-why-real-images-still-reign-supreme-for-vision-model-training 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/the-neurips-2024-preshow-more-than-meets-the-eye-how-transformations-reveal-the-hidden-biases-shaping-our-datasets 2025-05-22T04:20:18Z monthly 0.8 ... https://voxel51.com/blog/five-must-read-data-centric-ai-papers-from-neurips-2024 2025-05-22T04:20:15Z monthly 0.8 ... https://voxel51.com/blog/on-leaky-datasets-and-a-clever-horse 2025-05-22T04:20:15Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-computer-meetup-december-12-2024 2025-05-22T04:20:15Z monthly 0.8 ... https://voxel51.com/blog/what-ai-means-for-science-in-2025 2025-05-22T04:20:15Z monthly 0.8 ... https://voxel51.com/blog/why-2024-was-the-best-year-for-visual-ai-so-far 2025-05-22T04:20:15Z monthly 0.8 ... https://voxel51.com/blog/journey-into-visual-ai-exploring-fiftyone-together-part-ii-getting-started 2025-05-22T04:20:15Z monthly 0.8 ... https://voxel51.com/blog/journey-into-visual-ai-exploring-fiftyone-together-part-iii-preparing-a-computer-vision-challenge 2025-05-22T04:20:15Z monthly 0.8 ... https://voxel51.com/blog/bias-in-data-what-embeddings-reveal-about-real-vs-synthetic-data-distribution 2025-05-22T04:20:15Z monthly 0.8 ... https://voxel51.com/blog/how-to-make-the-best-self-driving-dataset 2025-05-22T04:20:15Z monthly 0.8 ... https://voxel51.com/blog/the-frontier-of-visual-ai-in-medical-imaging 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-24-2023 2025-05-22T04:23:04Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-0-20 2025-05-22T04:23:21Z monthly 0.8 ... https://voxel51.com/blog/announcing-fiftyone-teams-1-2 2025-05-22T04:23:04Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-computer-vision-meetup-sept-14-2023 2025-05-22T04:22:28Z monthly 0.8 ... https://voxel51.com/blog/caffe-computer-vision-glacial-mass-modeling 2025-05-22T04:22:28Z monthly 0.8 ... https://voxel51.com/blog/heatmaps-fiftyone-tips-and-tricks-october-6th-2023 2025-05-22T04:22:11Z monthly 0.8 ... https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-oct-5-2023 2025-05-22T04:22:11Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-community-update-october-2023 2025-05-22T04:22:11Z monthly 0.8 ... https://voxel51.com/blog/happy-5th-birthday-voxel51 2025-05-22T04:22:11Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-community-update-november-2023 2025-05-22T04:21:49Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-0-23-3-and-fiftyone-teams-1-5-4 2025-05-22T04:21:25Z monthly 0.8 ... https://voxel51.com/blog/fiftyone-computer-vision-community-update-february-2024 2025-05-22T04:21:19Z monthly 0.8 ... https://voxel51.com/blog/solving-the-ai-blindspot-using-data-to-drive-models-in-automotive 2025-05-22T04:20:31Z monthly 0.8 ... https://voxel51.com/blog/voxel51-filtered-views-newsletter-january-17-2025 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/how-to-tame-your-data-dragon 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/journey-into-visual-ai-exploring-fiftyone-together-part-iv-model-evaluation 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/elderly-action-recognition-no-one-should-age-alone-ais-promise-for-the-next-generation-of-elders 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/understanding-dataset-difficulty-with-class-wise-autoencoders 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/voxel51-and-ori-partner-to-accelerate-visual-ai-innovation-in-u-s-government 2025-05-20T15:42:44Z monthly 0.8 ... https://voxel51.com/blog/visual-understanding-with-aimv2 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/beyond-the-microscope-diving-into-bioscan-5m-a-new-dataset-for-insect-biodiversity-research 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/aimv2-outperforms-clip-on-synthetic-dataset-imagenet-d 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/webuot-1m-a-dataset-for-underwater-object-tracking 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/imagenet-d-new-synthetic-test-set-designed-to-rigorously-evaluate-the-robustness-of-neural-networks 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/can-vlms-hear-what-they-see 2025-05-22T04:20:04Z monthly 0.8 ... https://voxel51.com/blog/gaussian-splatting-from-research-to-reality 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/supercharge-your-visual-ai-workflow-fiftyone-new-plugin-for-janus-pro 2025-05-22T04:20:08Z monthly 0.8 ... https://voxel51.com/blog/memes-are-the-vlm-benchmark-we-deserve 2025-05-22T04:20:04Z monthly 0.8 ... https://voxel51.com/blog/this-visual-illusions-benchmark-makes-me-question-the-power-of-vlms 2025-05-22T04:20:04Z monthly 0.8 ... https://voxel51.com/blog/introducing-fiftyone-enterprise-visual-ai-workflows 2025-05-30T06:43:19Z monthly 0.8 ... https://voxel51.com/blog/new-data-and-model-workflows-from-voxel51-accelerate-visual-ai-development-for-enterprises 2025-05-20T15:42:44Z monthly 0.8 ... https://voxel51.com/blog/streamline-visual-data-discovery-with-fiftyone-data-lens 2025-05-22T04:20:04Z monthly 0.8 ... https://voxel51.com/blog/visualizing-model-certainty-in-the-unknown 2025-05-22T04:20:04Z monthly 0.8 ... https://voxel51.com/blog/build-better-visual-ai-datasets-with-the-fiftyone-data-quality-workflow 2025-05-22T04:20:04Z monthly 0.8 ... https://voxel51.com/blog/unified-model-insights-with-fiftyone-model-evaluation-workflows 2025-05-22T04:20:04Z monthly 0.8 ... https://voxel51.com/blog/computer-vision-for-earth-observation-from-manual-digitizing-to-ai-powered-analysis 2025-05-22T04:20:04Z monthly 0.8 ... https://voxel51.com/events/visual-ai-in-healthcare-june-26-2025 2025-08-03T20:00:32Z monthly 0.8 ... https://voxel51.com/events/best-of-cvpr-july-10-2025 2025-08-04T13:10:56Z monthly 0.8 ... https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-jan-22-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/visual-ai-for-geospatial-jan-29-2025 2025-08-05T09:24:45Z monthly 0.8 ... https://voxel51.com/events/visual-ai-for-geospatial-jan-31-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/visual-ai-hackathon-jan-31-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/berlin-ai-ml-computer-vision-meetup-feb-7-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/elderly-action-recognition-challenge-wacv-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/best-of-neurips-feb-6-2025 2025-08-05T09:18:34Z monthly 0.8 ... https://voxel51.com/events/visual-ai-hackathon-march-9-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/munich-ai-ml-computer-vision-meetup-feb-6-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/visual-ai-hackathon-feb-15-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/stuttgart-ai-ml-computer-vision-meetup-feb-5-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/boston-ai-ml-computer-vision-meetup-feb-28-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/best-of-neurips-feb-4-2025 2025-08-05T09:19:39Z monthly 0.8 ... https://voxel51.com/events/dusseldorf-ai-ml-computer-vision-meetup-feb-4-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/getting-started-with-fiftyone-workshop-feb-19-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/ai-machine-learning-computer-vision-meetup-feb-20-2025 2025-08-05T09:17:10Z monthly 0.8 ... https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-feb-26-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/visual-ai-hackathon-march-21-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/visual-ai-hackathon-march-22-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/visual-ai-in-agriculture-march-26 2025-08-04T15:29:31Z monthly 0.8 ... https://voxel51.com/events/visual-ai-hackathon-march-15-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/getting-started-with-fiftyone-workshop-march-12-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-workshop-march-19-2025 2025-08-04T16:17:49Z monthly 0.8 ... https://voxel51.com/events/visual-ai-hackathon-april-4-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/building-visual-ai-in-the-enterprise-workshop-march-27-2025 2025-08-06T23:36:44Z monthly 0.8 ... https://voxel51.com/events/world-agri-tech 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/foundations-of-computer-vision-workshop-march-4-2025 2025-08-04T14:27:52Z monthly 0.8 ... https://voxel51.com/events/neural-networks-fundamentals-multilayer-perceptrons-for-regression-workshop-march-11-2025 2025-08-04T14:28:54Z monthly 0.8 ... https://voxel51.com/events/nvidia-gtc-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/training-evaluation-of-classification-models-workshop-march-18-2025 2025-08-04T14:39:59Z monthly 0.8 ... https://voxel51.com/events/tech-ad-europe-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/raleigh-ai-machine-learning-and-computer-vision-meetup-april-17-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/vand-3-0-cvpr-2025-workshop 2025-05-26T10:41:40Z monthly 0.8 ... https://voxel51.com/events/ai-machine-learning-computer-vision-meetup-april-24-2025 2025-08-04T14:52:53Z monthly 0.8 ... https://voxel51.com/events/vand-3-0-challenge-at-cvpr-2025 2025-06-06T13:09:45Z monthly 0.8 ... https://voxel51.com/events/ai-machine-learning-computer-vision-meetup-march-20-2025 2025-08-04T16:15:51Z monthly 0.8 ... https://voxel51.com/events/convolutional-neural-networks-lenet5-workshop-march-25-2025 2025-08-04T14:41:03Z monthly 0.8 ... https://voxel51.com/events/training-techniques-for-convolutional-networks-workshop-april-1-2025 2025-08-04T14:42:24Z monthly 0.8 ... https://voxel51.com/events/multi-label-classification-with-binary-cross-entropy-amazon-satellite-images-workshop-april-8-2025 2025-08-04T14:43:40Z monthly 0.8 ... https://voxel51.com/events/interpretability-in-computer-vision-cam-grad-cam-workshop-april-15-2025 2025-08-04T14:43:35Z monthly 0.8 ... https://voxel51.com/events/convolutional-neural-networks-advanced-upsampling-u-net-for-semantic-segmentation-workshop-april-22-2025 2025-08-04T14:43:57Z monthly 0.8 ... https://voxel51.com/events/model-optimization-data-augmentation-regularization-workshop-april-29-2025 2025-08-04T14:44:22Z monthly 0.8 ... https://voxel51.com/events/image-embeddings-zero-shot-classification-with-clip-workshop-may-6-2025 2025-08-04T14:37:12Z monthly 0.8 ... https://voxel51.com/events/object-detection-instance-segmentation-yolo-in-practice-workshop-may-13-2025 2025-08-04T14:46:18Z monthly 0.8 ... https://voxel51.com/events/image-generation-diffusion-models-u-net-workshop-may-20-2025 2025-08-04T14:21:10Z monthly 0.8 ... https://voxel51.com/events/boston-ai-ml-and-computer-vision-meetup-april-18-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/deep-learning-fundamentals-with-pytorch-and-fiftyone-workshop-april-5-6-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/new-york-ai-ml-and-computer-vision-meetup-april-3-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/computer-vision-developer-hour-march-18 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/computer-vision-developer-hour-march-25 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/computer-vision-developer-hour-april-1 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/computer-vision-developer-hour-april-1-2 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/computer-vision-developer-hour-april-15 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/computer-vision-developer-hour-april-22 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/computer-vision-developer-hour-april-29 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/chicago-ai-ml-and-computer-vision-meetup-april-10-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/berlin-ai-ml-computer-vision-meetup-april-25-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-workshop-april-23-2025 2025-08-04T14:53:39Z monthly 0.8 ... https://voxel51.com/events/getting-started-with-fiftyone-workshop-april-16-2025 2025-08-06T23:34:55Z monthly 0.8 ... https://voxel51.com/events/building-visual-ai-in-the-enterprise-workshop-april-30-2025 2025-08-06T23:36:49Z monthly 0.8 ... https://voxel51.com/events/ai-ml-and-computer-vision-meetup-may-22-2025 2025-08-04T14:19:27Z monthly 0.8 ... https://voxel51.com/events/munich-ai-ml-and-computer-vision-meetup-april-24-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/sds-2025-workshop 2025-05-27T08:07:50Z monthly 0.8 ... https://voxel51.com/events/cologne-ai-ml-and-computer-vision-meetup-april-22-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/computer-vision-for-autonomous-driving-april-26-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/stuttgart-ai-ml-and-computer-vision-meetup-april-23-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/detecting-the-unexpected-practical-approaches-to-anomaly-detection-in-visual-data-workshop 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/best-of-wacv-may-29-2025 2025-08-04T14:13:23Z monthly 0.8 ... https://voxel51.com/events/best-of-wacv-may-30-2025 2025-08-04T14:12:16Z monthly 0.8 ... https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-workshop-may-21-2025 2025-08-04T14:20:07Z monthly 0.8 ... https://voxel51.com/events/getting-started-with-fiftyone-workshop-may-14-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/ai-ml-and-computer-vision-meetup-en-espanol-june-20-2025 2025-08-04T13:21:33Z monthly 0.8 ... https://voxel51.com/events/ann-arbor-ai-ml-and-computer-vision-meetup-may-14-2025 2025-05-16T21:05:55Z monthly 0.8 ... https://voxel51.com/events/visual-ai-in-healthcare-june-25-2025 2025-08-04T13:17:22Z monthly 0.8 ... https://voxel51.com/events/midl-2025-workshop 2025-06-05T09:37:59Z monthly 0.8 ... https://voxel51.com/events/mosaic-ai-fiftyone-scaling-physical-ai-for-mobility-and-autonomous-use-cases-june-17-2025 2025-08-04T14:10:09Z monthly 0.8 ... https://voxel51.com/events/computer-vision-developer-hour-may-13 2025-05-16T21:20:00Z monthly 0.8 ... https://voxel51.com/blog/author/mt\_admin 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/mt\_pierce 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/stgvoxel51 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/travis 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/kim 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/jason 2025-08-12T00:21:12Z monthly 0.8 ... https://voxel51.com/blog/author/eric 2025-07-15T07:07:11Z monthly 0.8 ... https://voxel51.com/blog/author/benjamin 2025-07-15T07:04:54Z monthly 0.8 ... https://voxel51.com/blog/author/dave 2025-07-15T07:06:26Z monthly 0.8 ... https://voxel51.com/blog/author/tarmily 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/vini 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/mt\_andrew 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/wpengine 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/yuheng-li 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/chien-vu 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/allen 2025-07-15T07:04:24Z monthly 0.8 ... https://voxel51.com/blog/author/leila 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/ritchie 2025-07-15T07:14:31Z monthly 0.8 ... https://voxel51.com/blog/author/jon 2025-07-15T07:09:57Z monthly 0.8 ... https://voxel51.com/blog/author/robert 2025-08-12T00:21:16Z monthly 0.8 ... https://voxel51.com/blog/author/prerna 2025-07-15T07:14:03Z monthly 0.8 ... https://voxel51.com/blog/author/markus 2025-07-15T07:12:39Z monthly 0.8 ... https://voxel51.com/blog/author/mt\_cathy 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/jimmy 2025-08-12T00:25:15Z monthly 0.8 ... https://voxel51.com/blog/author/jacob 2025-07-15T07:08:12Z monthly 0.8 ... https://voxel51.com/blog/author/dillon 2025-07-15T07:06:48Z monthly 0.8 ... https://voxel51.com/blog/author/dan 2025-08-12T00:23:17Z monthly 0.8 ... https://voxel51.com/blog/author/harpreet 2025-08-12T00:21:19Z monthly 0.8 ... https://voxel51.com/blog/author/paularamos 2025-08-12T00:21:22Z monthly 0.8 ... https://voxel51.com/blog/author/vallyn 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/brent 2025-08-12T00:35:27Z monthly 0.8 ... https://voxel51.com/blog/author/manushree 2025-08-12T00:41:05Z monthly 0.8 ... https://voxel51.com/blog/author/jacob-sela 2025-08-12T00:41:33Z monthly 0.8 ... https://voxel51.com/blog/author/mike 2025-07-15T07:13:10Z monthly 0.8 ... https://voxel51.com/blog/author/monica 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/sliday 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/amanda-seo-consultant 2025-07-15T07:04:19Z monthly 0.8 ... https://voxel51.com/blog/author/sofia-sliday 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/steve 2025-07-15T07:16:24Z monthly 0.8 ... https://voxel51.com/blog/author/oscarsidebo 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/oskarengstrom 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/author/kirtivoxel51-com 2025-07-15T07:10:46Z monthly 0.8 ... https://voxel51.com/blog/author/nickvoxel51-com 2025-08-12T00:25:19Z monthly 0.8 ... https://voxel51.com/blog/author/brianvoxel51-com 2025-08-12T00:27:25Z monthly 0.8 ... https://voxel51.com/blog/author/deanlee808gmail-com 2025-05-15T14:21:35Z monthly 0.8 ... https://voxel51.com/blog/category/uncategorized 2025-05-19T05:46:49Z monthly 0.8 ... https://voxel51.com/blog/category/tutorials 2025-05-19T05:46:49Z monthly 0.8 ... https://voxel51.com/blog/category/vector-search 2025-05-19T05:46:49Z monthly 0.8 ... https://voxel51.com/blog/category/industry-solutions 2025-05-19T05:46:49Z monthly 0.8 ... https://voxel51.com/blog/category/plugins 2025-05-19T05:46:49Z monthly 0.8 ... https://voxel51.com/blog/category/integrations 2025-05-19T05:46:49Z monthly 0.8 ... https://voxel51.com/blog/category/mlvoxel51 2025-05-19T05:46:49Z monthly 0.8 ... https://voxel51.com/blog/category/datasets 2025-05-19T05:46:49Z monthly 0.8 ... https://voxel51.com/blog/category/event-recaps 2025-05-19T05:46:49Z monthly 0.8 ... https://voxel51.com/blog/category/product-news 2025-05-20T23:39:17Z monthly 0.8 ... https://voxel51.com/blog/category/tips-tricks 2025-05-19T07:02:39Z monthly 0.8 ... https://voxel51.com/blog/category/computer-vision 2025-05-19T05:46:49Z monthly 0.8 ... https://voxel51.com/blog/tag/multi-modal-ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/transformers 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/year-in-review 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/carnegie-mellon 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/computer-vision-meetup 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/data-annotation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/qdrant 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/similarity-learning 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/wearable-vision-sensors 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/filtering 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dataset-curation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dataset-improvement 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/integrations 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/aggregations 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/pandas 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/pandas-style-queries 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/interactive-plots 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/autonomous-vehicles 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/end2end-learning 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/retail-use-case 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/supply-chain-use-case 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/synthetic-data 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/synthetic-data-generator 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/iou 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mistakenness 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/viewfield 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/exporting 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/community-rewards 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/anchor-boxes 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/embeddings 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/metadata 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-brain 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/persisting-datasets 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/video-datasets 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/voxel51-birthday 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/voxel51-milestone 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mongodb 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/pytorch 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/sorting 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/hacktoberfest 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/open-source 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-0-17 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/custom-datasets 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/tiff-images 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/funding 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/series-a 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/custom-plugins 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/importing 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/oss-community 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/success-story 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/labels 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/bdd100k 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/berkeley-deep-drive 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/clip 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/natural-language-processing 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/nlp 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/openai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/pinecone 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/classification 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/nearest-neighbors 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/data-centric-machine-learning 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/data-centric-ml 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mlops 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mlops-meetup 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/kinetics 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/video-classification-models 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/clearml 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/deep-learning 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/experiment-tracking 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/object-detection 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/computer-vision 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/activitynet 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/model-evaluation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/view-stages 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/viewstage 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/jupyter-notebook 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/evaluation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/adding-data 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/merging-data 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-app 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/tensorflow 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/annotation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cifar-10 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/lightning-flash 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mnist 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/coco 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/labelbox 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mobilenet 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/model-zoo 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/umap 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/avatar 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/frame-interpolation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/kitti-dataset 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/performance-capture 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/sensor-fusion 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/stereoscopic-vision 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/hugging-face 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/hyperparameter-scheduling 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/hyperparameters 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vision-transformers 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vit 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/anniversary 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/image-dataset 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/images 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/open-images 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/open-images-v6 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/colab-notebook 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/notebooks 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/company-culture 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/developer-relations 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/hiring 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/open-positions 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/open-roles 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/people-at-voxel51 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/agriculture 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/agriculture-use-case 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/industry-spotlight 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/use-case 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/computer-vision-news-recap 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cookiecutter 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/docker 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/github-actions 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/poetry 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/geojson 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/labeling-mistakes 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/instance-segmentation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ms-coco 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ultralytics 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/yolo 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/yolov8 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mean-average-precision 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/albumentations 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/diffusion-models 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fine-tune-models 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/gans 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/anomalib 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/asr 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/edge-ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/openai-whisper-model 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/openvino 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/speech-recognition 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/whisper 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/whisper-model 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-0-19 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/on-disk-segmentations 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/saved-views 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/spaces 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ui-filtering 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/community-update 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/action-recognition 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ucf101 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/heatmaps 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/segmentations 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/google 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/keypoints 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/open-images-v7 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/point-labels 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/anomaly-detection 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/bin-picking 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cognex 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/datalogic 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/defect-detection 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/depalletizing 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/industrial-automation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/instrumental-ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/machine-tending 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/machine-vision 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/manufacturing 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/matroid 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mech-mind-robotics 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/palletizing 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/pickit-3d 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/predictive-maintenance 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/preml-gmbh 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/prophesee 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/protex-ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/rios-intelligent-machines 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/stemmer-imaging 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/quickstart-dataset 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/quickstart-video-dataset 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/image-restoration 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/low-light-image-enhancement 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/models-in-production 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/recycling-max-pooling 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/rmp 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cityscapes 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/semantic-urban-scene-understanding 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/events 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/getting-started-with-fiftyone 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/training 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/workshop 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/workshop-series 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/3d-point-cloud 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/birds-eye-view 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dbscan 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/point-cloud 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/point-cloud-synthesis 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/point-e 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/quickstart-groups-dataset 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/natural-language-search 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/semantic-search 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/similarity-search 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vector-database 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/regex 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/3d 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-0-20 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/nlp-search 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vector-search-database 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vector-search-engines 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-teams-1-2 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/getting-started-with-fiftyone-workshop 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/computer-vision-events 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/video-data-analysis 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/gligen 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/grounded-language-to-image-generation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/text-to-image 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/text-to-image-diffusion-models 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/image-matching-challenge-2023 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/kaggle 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/kaggle-competition 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/datasets 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/filtered-views 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/machine-learning 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/newsletter 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cotracker3 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/point-tracking 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/gpt 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/human-motion-synthesis 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mocap 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/motion-capture 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/motion-estimation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/motion-prediction 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/t2m-gpt 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vq-vae 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/hyperparameter-sweep 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/model-selection 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/model-training 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/weights-biases 2025-05-22T04:26:45Z monthly 0.8 ... https://voxel51.com/blog/tag/attention 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/amazon-dataset 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/armbench 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dataset-visualization 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/object-segmentation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/pick-and-place-robots 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/robotics 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/robots 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/deci-ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/supergradient 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/yolo-nas 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/model-predictions 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/uniqueness 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/voc-annotation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/voc-label 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cvf 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cvpr 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ieee 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/midjourney 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/nerf 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dataset-views 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/object-patches 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/sidebar 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dreambooth 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/f2-nerf 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/imagebind 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mask-dino 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mobilenerf 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/sadtalker 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/soldier 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/videofusion 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/opendoor 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/real-estate 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/wildlife-ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/color-schemes 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dynamic-groups 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-0-21 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/operators 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/plugins 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/langchain 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/voxelgpt 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cvpr-2023 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/tradeshows 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/sama-coco 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/yolov5 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai-stack 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/data-centric-ai-tooling 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/try-fiftyone-ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/compute\_hardness 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/custom-color-schemes 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/webinar 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/arkittrack 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/geonet 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/imagenet-e 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/jrdb-pose 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/llcm 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mobile-hdr 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/mvimgnet 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/spring 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/synsl-120k 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/caption 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/controlnet 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/google-conceptual-captions 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/multimodal 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/lancedb 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/milvus 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/meetup 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dreamsim 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/facial-video-representation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/human-perception 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/marlin 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/nights-dataset 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vector-search 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/lpips 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/nights 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/perceptual-metrics 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/similarity 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/keypoint-skeletons 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fields 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/sample-fields 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/medical-imaging 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/neural-congealing 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/radiotherapy 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/data-curation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/competition 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/opencv 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-community 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai-art 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dalle2 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/genai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/generative-ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/replicate 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/stable-diffusion 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vqgan 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/artificial-intelligence 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/computer-aided-diagnosis 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/disease-detection 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/healthcare 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/indiustry-spotlight 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/medicine 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/segment-anything 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/surgical-guidance 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cli 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/blip 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/visual-question-answering 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vqa 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/groups 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/drones 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/javascript 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/material-ui 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/react 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/youtube 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/group-datasets 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai-machine-learning-data-science-meetup 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai-ml-ds-meetup 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/meetups 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/benchmark 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/bias 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/disparity 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fairness 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/meta 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/sa1b 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/deduplication 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/polylines 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/iccv 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/lidar 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/iccv23 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/badger 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/custom-badges 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/python-library 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/labeling 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/pose-estimation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/tag-1 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/asl 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/zilliz 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/tips-tricks 2025-05-22T04:27:40Z monthly 0.8 ... https://voxel51.com/blog/tag/video-object-tracking 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/neurips 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/partnership 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/v7 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai-referee 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fitness 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/industry-use-cases 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/injury-prevention 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/rehabilitation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/sports 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/sports-analytics 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/sports-fan-enhancement 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/autonomous-checkout 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/customer-behavior-analysis 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/inventory-management 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/product-recommendations 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/recommender-engines 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/retail 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/retail-industry 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/supply-chain 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/virtual-try-ons 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dimensionality-reduction 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/pca 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/resnet50 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/tsne 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/visualization 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dynamic-attributes 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/facial-recognition 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/intruder-detection 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/safety-security 2025-05-22T04:27:51Z monthly 0.8 ... https://voxel51.com/blog/tag/safety-equipment-recognition 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/search-and-rescue 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/security-checkpoint-screening 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/security-industry 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/nvidia 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/nvidia-omniverse 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai-excellence-award 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/awards 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/detector-model 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/rios 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/iso-27001 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/iso-27001-certification 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/adt-commercial 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dataset-management 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/everon 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/autonomous-navigation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/environmental-monitoring 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/object-recognition 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/quality-control 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/robotics-industry 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/surveillance 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/visual-inspection 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/text2image 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/llms 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/series-b 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/bessemer-venture-partners 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/bvp 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/3d-object-generation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/shap-e 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/3d-geometries 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/3d-meshes 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/custom-workspaces 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-0-24 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/llama2 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cvpr-2024 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/adas 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/automotive 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/driving-use-case 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/in-cabin-perception 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/zero-shot 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/medical-imagery 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/hugging-face-competition 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/visual-ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/what-is-visual-ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/rag-models 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/sam-2 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/custom-dashboards 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/elasticsearch 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-0-25 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/python-panels 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/data-annotation-workflows 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/segments-ai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-1-0 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/data-quality 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/high-quality-data 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/eccv 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/naturalbench 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/neurlps 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vlm 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vqa-benchmarks 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vqav2 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/image-analysi 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/knobo 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/medical 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/medical-ai-model 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/medical-education 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/neurlps-2024 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dataset-distillation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/soft-labels 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/spiqa 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/visualai 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/builtin-compute 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/data-lens 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/panels 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/query-performance 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/leaky-splits 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/leaky-splits-analysis 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai-in-chemistry 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai-in-natural-sciences 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai-in-physics 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai-in-science 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cifake 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/data-bias 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/embeddings-comparison 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/synthetic-vs-real-data 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/health 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/coreset-selection 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/zcore 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/zero-shot-coreset-selection 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/zero-shot-data-reduction 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/class-wise-autoencoders 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ml-research 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/reconstruction-error-ratios 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/rers 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/janus-pro 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/gaussian-splatting 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/gaussian-splats 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/3d-reconstruction 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/imagenet 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/imagenet-d 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/aimv2 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/zero-shot-classification 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/bioscan-5m 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/bioclip 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/barcodebert 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/geolocation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/webuot-1m 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/video-embeddings 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/hiera 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/text-embeddings 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/jina-embeddings-v3 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/visual-spectrogram-classification 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/vsc 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/esc-50 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/esc-50-dataset 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/music2latent 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/clap 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/spectrograms 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/moondream2 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/meme-understanding 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ocr 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/contextual-caption-generation 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/siglip-2 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-enterprise 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/model-certainty-visualization 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/dataset-zoo 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/families-in-the-wild 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiw 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-0-18 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/webinar-recap 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/case-study 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/fiftyone-teams 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/product-release 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/bounding-boxes 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/exporting-with-splits 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/faq 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/large-datasets 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/cvat 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/detections 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/grouped-datasets 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/label-studio 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/point-clouds 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/chatgpt 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/ai-generated-artwork 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/computer-vision-applications 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/computer-vision-research 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/computer-vision-trends 2025-05-19T05:50:26Z monthly 0.8 ... https://voxel51.com/blog/tag/data-centric-computer-vision 2025-05-19T05:50:26Z monthly 0.8 ... ... <|firecrawl-page-10-lllmstxt|> ## About Voxel51 [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # AboutVoxel51 We’re on a mission to bring transparency and clarity to the world’s data. We empower hundreds of thousands of AI builders to unlock data insights to maximize model performance and make visual AI a reality. [Apply for open roles](https://voxel51.com/careers) ![](https://cdn.sanity.io/images/h6toihm1/production/cd7fe79ef465aa75489f4c049ce2af346f0985b4-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/720f6702611411baf6a274b1010856c78c07235d-240x96.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/57564abbae97325ddeb3bc3b8b7b5690d61cddc5-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/be26ca6233b020b9201ddd482f7d34f53555b51d-276x90.png?auto=format&dpr=2&fit=max&q=75&w=100) Maintained 99% fall detection rates for model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/636eaa49ef62ea3d07ca264f7baa4fe0e494e125-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/3911d157da27a0e996468bf45393d4e01e522934-384x39.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/a54f5857074a8e45684334b38c3e4f1f1528274e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/e845660699edf4deed110e60425659ba8ae73e6e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Foundation for Florence-2 VLM development](https://voxel51.com/plugins/?search=florence) ![](https://cdn.sanity.io/images/h6toihm1/production/c8be8bb13c2be64cd70dda28334501a812eb0bc1-520x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/8c596efcf17de5a9bc02b05b44f55474802f0abe-90x90.png?auto=format&dpr=2&fit=max&q=75&w=90) ![](https://cdn.sanity.io/images/h6toihm1/production/1f4c7beab6c76544152d4f3d990059bc62eb90cb-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/aee0fde76666dd2d134875eb8d5247ee75dab896-153x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cb778b338fac262a1c8a9de0334b256747cfbc5b-500x261.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Eliminated repetitive manual transformations on 20 TB+ of visual data](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/2eed779bf08ffa8ea43cb822cb013674cc05e546-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/6f935624676c89371a1b094385b7b60439e09824-351x60.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/d66cf4d8219472fca085434fac624f80f23d5339-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/31cf802f8e638e4fe47146b1a97fb8e2a315d77d-309x54.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg)\\ \\ Official partner for visualizing Open Images Dataset V7](https://voxel51.com/blog/exploring-google-open-images-v7/) ![](https://cdn.sanity.io/images/h6toihm1/production/cd7fe79ef465aa75489f4c049ce2af346f0985b4-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/720f6702611411baf6a274b1010856c78c07235d-240x96.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/57564abbae97325ddeb3bc3b8b7b5690d61cddc5-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/be26ca6233b020b9201ddd482f7d34f53555b51d-276x90.png?auto=format&dpr=2&fit=max&q=75&w=100) Maintained 99% fall detection rates for model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/636eaa49ef62ea3d07ca264f7baa4fe0e494e125-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/3911d157da27a0e996468bf45393d4e01e522934-384x39.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/a54f5857074a8e45684334b38c3e4f1f1528274e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/e845660699edf4deed110e60425659ba8ae73e6e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Foundation for Florence-2 VLM development](https://voxel51.com/plugins/?search=florence) ![](https://cdn.sanity.io/images/h6toihm1/production/c8be8bb13c2be64cd70dda28334501a812eb0bc1-520x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/8c596efcf17de5a9bc02b05b44f55474802f0abe-90x90.png?auto=format&dpr=2&fit=max&q=75&w=90) ![](https://cdn.sanity.io/images/h6toihm1/production/1f4c7beab6c76544152d4f3d990059bc62eb90cb-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/aee0fde76666dd2d134875eb8d5247ee75dab896-153x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cb778b338fac262a1c8a9de0334b256747cfbc5b-500x261.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Eliminated repetitive manual transformations on 20 TB+ of visual data](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/2eed779bf08ffa8ea43cb822cb013674cc05e546-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/6f935624676c89371a1b094385b7b60439e09824-351x60.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/d66cf4d8219472fca085434fac624f80f23d5339-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/31cf802f8e638e4fe47146b1a97fb8e2a315d77d-309x54.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg)\\ \\ Official partner for visualizing Open Images Dataset V7](https://voxel51.com/blog/exploring-google-open-images-v7/) ![](https://cdn.sanity.io/images/h6toihm1/production/cd7fe79ef465aa75489f4c049ce2af346f0985b4-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/720f6702611411baf6a274b1010856c78c07235d-240x96.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/57564abbae97325ddeb3bc3b8b7b5690d61cddc5-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/be26ca6233b020b9201ddd482f7d34f53555b51d-276x90.png?auto=format&dpr=2&fit=max&q=75&w=100) Maintained 99% fall detection rates for model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/636eaa49ef62ea3d07ca264f7baa4fe0e494e125-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/3911d157da27a0e996468bf45393d4e01e522934-384x39.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/a54f5857074a8e45684334b38c3e4f1f1528274e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/e845660699edf4deed110e60425659ba8ae73e6e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Foundation for Florence-2 VLM development](https://voxel51.com/plugins/?search=florence) ![](https://cdn.sanity.io/images/h6toihm1/production/c8be8bb13c2be64cd70dda28334501a812eb0bc1-520x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/8c596efcf17de5a9bc02b05b44f55474802f0abe-90x90.png?auto=format&dpr=2&fit=max&q=75&w=90) ![](https://cdn.sanity.io/images/h6toihm1/production/1f4c7beab6c76544152d4f3d990059bc62eb90cb-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/aee0fde76666dd2d134875eb8d5247ee75dab896-153x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cb778b338fac262a1c8a9de0334b256747cfbc5b-500x261.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Eliminated repetitive manual transformations on 20 TB+ of visual data](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/2eed779bf08ffa8ea43cb822cb013674cc05e546-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/6f935624676c89371a1b094385b7b60439e09824-351x60.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/d66cf4d8219472fca085434fac624f80f23d5339-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/31cf802f8e638e4fe47146b1a97fb8e2a315d77d-309x54.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg)\\ \\ Official partner for visualizing Open Images Dataset V7](https://voxel51.com/blog/exploring-google-open-images-v7/) Since 2018 ## By the numbers Backed by Bessemer Venture Partners, Drive Capital, Tru Arrow Partners, Top Harvest Capital, Shasta Ventures, and ID Ventures. $0M total funding 0M FiftyOne OSS downloads 0+ enterprise customers ## We unlock breakthroughs in visual AI and transform industries from automotive to robotics ![](https://cdn.sanity.io/images/h6toihm1/production/22b434030668d87caf6a19635da5cab49dfe08d4-2048x1365.jpg?auto=format&dpr=2&fit=max&q=75&w=600) ### Our founding story Voxel51 was founded in 2018 at the University of Michigan when professor Jason Corso teamed up with his PhD Brian Moore to turn their research into developer-friendly tooling for computer-vision data. The explosive adoption of FiftyOne has made us the go-to platform for ML teams who work with multimodal data at scale. ![](https://cdn.sanity.io/images/h6toihm1/production/282a93d2b293247c11ae9f2b05bf0084ce5178d4-768x620.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,768,620&w=384) ![](https://cdn.sanity.io/images/h6toihm1/production/a7a1085305de96e5ec83621e54a23dfed5cb725c-2560x640.png?auto=format&dpr=2&fit=max&q=75&w=1280) Customer Testimonials ## Developers love FiftyOne > “From vehicle safety and autonomy to security systems to robotics, Bosch is a leader in artificial intelligence solutions utilizing computer vision. Voxel51’s solutions help us organize, evaluate and refine our data and models, enabling us to develop robust, reliable AI applications across multiple teams and projects. ” > > **Arvind Kumar Shekar** > > Lead Expert AI Validation, Bosch ![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) > “The biggest benefit of FiftyOne has been the speed of development. What used to take weeks or even months can now be done in days, with fewer people and a 7% increase in model performance. It’s freed up our team to focus on what they do best while accelerating our computer vision pipeline.”” > > **Kermal Eren** > > Lead Computer Vision Engineer, Ancera ![](https://cdn.sanity.io/images/h6toihm1/production/17832cc5c128034561711a5d43942252792462b4-479x101.webp?auto=format&dpr=2&fit=max&q=75&rect=0,4,479,95&w=100) > “As we dive into the development of Florence-5B, we’re relying on FiftyOne more than ever. The tool’s intuitive interface and rich feature set are essential for effectively managing our large datasets and gaining critical insights. ” > > **Bin Xiao** > > AI Researcher, Florence-2 Visual Language Model ![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/17832cc5c128034561711a5d43942252792462b4-479x101.webp?auto=format&dpr=2&fit=max&q=75&rect=0,4,479,95&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) About Voxel51 <|firecrawl-page-11-lllmstxt|> ## Voxel51 Blog Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/14713df0d4dec67cd3bb5e9c292c607820df061d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=540&q=75&w=960) Featured Databricks and Voxel51: Scaling Data-Centric Visual AI on the Data Intelligence Platform [Read article](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/7d027bddb314b23d6afedfbcbdd8e8770784661a-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured NVIDIA AI Podcast: ADAS, Visual AI, and the Road Ahead with Porsche and Voxel51 [Read article](https://voxel51.com/blog/nvidia-ai-podcast-adas-with-porsche) ![Your data, your advantage - the hidden cost of outsourced data annotation](https://cdn.sanity.io/images/h6toihm1/production/9733b62da57bf722c6ab7ec2dee89ec9353158a4-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured Your data, your advantage: the hidden cost of outsourced data annotation [Read article](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation) ![](https://cdn.sanity.io/images/h6toihm1/production/6ef6221c55258d5132bdb6deaa8c5494cbbd7cda-3840x2161.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.52&h=540&q=75&rect=11,30,3827,2116&w=960) Featured Zero-shot auto-labeling rivals human performance [Read article](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance) ![](https://cdn.sanity.io/images/h6toihm1/production/14713df0d4dec67cd3bb5e9c292c607820df061d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=540&q=75&w=960) Featured Databricks and Voxel51: Scaling Data-Centric Visual AI on the Data Intelligence Platform [Read article](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/7d027bddb314b23d6afedfbcbdd8e8770784661a-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured NVIDIA AI Podcast: ADAS, Visual AI, and the Road Ahead with Porsche and Voxel51 [Read article](https://voxel51.com/blog/nvidia-ai-podcast-adas-with-porsche) ![Your data, your advantage - the hidden cost of outsourced data annotation](https://cdn.sanity.io/images/h6toihm1/production/9733b62da57bf722c6ab7ec2dee89ec9353158a4-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured Your data, your advantage: the hidden cost of outsourced data annotation [Read article](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation) ![](https://cdn.sanity.io/images/h6toihm1/production/6ef6221c55258d5132bdb6deaa8c5494cbbd7cda-3840x2161.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.52&h=540&q=75&rect=11,30,3827,2116&w=960) Featured Zero-shot auto-labeling rivals human performance [Read article](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance) [Case Studies](https://voxel51.com/customers) [Events](https://voxel51.com/events) [Glossary](https://voxel51.com/glossary) [Documentation](https://docs.voxel51.com/) Product & News [View all](https://voxel51.com/blog/category/product-news) [![](https://cdn.sanity.io/images/h6toihm1/production/35e10bff3f7e49854806cbbd163022932bf07548-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How Voxel51 is Powering Physical AI with Databricks\\ \\ Product & News\\ \\ • \\ \\ Aug 20, 2025](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks) [![](https://cdn.sanity.io/images/h6toihm1/production/6aafb2b5fa699824c252fabfe2607eaeb820616a-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Enabling the AV Datasets of the Future with NVIDIA NuRec and FiftyOne\\ \\ Product & News\\ \\ • \\ \\ Aug 11, 2025](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/7d027bddb314b23d6afedfbcbdd8e8770784661a-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ NVIDIA AI Podcast: ADAS, Visual AI, and the Road Ahead with Porsche and Voxel51\\ \\ Product & News\\ \\ • \\ \\ Jul 30, 2025](https://voxel51.com/blog/nvidia-ai-podcast-adas-with-porsche) Machine Learning Research [View all](https://voxel51.com/blog/category/mlvoxel51) [![](https://cdn.sanity.io/images/h6toihm1/production/6ef6221c55258d5132bdb6deaa8c5494cbbd7cda-3840x2161.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.52&h=270&q=75&rect=11,30,3827,2116&w=480)\\ \\ Zero-shot auto-labeling rivals human performance\\ \\ ML@Voxel51\\ \\ • \\ \\ Jun 4, 2025](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance) [![](https://cdn.sanity.io/images/h6toihm1/production/62d6fe5dbbd89564bba1457900224b243fbd1a91-1200x675.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visualizing Model Certainty in the Unknown\\ \\ ML@Voxel51\\ \\ • \\ \\ Apr 1, 2025](https://voxel51.com/blog/visualizing-model-certainty-in-the-unknown) [![](https://cdn.sanity.io/images/h6toihm1/production/4f4b3ab874b2159c10b6ac360c5f72b3baf7ade4-2500x1406.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Understanding Dataset Difficulty with Class-Wise Autoencoders\\ \\ ML@Voxel51\\ \\ • \\ \\ Feb 4, 2025](https://voxel51.com/blog/understanding-dataset-difficulty-with-class-wise-autoencoders) All posts [![](https://cdn.sanity.io/images/h6toihm1/production/35e10bff3f7e49854806cbbd163022932bf07548-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How Voxel51 is Powering Physical AI with Databricks\\ \\ Product & News\\ \\ • \\ \\ Aug 20, 2025](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks) [![](https://cdn.sanity.io/images/h6toihm1/production/6aafb2b5fa699824c252fabfe2607eaeb820616a-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Enabling the AV Datasets of the Future with NVIDIA NuRec and FiftyOne\\ \\ Product & News\\ \\ • \\ \\ Aug 11, 2025](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/65dd49021e60c3c8d3f0982f888781871ba4ed26-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ From Prototype to Production: What it Really Takes to Deploy Computer Vision in Manufacturing\\ \\ Computer Vision\\ \\ • \\ \\ Aug 6, 2025](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing) [![](https://cdn.sanity.io/images/h6toihm1/production/7d027bddb314b23d6afedfbcbdd8e8770784661a-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ NVIDIA AI Podcast: ADAS, Visual AI, and the Road Ahead with Porsche and Voxel51\\ \\ Product & News\\ \\ • \\ \\ Jul 30, 2025](https://voxel51.com/blog/nvidia-ai-podcast-adas-with-porsche) [![](https://cdn.sanity.io/images/h6toihm1/production/c69415df557facd00411250d071e976a2bbff233-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ A Comprehensive Guide to Working with Point Cloud Data\\ \\ Learn\\ \\ • \\ \\ Jul 25, 2025](https://voxel51.com/blog/comprehensive-guide-point-cloud-data) [![](https://cdn.sanity.io/images/h6toihm1/production/6af33def6d297e2382d387e224e16451c95876af-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Why the “Annotate Everything” Era in Automotive AI Is Over\\ \\ Computer Vision\\ \\ • \\ \\ Jul 24, 2025](https://voxel51.com/blog/smarter-automotive-datasets-selection) [![](https://cdn.sanity.io/images/h6toihm1/production/14713df0d4dec67cd3bb5e9c292c607820df061d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Databricks and Voxel51: Scaling Data-Centric Visual AI on the Data Intelligence Platform\\ \\ Product & News\\ \\ • \\ \\ Jul 22, 2025](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/289034022fb1785d1f1960793ae875dd324e13d8-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Import Kaggle Datasets into FiftyOne & Publish to Hugging Face Hub — Step‑by‑Step Tutorial with ASL‑MNIST\\ \\ Tutorials\\ \\ • \\ \\ Jul 17, 2025](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub) [![](https://cdn.sanity.io/images/h6toihm1/production/7ab7bffdf1d0c321751e00fe91969c1fcf8f40d6-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ What Makes ‘Good’ Data? A View from the Front Lines of AI\\ \\ Datasets\\ \\ • \\ \\ Jul 17, 2025](https://voxel51.com/blog/what-makes-good-data-a-view-from-the-front-lines-of-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/55292ec1fada0552df3d72fb684759407b67f5bc-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Visual AI in Manufacturing: 2025 Landscape\\ \\ Industry Solutions\\ \\ • \\ \\ Jul 16, 2025](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape) [![](https://cdn.sanity.io/images/h6toihm1/production/ebd118c3d8181dcba0532200d024759e23c7ec1e-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Computer vision in healthcare: 12 breakthrough case studies\\ \\ Industry Solutions\\ \\ • \\ \\ Jul 15, 2025](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies) [![Your data, your advantage - the hidden cost of outsourced data annotation](https://cdn.sanity.io/images/h6toihm1/production/9733b62da57bf722c6ab7ec2dee89ec9353158a4-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Your data, your advantage: the hidden cost of outsourced data annotation \\ \\ Product & News\\ \\ • \\ \\ Jul 14, 2025](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation) Load more ## Enough data wrangling.
 Request a demo. [Get started](https://voxel51.com/link-catcher) [Explore the Demo](https://voxel51.com/link-catcher) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-12-lllmstxt|> ## Best of CVPR 2025 [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision) The Best of CVPR 2025 Series – Day 2 May 29, 2025 • 10 min read Article content In this article [Elevating Research Voices in Vision AI](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#ea993ae3e680) [Teaching AI What Matters at Home \[1\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#9205ddf9beff) [Teaching AI to Listen: Doctor-in-the-Loop Diagnosis with Concept Reasoning \[2\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#85e65e9193e0) [OFER: Filling in the Blanks in Occluded Faces \[3\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#63fb5ecffb7a) [Multi-Flow: Smarter Eyes on the Factory Floor \[4\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#3b091d813ec9) [Why These Papers Matter — A New Era in Computer Vision](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#d490b43df918) [What is next?](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#9e7261c72a86) [References](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#1748348e084b) In this article [Elevating Research Voices in Vision AI](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#ea993ae3e680) [Teaching AI What Matters at Home \[1\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#9205ddf9beff) [Teaching AI to Listen: Doctor-in-the-Loop Diagnosis with Concept Reasoning \[2\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#85e65e9193e0) [OFER: Filling in the Blanks in Occluded Faces \[3\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#63fb5ecffb7a) [Multi-Flow: Smarter Eyes on the Factory Floor \[4\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#3b091d813ec9) [Why These Papers Matter — A New Era in Computer Vision](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#d490b43df918) [What is next?](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#9e7261c72a86) [References](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2#1748348e084b) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Elevating Research Voices in Vision AI Computer vision is evolving fast, but too often, powerful research sits unread in conference proceedings. That’s why we created the Best of CVPR [virtual meetup](https://voxel51.com/events/best-of-cvpr-july-10-2025)series—to bring researchers to the forefront, spotlight how their work can solve real-world problems, and open doors for future collaboration. Research needs room to breathe, be explained, and celebrated. This blog is the second in a three-part series highlighting papers that go beyond technical novelty — they address safety, trust, fairness, and usability across industries. From smart homes and clinical decision support to expressive avatars and factory inspection, Day 2 shows how vision AI is growing more aware, robust, and practical. Check out [Day 1](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1) and [Day 3](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3) blog posts as well. ![](https://cdn.sanity.io/images/h6toihm1/production/60ea9f61452dac085454223ef44963e940f6b3f0-1400x515.webp?auto=format&dpr=2&fit=max&q=75&w=1400) Let’s dive into the impact of the algorithms. ## Teaching AI What Matters at Home \[1\] **_Paper title:_** _SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models._ **_Paper Authors:_** _Xinyi Zhao, Congjing Zhang, Pei Guo, Wei Li, Lin Chen, Chaoyue Zhao, Shuai Huang._ **_Institutions:_** _University of Washington, Wyze Labs_ Can we build AI that understands anomalies in smart homes, not just detecting strange events but explaining them in ways humans trust? ![](https://cdn.sanity.io/images/h6toihm1/production/7a1ec101d1229f0725f41c8b8b45e993213e3ef7-900x620.webp?auto=format&dpr=2&fit=max&q=75&w=900) ### What It’s About: This paper introduces SmartHome-Bench, the first benchmark specifically designed for video anomaly detection (VAD) in smart home environments. The benchmark includes a dataset of 1,203 smart home videos annotated with anomaly types, detailed descriptions, and reasoning, and it evaluates the performance of multi-modal large language models (MLLMs) using various prompting strategies. ### Why It Matters: Previous benchmarks for VAD were geared toward public settings and lacked relevance for private environments like homes. Smart home anomalies — such as pet escapes, elder falls, or child safety issues — are unique and sensitive. This work fills a gap by enabling the evaluation of VAD in personal, safety-critical domains, emphasizing trust, transparency, and reasoning. ### How It Works: - The authors propose a taxonomy of 7 smart home event categories (e.g., wildlife, baby monitoring, senior care). - They annotate videos with binary anomaly labels, rich textual descriptions, and reasoning explanations (e.g., why a behavior is abnormal). - The paper evaluates six MLLMs (e.g., Gemini-1.5, GPT-4o, Claude-3.5-sonnet, VILA-13b) under multiple prompting setups: zero-shot, chain-of-thought (CoT), few-shot CoT, and in-context learning (ICL). - They introduce a new method: Taxonomy-Driven Reflective LLM Chain (TRLC) — a multi-step process combining rule generation, initial prediction, and self-reflection. ### Key Result: The proposed TRLC framework achieves a notable 11.62% improvement in anomaly detection accuracy over zero-shot prompting. It delivers the highest accuracy of 79.05% with Claude-3.5-sonnet and improves model performance across ambiguous (vague abnormal) and standard scenarios. ### Broader Impact: SmartHome-Bench paves the way for trustworthy, explainable AI in home monitoring. It highlights the limitations of current MLLMs in capturing nuanced private-space anomalies and introduces strategies for improvement. The benchmark and TRLC methodology could influence safety tech, elder care systems, and smart home automation while offering tools to improve human-AI alignment in real-world video understanding. ## Teaching AI to Listen: Doctor-in-the-Loop Diagnosis with Concept Reasoning \[2\] **_Paper title:_**_[Interactive Medical Image Analysis with Concept-based Similarity Reasoning](https://arxiv.org/abs/2503.06873)._ **_Paper Authors:_** _Ta Duc Huy, Sen Kim Tran, Phan Nguyen, Nguyen Hoang Tran, Tran Bao Sam, Anton van den Hengel, Zhibin Liao, Johan W. Verjans, Minh-Son To, Vu Minh Hieu Phan._ **_Institutions:_** _Australian Institute for Machine Learning — University of Adelaide, Flinders University_ ![](https://cdn.sanity.io/images/h6toihm1/production/426f5224ed27603c30f40051963a9759a5cb4cee-1400x922.webp?auto=format&dpr=2&fit=max&q=75&w=1400) What if doctors could interact directly with a diagnostic AI model — not only to view its reasoning, but to correct it in real time and teach it what truly matters in medical images? ### What It’s About: This paper introduces CSR (Concept-based Similarity Reasoning), an interpretable and interactive model for medical image analysis. CSR addresses limitations in current concept-based and prototype-based methods by offering patch-level interpretability, localized concept grounding, and real-time doctor interaction to refine predictions without relying on post-hoc analysis ### Why It Matters: Existing AI models often act as black boxes, making them unreliable in safety-critical domains like healthcare. CSR improves trustworthiness by offering transparency, letting doctors understand the why behind model predictions and intervene during both training and test time. This enhances both accuracy and adoption of AI tools in clinical workflows. ### How It Works: CSR combines: - Concept-based reasoning: It computes similarity between input image patches and learned concept prototypes. - Prototype learning: Uses contrastive learning to ensure semantic consistency and compactness across concepts. - Doctor-in-the-loop interaction: Doctors can reject irrelevant concepts or guide focus to specific regions using bounding boxes (spatial and concept-level feedback). - Patch-level explanations: For each concept, CSR returns visual maps indicating where it found supporting evidence ### Key Result: CSR outperforms existing interpretable methods across three biomedical datasets (TBX11K, VinDr-CXR, ISIC), achieving up to 94.4% F1-score. It also demonstrated higher trustworthiness, with a Pointing Game hit rate of 79.5% when refined by doctors, surpassing previous models like ProtoPNet and CBM. ### Broader Impact: CSR offers a trust-enhancing framework for medical AI by combining interpretability with interactivity. It empowers clinicians to inspect, correct, and guide AI systems, addressing key concerns around explainability and clinical safety. The proposed method moves beyond performance alone, emphasizing collaborative intelligence between human and machine in medical diagnostics. ## OFER: Filling in the Blanks in Occluded Faces \[3\] **_Paper title:_**_[OFER: Occluded Face Expression Reconstruction](https://arxiv.org/abs/2410.21629)_ _._ **_Paper Authors:_** _Pratheba Selvaraju, Victoria F. Abrevaya, Timo Bolkart, Rick Akkerman, Tianyu Ding, Faezeh Amjadi, Ilya Zharkov._ **_Institutions:_** _University of Massachusetts Amherst, MPI-IS, Google Research, University of Amsterdam, Microsoft Research_ How do we reconstruct 3D facial expressions from a single image when part of the face is covered by a hand, mask, or hair, yet still produce multiple plausible, expressive results? ![](https://cdn.sanity.io/images/h6toihm1/production/3fa15e73bff40f9751f405ba284e9426683bc4c6-1400x734.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ### What It’s About: The paper introduces OFER, a novel method for reconstructing 3D faces with diverse expressions from single occluded images. OFER is the first to combine two conditional diffusion models for generating shape and expression coefficients of a face model (FLAME), and introduces a ranking mechanism to select the most accurate facial identity among multiple generated hypotheses. ### Why It Matters: Single-image 3D face reconstruction is already challenging; occlusions introduce ambiguity and variability. Traditional methods often fail under occlusion or generate unrealistic results. OFER offers a plausible, diverse, and structurally consistent alternative, which is crucial for applications in telepresence, AR/VR, medical imaging, and biometric systems. ### How It Works: - IdGen: A conditional diffusion model generates multiple candidate neutral 3D face shapes. - IdRank: A novel ranking network scores and selects the best-fitting shape sample based on the visible parts of the face. - ExpGen: A second diffusion model generates multiple expression variations for the selected shape. - The result is a set of expressive 3D face reconstructions that preserve identity while varying the expression plausibly. ### Key Result: OFER outperforms state-of-the-art methods (e.g., Diverse3D, EMOCA) in both quality and diversity of expression under occlusion. It achieves lower reconstruction errors on the NoW benchmark and improved results on the new CO-545 dataset introduced by the authors. ### Broader Impact: OFER sets a new standard for multi-hypothesis 3D face reconstruction under occlusion, unlocking real-world applications in challenging environments. The proposed ranking mechanism also opens new research directions for sample selection in diffusion models, and the CO-545 dataset provides a valuable benchmark for future research. ## Multi-Flow: Smarter Eyes on the Factory Floor \[4\] **_Paper title:_**_[Multi-Flow: Multi-View-Enriched Normalizing Flows for Industrial Anomaly Detection](https://arxiv.org/pdf/2504.03306)_ _._ **_Paper Authors:_** _Mathis Kruse, Bodo Rosenhahn._ **_Institutions:_** _Institute for Information Processing, L3S — Leibniz University Hannover_ Traditional visual inspection models can miss defects by relying on just one camera angle. What if AI could reason across multiple views to spot anomalies — no matter where they hide? ![](https://cdn.sanity.io/images/h6toihm1/production/58bf8d8b4f49e500293c970058dedd8611066b15-870x894.webp?auto=format&dpr=2&fit=max&q=75&w=870) ### What It’s About: The paper introduces Multi-Flow, a novel multi-view industrial anomaly detection architecture using normalizing flows. It enhances anomaly detection by integrating multiple views of an object and estimating the exact likelihood across them. The method is evaluated explicitly on the challenging Real-IAD dataset, which features five fixed views per object. ### Why It Matters: Most existing methods assume that a single image is enough to detect product anomalies, which is unrealistic in industrial scenarios where defects may only be visible from certain angles. Multi-Flow addresses this limitation by combining cross-view reasoning with a robust statistical modeling framework, offering better reliability in real-world settings where quality assurance is critical. ### How It Works: - Uses a RealNVP-based normalizing flow to model the distribution of normal (non-defective) features extracted from images. - Introduces a multi-view coupling block that allows cross-view message passing using 2D convolutions between adjacent and top-view camera images. - Employs foreground-background segmentation (MVANet) to exclude background noise and focus learning on the object itself. - Regularizes training via noise-conditioned data augmentation, using a modified variant of the SoftFlow and SimpleNet conditioning strategies. - Trained in a semi-supervised setup, using only defect-free training images, and evaluated both image-wise and sample-wise. ### Key Result: Achieves state-of-the-art detection on the Real-IAD dataset: - Sample-wise AUROC: 95.85 (↑ from 94.9 by SimpleNet) - Image-wise AUROC: 90.27 Outperforms all prior baselines in detecting defects aggregated across multiple views. Ablation studies show the importance of cross-view connections and background removal for boosting performance. ### Broader Impact: Multi-Flow contributes to building more trustworthy, robust AI systems for visual quality inspection in manufacturing. Its architecture applies to real-world production lines where multi-view camera setups are standard. The method reduces reliance on idealized single-view assumptions and paves the way for industry-scale, view-agnostic anomaly detection systems. The authors provide open-source code to facilitate further adoption and research. ## Why These Papers Matter — A New Era in Computer Vision The research we explored today gives us a glimpse into how vision AI is starting to show up in our daily lives — in homes, hospitals, and industrial settings. Each project brings something meaningful to the table: tools that help doctors understand AI decisions, systems that watch over smart homes with care, models that handle real-world imperfections like occlusions, and methods that make quality control more innovative and more reliable. What stands out is the intention behind these efforts. The focus isn’t just on making things work — it’s on making them work in clear, usable, and trustworthy ways. We’re excited to continue sharing the stories behind this work. We’ll see you soon on [Day 3](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3) or at the [virtual meetup](https://voxel51.com/events/best-of-cvpr-july-10-2025). ## What is next? If you’re interested in following along as I dive deeper into the world of AI and continue to grow professionally, feel free to connect or follow me on [LinkedIn](https://www.linkedin.com/in/paula-ramos-phd/). Let’s inspire each other to embrace change and reach new heights! You can find me at some [Voxel51 events](https://voxel51.com/events), or if you want to join this fantastic team, it’s worth taking a look at this page: [https://voxel51.app/careers](https://voxel51.com/careers) ![](https://cdn.sanity.io/images/h6toihm1/production/571af7476f44954ee67389529f23791f4d50ff97-990x990.png?auto=format&dpr=2&fit=max&q=75&w=990) ## References \[1\] Xinyi Zhao, Congjing Zhang, Pei Guo, Wei Li, Lin Chen, Chaoyue Zhao, and Shuai Huang, “SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models,” in _Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)_, 2025. \[2\] Ta Duc Huy, Sen Kim Tran, Phan Nguyen, Nguyen Hoang Tran, Tran Bao Sam, Anton van den Hengel, Zhibin Liao, Johan W. Verjans, Minh-Son To, and Vu Minh Hieu Phan, “Interactive Medical Image Analysis with Concept-based Similarity Reasoning,” in _Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)_, 2025\. Temporal link: [https://arxiv.org/abs/2503.06873](https://arxiv.org/abs/2503.06873) \[3\] Pratheba Selvaraju, Victoria Fernandez Abrevaya, Timo Bolkart, Rick Akkerman, Tianyu Ding, Faezeh Amjadi, and Ilya Zharkov, “OFER: Occluded Face Expression Reconstruction,” in _Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)_, 2025\. Temporal link: [https://arxiv.org/abs/2410.21629](https://arxiv.org/abs/2410.21629) \[4\] Mathis Kruse and Bodo Rosenhahn, “Multi-Flow: Multi-View-Enriched Normalizing Flows for Industrial Anomaly Detection,” in _Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)_, 2025.Temporal link: [https://arxiv.org/pdf/2504.03306](https://arxiv.org/pdf/2504.03306) [CVPR](https://voxel51.com/blog/tag/cvpr) [Visual AI](https://voxel51.com/blog/tag/visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/e926c07c7d1426c0fde8fdefa637c528d47b16f4-512x512.webp?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/db7358784a18a2ffa365704e7d941e73fbdf1fcd-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 1\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1) [![](https://cdn.sanity.io/images/h6toihm1/production/3a661345dfbc596f7118a7b8ec8375ea1f73d138-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 3\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3) [![](https://cdn.sanity.io/images/h6toihm1/production/b9e8bcb7b44ddb46a05144cd1e370b837e6aa401-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Visual Agents at CVPR 2025\\ \\ Computer Vision\\ \\ • \\ \\ May 28, 2025](https://voxel51.com/blog/visual-agents-at-cvpr-2025) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-13-lllmstxt|> ## Voxel51 Series B Funding [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Press](https://voxel51.com/blog/category/press) Voxel51 Raises $30M Series B Funding to Make Visual AI a Reality May 16, 2024 • 4 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _New financing led by Bessemer Venture Partners will expand go-to-market operations and accelerate product roadmap_ ANN ARBOR, Mich., May 16, 2024 — **Voxel51**, the leader in visual AI, today announced that it has closed a $30M Series B funding round. The round was led by Bessemer Venture Partners, with participation from new investor Tru Arrow Partners and existing investors Drive Capital, Top Harvest Capital, Shasta Ventures and ID Ventures. Voxel51 will use this funding to scale up its go-to-market organization, expand its community, invest in AI research and accelerate its roadmap to meet the growing market demand for solutions that unlock the value of [visual AI](https://voxel51.com/blog/what-is-visual-ai-going-beyond-computer-vision/). Voxel51 was founded by a team of machine learning and computer vision experts to help address the distressing failure rate of AI projects. While investments in AI have exploded, a [recent article in Harvard Business Review](https://hbr.org/2023/11/keep-your-ai-projects-on-track) reported that the failure rate of AI projects is estimated to be as high as 80%. With visual data, such as the image and video data that [already make up over 60% of all data traffic](https://www.prnewswire.com/news-releases/sandvines-2023-global-internet-phenomena-report-shows-24-jump-in-video-traffic-with-netflix-volume-overtaking-youtube-301723445.html), becoming an increasingly important part of AI projects, the complexity and difficulty of building successful AI applications only increases. All too often, the culprit behind these failures is the struggle to put visual data to work in building production-ready AI models and applications—today’s fragmented, inflexible tools and clunky manual processes lead to lengthy development cycles and painful errors in production. Voxel51 offers the antidote to those headaches, transforming the way that AI engineers and teams build visual AI. Voxel51 provides the leading refinery for data + models, where AI builders can simplify and automate their workflows to explore, visualize, curate and test datasets and put them to use to build models and applications. Tens of thousands of AI builders rely on Voxel51’s open source [FiftyOne](https://docs.voxel51.com/) and enterprise [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) offerings to build production-ready visual AI that is dramatically more accurate and robust, improving team productivity by up to 50% and model accuracy by up to 30%. Innovative organizations including LG Electronics, Berkshire Grey, Precision Planting, RIOS Intelligent Machines and Forsight use Voxel51 solutions to help them build innovative visual AI solutions that improve product quality, ensure safety and increase efficiency. “Traditional enterprises have increasingly sophisticated AI and ML teams with complex workflows that still include painful manual processes,” said Lindsey Li, investor at Bessemer Venture Partners. “As soon as Brian described his vision at Voxel51, we could see the platform’s role as core enterprise AI infrastructure of the future. We’ve been impressed by the entire team and their elegant and differentiated solution as a true orchestration layer for customers to build AI-forward and AI-native applications, faster and more easily. We are thrilled to partner with the team to build and scale their continued success.” “No one would build and tune a race car on a grass field and expect to win a Formula 1 race,” said Brian Moore, CEO and co-founder of Voxel51. “Yet, organizations building visual AI are all too often forced to build and fine-tune models with mislabeled, inadequate and under-representative data and then hope for success in production. We’re helping our customers and community build better AI applications by bringing their models and data together in one place, and we’re thrilled to have the support of great investors to accelerate our mission to make visual AI a reality.“ This new fundraising round is a milestone enabled by the momentum demonstrated since the close of its Series A funding round: - 4x increase in community membership and engagement - 6x growth to over 2 million downloads of the open source FiftyOne project - 10x increase in ARR for the commercial FiftyOne Teams offering That growth has been supported by continued innovation delivered by the Voxel51 team, including the following recent highlights: - [Integration of FiftyOne and NVIDIA Omniverse](https://www.prnewswire.com/news-releases/voxel51-accelerates-autonomous-vehicle-development-with-nvidia-omniverse-integration-302091979.html) to help autonomous vehicle developers create, curate, and visualize robust synthetic training data to maximize AI model performance - [Native vector search integrations](https://voxel51.com/vector-search/) with engines including Pinecone, MongoDB, Redis, Qdrant and more to enable fast, easy search through billions of images - [Introduction of VoxelGPT](https://www.prnewswire.com/news-releases/introducing-voxelgpt-ai-powered-computer-vision-insights-delivered-through-chat-301844496.html), a natural-language interface to FiftyOne that makes it possible for users to leverage large language models (LLMs) to rapidly and easily gain insights about visual data “FiftyOne Teams is critical to getting our AI models into the real world,” said Matt Shaffer, VP of Artificial Intelligence and Co-Founder at RIOS Intelligent Machines. “It’s the hub that powers our AI workflows, from data management to model refinement, enabling us to deliver AI-powered robotics into production lines rapidly and reliably to help our customers improve food packaging, reduce waste, and improve safety for any person on the factory floor.” To learn more: - Visit us online at [https://voxel51.com](https://voxel51.com/) to learn more about FiftyOne and FiftyOne Teams - Book a demo at [https://voxel51.com/book-a-demo/](https://voxel51.com/book-a-demo/) - Join the community at [https://voxel51.com/community-resources/](https://voxel51.com/community-resources/) - Follow us on [LinkedIn](https://www.linkedin.com/company/voxel51/), [X](https://twitter.com/voxel51), [Slack](https://slack.voxel51.com/) and [GitHub](https://github.com/voxel51/fiftyone) **About Voxel51** Voxel51 is helping organizations make visual AI a reality. Our open source and commercial software enables teams to build high-quality datasets and computer vision models that power leading machine learning and artificial intelligence applications. Tens of thousands of engineers and scientists have integrated open source FiftyOne into their workflows, and enterprise customers spanning from automotive, robotics, security and retail to healthcare rely on FiftyOne Teams to securely collaborate on datasets and models. We’re building a fully-remote team of exceptional and diverse people who want to bring visual AI to life. To learn more, visit [**voxel51.com**](https://voxel51.com/). ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0a21c6ca2af5253f72f6b88f9d8f6dafa5fb4ab3-2560x1390.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51 & Google Collaborate to Make Downloading & Visualizing Open Images A Breeze\\ \\ Press\\ \\ • \\ \\ May 13, 2021](https://voxel51.com/blog/fiftyone-open-images-collaboration) [![](https://cdn.sanity.io/images/h6toihm1/production/62b79d1d13bfc9926b5f560049adc33fd5d2cbf3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 Launches Computer Vision Industry’s First Open-Source Rapid Dataset Experimentation Tool\\ \\ Press\\ \\ • \\ \\ Aug 12, 2020](https://voxel51.com/blog/fiftyone-open-source-launch) [![](https://cdn.sanity.io/images/h6toihm1/production/b2822f56ac527e57cfa5a5e6dbdd2fea96ae2817-1934x1110.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51’s Coronavirus Physical Distancing Index Tracks Reaction to Social Distancing Around the World\\ \\ Press\\ \\ • \\ \\ Apr 1, 2020](https://voxel51.com/blog/voxel51-physical-distancing-index) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-14-lllmstxt|> ## Terms of Service [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # VOXEL51, INC TERMS OF SERVICE AGREEMENT **Effective Date: October 18, 2018** Voxel51, INC (“Voxel51”, “our”, “us”, or “we”) is Delaware Corporation that provides robust and customized computer vision and machine learning capabilities to enable advanced analytics that enhance societal welfare in a variety of business sectors. Voxel51 owns and operates the voxel51.com domain and associated subdomains to provide websites, associated online and/or mobile services, APIs, documentation, platforms, features, software, updates, and content (collectively “Services”), Voxel51 has adopted this Terms of Service Agreement (“Agreement”) to inform you (“User(s)”) of your rights and duties when using the Services. If you do not agree with the terms and conditions of this Agreement, you are expressly prohibited from accessing or using the Services and must discontinue your use immediately. PLEASE READ THIS AGREEMENT CAREFULLY BEFORE ACCESSING OR USING THE SERVICES. BY ACCESSING OR USING THE SERVICES, YOU AGREE TO BE BOUND BY THE TERMS AND CONDITIONS OF THIS AGREEMENT. VOXEL51 MAY, FROM TIME TO TIME, AND RESERVES THE RIGHT, IN ITS SOLE AND ABSOLUTE DISCRETION, TO MODIFY, LIMIT, CHANGE, DISCONTINUE, OR REPLACE THE SERVICES OR THIS AGREEMENT. IN THE EVENT VOXEL51 MODIFIES, LIMITS, CHANGES, OR REPLACES THE SERVICES OR THIS AGREEMENT, WE WILL BRING IT TO YOUR ATTENTION BY PLACING A NOTICE ON A VOXEL51.COM WEBSITE, BY SENDING YOU AN EMAIL, AND/OR BY SOME OTHER MEANS. IF YOU DON’T AGREE WITH THE UPDATED TERMS, YOU ARE FREE TO REJECT THEM, BUT YOU WILL THEN NO LONGER BE ABLE TO USE THE SERVICES. YOUR USE OF ANY OF THE SERVICES AFTER SAID MODIFICATION, LIMITATION, CHANGE, OR REPLACEMENT CONSTITUTES YOUR MANIFESTATION **Definitions** ● **“Data Protection Laws”** means all data protection and privacy laws, rules, and regulations applicable to a party in the performance of its obligations under this Agreement, including, where applicable the General Data Protection Regulation (“GDPR”). ● **“Derived Data”** means data and information generated from the processing and analysis of the User Data through the Services. ● **“Fees”** means the applicable fees as set forth on the Order Form. ● **“Intellectual Property Rights”** means any and all registered and unregistered rights granted, applied for, or otherwise now or hereafter in existence under or related to any patent, copyright, trademark, trade secret, database protection, or other intellectual property rights laws, and all similar or equivalent rights or forms of protection, in any part of the world. ● **“Fees”** means the applicable fees as set forth on the Order Form. ● **“Order Form”** means the invoice or other forms from Voxel51 for the initial order for the Service, and any subsequent invoice, or other forms from Voxel51, specifying, among other things, the amount of User interaction with the Services, the number of Users, the amount of User Data processed using the Services, and the size and length of time that User Data is maintained and stored by the Services. ● **“Third-Party Materials”** means materials and information, in any form or medium, including any open-source or other software, documents, data, content, specifications, products, equipment, or components of relating to the Services that are not proprietary to Voxel51. ● **“User Data”** means information, data, and other content, in any form or medium, that is collected, downloaded, or otherwise received, directly or indirectly from a user by or through the Services. For the avoidance of doubt, User Data includes Derived Data. ● **“User Personal Data”** means any User Data that is personal data, as defined by under the applicable Data Protection Laws ● **“Voxel51’s Materials”** means the Services and any and all other information, data, documents, materials, works, and other content, methods, processes, hardware, software, and other technologies and inventions, including any deliverables, technical descriptions, plans, or reports, that are provided or used by Voxel51 in connection with the Services or otherwise comprise or relate to the Services. For the avoidance of doubt, Voxel51’s Materials do not include Derived Data or User Data. **Warranties and Representations** You warrant and agree that you have the right and legal capacity to enter into this Agreement and to adhere to its terms and conditions. You warrant that you are a human individual that is eighteen (18) years of age or older. If you are under eighteen (18) years of age but at least thirteen (13) years of age, you must present this Agreement to your parent or legal guardian for their review. You warrant that you are not prohibited from assenting to this Agreement by any preexisting Agreement. You warrant and represent that any and all information that you provide to Voxel51 is accurate and valid. You agree to comply in good faith with the terms of this Agreement. You will not use the Services in any way that violates the rights of third parties, and you agree to comply with any and all applicable local, national, state, provincial, and international laws, treaties, and regulations. Given the global nature of the Internet, you agree to comply with all laws and rules where you reside or where you use the Services. The Services are operated in the United States and Voxel51 makes no representation that its Services or corresponding products are appropriate, lawful, or available for use in other locations. You may not use the Services if you are a resident of a country embargoed by the United States, or a foreign person or entity blocked or denied by the United States government. **User Data and Derived Data** You may submit User Data to Voxel51, including, but not limited to video files in any commonly-used format (e.g. mp4, avi, mov, etc.), images in any commonly-used format (e.g. png, jpg, tiff, etc.), text, annotations, and spreadsheets. User Data may be uploaded directly to Voxel51’s cloud platform from your personal or company computers or you may provide Voxel51 with permission to access to your User Data that is stored in your private internet-enabled network or a third-party cloud storage solution. Voxel51 shall not be responsible for any costs associated with the upload, download, or other method of User Data transfer to Voxel51. Upon obtaining access of your User Data, your User Data will be processed using the Services in order to provide you with the requested analytics. The processing of your User Data will result in the generation of Derived Data, which contain certain non-obvious information about the User Data, such as the location of objects, people, actions, activities, and other descriptions of the semantic content of the User Data. Voxel51 does not guarantee that the Derived Data will be error-free or completely accurate. As such, you acknowledge and agree that Voxel51 will be not be held liable for such errors and/or inaccuracies in the Derived Data. You have and retain sole ownership rights to all of your User Data and all of your Derived Data. As a result, you may upload, modify, and download your User Data and your Derived Data from Voxel51’s platform at your discretion. However, the storage and maintenance of User Data by Voxel51 will involve associated Fees, as further described in the **Payment and Account Termination** section below. You have and will retain sole responsibility for: (a) all of your User Data and Derived Data, including the respective content and use; (b) all information and materials that you provide to Voxel51 in connection with the Services; (c) your technology infrastructure, including computers, software, databases, electronic systems, and networks; (d) your access to and use of the Services and Voxel51’s Materials directly or indirectly by or through your system, including all results obtained from, and all decisions and actions based on such access or use. However, by submitting the User Data to Voxel51, you hereby grant Voxel51 a limited, irrevocable, worldwide, perpetual, non-exclusive, royalty free, and transferable license to use, access, display, perform and distribute your User Data and your Derived Data for the purpose of providing you the Services, improving our Services, conducting quality assurance, development, testing, and validation You acknowledge and agree that you shall at all times be the data controller and Voxel51 shall be the data processor with respect to the processing of User Personal Data in connection with your use of the Services. Solely if, and to the extent that Voxel51 is processing personal data, as defined the General Data Protection Regulation (“GDPR”), that is contained in the User Data, then the terms of the data processing agreement available at the Voxel51 website shall apply to such processing and be incorporated into this Agreement. As the data controller of User Personal Data, User represents and warrants to Voxel51 that its provision of personal data to Voxel51 and instructions for processing such personal data in connection with the Services shall comply with all Data Protection Laws. You also agree that the User Data you submit to Voxel51 will not contain third party copyrighted material, or material subject to other third party proprietary rights, unless you have permission from the rightful owner of the material or you are otherwise legally entitled to post the material and to grant Voxel51 all of the license rights granted herein. In accordance with applicable Data Protection Laws, Voxel51 shall take all commercially reasonable measures to protect the security and confidentiality of User Personal Data against any accidental or illicit alteration, destruction, or unauthorized access or disclosure to third parties. **Voxel51’s Rights in the Services and Intellectual Property Rights** Except for your User Data and your Derived Data, you acknowledge and agree that all right, title, and interest in and to Voxel51’s Materials, including all Intellectual Property Rights therein, are and will remain with Voxel51 and, with respect to Third-Party Materials, the applicable third-party providers own all right, title, and interest, including all Intellectual Property Rights, in and to the Third-Party Materials. Specifically, all Voxel51 marks are the property of Voxel51, including, but not limited to VOXEL51 and all Voxel51 logos. Except as explicitly provided herein, nothing in this Agreement shall be deemed to create a license, right, or authorization with respect to any of Voxel51’s Materials. All other rights in and to Voxel51’s Materials are expressly reserved by Voxel51. Absent prior written permission from Voxel51, you are not permitted to reproduce, prepare derivative works, distribute copies, perform, display, or use for commercial purposes Voxel51’s Materials or any domains/subdomains owned or controlled by Voxel51. **Use of Services & Account Registration** To use the Services, you must register a free account and create a user profile. Users can obtain accounts to use some of the Services by requesting an invitation from a voxel51.com website. Users will then be prompted to provide Voxel51 with the required information, which may include data, such as name, email address, company/personal website, phone number, recovery email address, password, and payment information. Please see our Privacy Policy, which is incorporated into this Agreement by reference, regarding the collection and use of this and other information about you. Voxel51 does not endorse you or discriminate based upon any information provided by you or made available for population on your account. We maintain different types of accounts for different types of users. You may never use another user’s account without permission. You have a duty to ensure that the information provided through your account is truthful, current, complete, and accurate. Users understand and agree that they have an ongoing duty to update and keep current the information provided through their account if and when that information changes. Users are expressly prohibited from creating an account that impersonates another person or entity, contains offensive or obscene language, or otherwise violates the rights of a third party. Users expressly agree that they will not use their account to interfere with or disrupt a third party’s enjoyment and use of the Services. Voxel51 reserves the right to restrict access to, monitor, suspend, disable, or delete accounts at any time, in its sole discretion, and without prior warning. Users agree to keep their account secure from unauthorized access. Users should not reveal their password to others. Users agree that they alone are responsible for their account. Users accept full responsibility for any and all use of their account, whether authorized or unauthorized. Users must notify Voxel51 immediately of any breach of security or unauthorized use of their account. Users agree to hold harmless and indemnify Voxel51 and to pay any associated Fees for any damages that arise out of or in relationship to the use of their account. Your use of any Third-Party Materials shall be governed solely by the terms and conditions applicable to the respective third parties providing the Third-Party Materials, as agreed to between you and the third parties. Voxel51 is not responsible for any and disclaims all liability with respect to Third-Party Materials, including without limitation, the privacy practices, data security processes, or other policies related to the Third-Party Materials. You agree to waive any claim against Voxel51 with respect to the use of any Third-Party Materials in connection with the Services. **Payment and Account Termination** Users will pay Voxel51 the Fees plus all applicable taxes, duties, levies, or charges imposed by any governmental entity anywhere in the world in connection with their use of the Services in accordance with the payment terms set forth in the Order Form. Users shall be responsible for all taxes related to the Services and this Agreement, exclusive of taxes on Voxel51’s income. Except as otherwise indicated in the applicable Order Form, all fees and expenses shall be in U.S. dollars. Unpaid and due Fees may be subject to a finance charge of one percent (1.0%) per month, or the maximum permitted by law, whichever is lower, plus all expenses associated with collection, including reasonable attorneys’ fees. Payment made to Voxel51 are processed internally via credit card, debit card, and bank account transfers. Payment is due to Voxel51 within thirty (30) days of receipt of invoice. If the method of payment is by credit card, User agrees to (i) keep User’s credit card information updated and (ii) authorize charging User’s credit card the Fees when due. Voxel51 may change any terms and conditions regarding payment for the Services at any time. In the event Voxel51 modifies, limits, changes, or replaces the services or this agreement, we will bring it to your attention by placing a notice on a voxel51.com website, by sending you an email, and/or by some other means. Your continued use of the Services shall constitute acceptance of any changes made to this Agreement. Users may terminate their account by using the Delete Account functionality on the voxel51.com console website or by emailing support@voxel51.com. If you cancel your account, Voxel51 is under no obligation to preserve your User Data or Derived Data for any length of time and will not be responsible for any loss of User Data or Derived Data. Voxel51 is under no obligation to provide you with the data associated with your account and/or user profile after cancellation of your account, except as otherwise provided in the Privacy Policy. Voxel51 recommends that you maintain your own backup of account and user profile data. **Use Restrictions** You expressly agree that you will not use the Services to violate any law, statute, ordinance, regulation, or treaty, to violate the rights of third parties, or for a use outside of the customary and intended purposes of the Services. Voxel51 reserves the right to remove any user at any time without notice. Users shall not access or use the Services or Voxel51’s Materials beyond the scope of the authorization granted under this Section. Specifically, users shall not, except as this Agreement expressly permits: ● copy, modify, or create derivative works or improvements of the Services, Third-Party Materials, or Voxel51’s Materials; ● rent, lease, sell, sublicense, assign, distribute, publish, transfer, or otherwise make available any Services or Voxel51’s Materials to any person on or in connection with the internet or any software as a service, cloud, or other technology or service; ● reverse engineer, disassemble, decompile, decode, adapt, or otherwise attempt to derive or gain access to the source code of the Services, Third-Party Materials, or Voxel51’s Materials, in whole or in part; ● bypass or breach any security device or protection used by the Services or Voxel51’s Materials; ● input, upload, transmit, or otherwise provide to or through the Services or Voxel51’s Materials, any information or materials that are harmful or injurious, or contain, transmit, or activate any software, hardware, or other technology, or means, including any virus, worm, malware, or other malicious computer code, the purpose or effect of which is to: (a) permit unauthorized access to, or to destroy, disrupt, disable, distort, or otherwise harm or impede in any manner any computer, software, firmware, hardware, system, or network; or (b) prevent a user from accessing or using the Services or Voxel51’s Materials as intended by this Agreement. ● damage, destroy, disrupt, disable, impair, interfere with, or otherwise impede or harm in any manner the Services or Voxel51’s provision of Services to any third party, in whole or in part; ● remove, delete, alter, or obscure any trademarks, warranties, or disclaimers, or any copyright, trademark, patent, or other intellectual property or proprietary rights notices from any Services or Voxel51’s Materials, including any copy thereof; ● access or use the Services or Voxel51’s Materials in any manner or for any purpose that infringes, misappropriates, or otherwise violates any Intellectual Property Rights or other rights of any third party, or that violates any applicable statute, law, ordinance, regulation, rule, code, order, constitution, treaty, common law, judgment, or other requirement of any federal, state, local, or foreign government, or any court; or ● access or use the Services or Voxel51’s Materials for purposes of competitive analysis of the Services or Voxel51’s Materials, the development, provision, or use of a competing software service or product or any other purpose that is to Voxel51’s detriment or commercial disadvantage. If you encounter content or witness behavior that you believe is inappropriate and violates this Agreement, you may report it to Voxel51 by sending an email to support@voxel51.com. **Section 230 of Communications Decency Act** You acknowledge and agree that Voxel51 is an interactive computer service provider under Section 230 of the Communications Decency Act. Though Voxel51 may edit, remove, or control the content displayed through the Services, you agree that Voxel51 will not be considered an information content provider and will not be held liable for the republication of defamatory or tortious content created by third parties, whether through the Services or otherwise. **Third Party Links** You understand that the Services may contain links to third party websites, applications, or services that Voxel51 does not own or control. You agree that Voxel51 will not be held responsible or liable for the content of third party websites, applications, or services and that Voxel51’s inclusion of those websites, applications, or services within its Services does not constitute Voxel51’s endorsement of, recommendation of, or affiliation with any of those websites, applications, or services. **Term and Termination** This Agreement will remain in full force and effect so long as the Services are in operation. Voxel51 may terminate this Agreement without liability at any time, without notice, and for any reason, including but not limited to for your violation of a term or condition of this Agreement. **Disclaimer of Warranties** ALL SERVICES AND VOXEL51’S MATERIALS ARE PROVIDED “AS IS.” VOXEL51 DISCLAIMS ALL IMPLIED WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE, AND NON-INFRINGEMENT, AND ALL WARRANTIES ARISING FROM COURSE OF DEALING, USAGE, OR TRADE PRACTICE. WITHOUT LIMITING THE FOREGOING, VOXEL51 MAKES NO WARRANTY OF ANY KIND THAT THE SERVICES OR VOXEL51’S MATERIALS OPERATE WITHOUT INTERRUPTION, ACHIEVE ANY INTENDED RESULT, BE COMPATIBLE OR WITH ANY SOFTWARE, SYSTEM, OR OTHER SERVICES, OR BE SECURE, ACCURATE, COMPLETE, FREE OF HARMFUL CODE, OR ERROR FREE. ALL THIRD-PARTY MATERIALS ARE PROVIDED “AS IS” AND ANY REPRESENTATIONS OR WARRANTY OF CONCERNING ANY THIRD-PARTY MATERIALS IS STRICTLY BETWEEN USERS AND THIRD-PARTY OWNERS OF THE THIRD-PARTY MATERIALS. **Limitation of Liability** IN NO EVENT WILL VOXEL51 BE LIABLE UNDER OR IN CONNECTION WITH THIS AGREEMENT OR ITS SUBJECT MATTER UNDER ANY LEGAL OR EQUITABLE THEORY, INCLUDING BREACH OF CONTRACT, TORT, STRICT LIABILITY, AND OTHERWISE, FOR ANY: (a) LOSS OF PRODUCT, USE, BUSINESS, REVENUE, OR PROFIT; (b) IMPAIRMENT, INABILITY TO USE OR LOSS, INTERRUPTION OR DELAY OF THE SERVICES; (c) LOSS, DAMAGE, CORRUPTION OR RECOVERY OF DATA, OR BEACH OF DATA OR SYSTEM SECURITY; (d) COST OF REPLACEMENT GOODS OR SERVICES; (e) LOSS OF GOODWILL OR REPUTATION; OR (f) CONSEQUENTIAL, INCIDENTAL, INDIRECT, EXEMPLARY, SPECIAL, ENHANCED, OR PUNITIVE DAMAGES, REGARDLESS OF WHETHER SUCH PERSONS WHERE ADVISED OF THE POSSIBILITY OF SUCH LOSSES OR DAMAGES OR SUCH LOSSES OR DAMAGES WERE OTHERWISE FORESEEABLE, AND NOTWITHSTANDING THE FAILURE OF ANY AGREED OR OTHER REMEDY OF ITS ESSENTIAL PURPOSE. IN NO EVENT WILL THE AGGREGATE LIABILITY OF VOXEL51 ARISING OUT OF OR RELATED TO THIS AGREEMENT, WHETHER ARISING UNDER OR RELATED TO BREACH OF CONTRACT, TORT, STRICT LIABILITY, OR ANY OTHER LEGAL OR EQUITABLE THEORY, EXCEED THE TOTAL AMOUNTS PAID TO VOXEL51 UNDER THIS AGREEMENT IN THE SIX MONTHS PRECEDING THE EVENT GIVING RISE TO THE CLAIM. SOME JURISDICTIONS DO NOT ALLOW THE EXCLUSION OR LIMITATION OF DAMAGES. IF YOUR JURISDICTION DOES NOT ALLOW THE EXCLUSION OR LIMITATION OF DAMAGES, YOU SHOULD SEEK LEGAL COUNSEL TO UNDERSTAND YOUR LEGAL RIGHTS UNDER THE LAW. **Indemnification** You agree to hold harmless, indemnify, and defend Voxel51, its officers, employees, agents, successors, and assigns, from and against any and all claims, demands, losses, damages, rights, and actions of any kind, including, but not limited to, property damage, infringement, personal injury, and death, that either directly or indirectly arise out of or are related to your use of the Services or Voxel51’s Materials, your use or provision of any services or monetary contributions made through the Services, your violation of any term or condition of this Agreement, your violation of any applicable law, statute, ordinance, regulation, or treaty, whether local, state, national, or international, or your violation of the rights of a third party. Your obligation to defend Voxel51 under the terms of this Agreement will not provide you with the right to control Voxel51’s defense, and Voxel51 reserves the right to control its defense and choose its counsel regardless of your contractual requirement to indemnify Voxel51. **No Assignment** You acknowledge and agree that you are prohibited from assigning, delegating, or transferring your rights and obligations under this Agreement, without Voxel51’s prior written consent. Voxel51 may assign its rights and obligations under this Agreement at any time and without consent. **Jurisdiction, Governing Law, and Resolution of Disputes** This Agreement will be interpreted, governed, construed, and enforce in accordance with the laws of the United States of America and the State of Michigan without giving effect to any conflicts of laws principles. The parties submit to and agree to personal jurisdiction in Michigan, with venue proper in Ann Arbor, Michigan. YOU AND VOXEL51 AGREE THAT ARBITRATION WILL BE THE EXCLUSIVE FORUM AND REMEDY AT LAW FOR ANY DISPUTES ARISING OUT OF OR RELATING TO THIS AGREEMENT, YOUR USE OF THE SERVICES, OR THE PURCHASE OF SERVICES FROM VOXEL51, INCLUDING ANY DISPUTES CONCERNING THE VALIDITY, INTERPRETATION, VIOLATION, BREACH, OR TERMINATION OF THIS AGREEMENT. ARBITRATION UNDER THIS AGREEMENT WILL BE HELD IN ANN ARBOR, MICHIGAN AND IN ACCORDANCE WITH THE MOST RECENTLY EFFECTIVE COMMERCIAL ARBITRATION RULES OF THE AMERICAN ARBITRATION ASSOCIATION. THE ARBITRATION PROCEEDING WILL BE DECIDED BY A SINGLE ARBITRATOR AND THE ARBITRATOR WILL DECIDE THE ARBITRATION PROCEEDING BY APPLYING THE LAWS AND LEGAL PRINCIPLES OF THE STATE OF MICHIGAN AND THE FEDERAL LAWS OF THE UNITED STATES. THE LOSING PARTY WILL BE REQUIRED TO PAY THE PREVAILING PARTY’S REASONABLE ATTORNEYS’ FEES. YOU AND VOXEL51 AGREE THAT THE SITUS OF THIS AGREEMENT IS IN THE STATE OF MICHIGAN. YOU AND VOXEL51 AGREE TO SUBMIT TO THE EXCLUSIVE PERSONAL JURISDICTION OF ANY SUCH ARBITRATOR OR ARBITRATION PROCEEDING. **Severability** If any provision of this Agreement is found to be invalid or unenforceable for any reason whatsoever, the remaining provisions will remain valid and unimpaired and will continue in full force and effect. **Integration** Voxel51 hereby incorporates its Privacy Policy into this Agreement. This Agreement and its incorporated Privacy Policy constitutes the entire agreement between the parties with respect to the use of the Services and Voxel51’s Materials. You acknowledge and agree that any additional provisions that may appear in any communication from you will not bind Voxel51. **No Waiver** You understand and agree that no term or provision of this Agreement will be deemed to have been waived and no breach will be deemed to have been consented to unless said waiver or consent is in writing and signed by the party to be charged. **Child Online Privacy Protection Act** The Services are not directed to persons under the age of eighteen (18) and Voxel51 will not knowingly collect personally identifiable information from children under the age of eighteen (18). If Voxel51 inadvertently collects such personally identifiable information, Voxel51 will delete the personally identifiable information in accordance with its security protocols. **Limitation on Actions** VOXEL51 AND YOU BOTH AGREE THAT ANY CAUSE OF ACTION ARISING OUT OF OR RELATED TO THE SERVICES MUST COMMENCE WITHIN ONE YEAR AFTER THE CAUSE OF ACTION ACCRUES. FAILURE TO ASSERT SAID CAUSE OF ACTION WITHIN ONE YEAR WILL PERMANENTLY BAR ANY AND ALL RELIEF. YOU WILL ONLY BE PERMITTED TO PURSUE CLAIMS AGAINST VOXEL51 ON AN INDIVIDUAL BASIS, NOT AS A PLAINTIFF OR CLASS MEMBER IN ANY CLASS OR REPRESENTATIVE ACTION OR PROCEEDING AND YOU WILL ONLY BE PERMITTED TO SEEK RELIEF (INCLUDING MONETARY, INJUNCTIVE, AND DECLARATORY RELIEF) ON AN INDIVIDUAL BASIS. **Reservation of Rights** All rights not expressly granted herein are reserved to Voxel51. **Notice** Any notice required by this Agreement must be in writing, and must be either mailed or emailed to: Voxel51, Inc. 330 E. Liberty St. Ann Arbor, MI 48104 [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-15-lllmstxt|> ## Voxel51 Privacy Policy [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # VOXEL51, INC PRIVACY POLICY **Effective Date: October 18, 2018** Voxel51, INC (“Voxel51”, “our”, “us”, or “we”) is a Delaware Corporation that owns and operates the voxel51.com domain and associated subdomains to provide websites, associated online and/or mobile services, APIs, documentation, platforms, features, software, updates, and content (collectively “Services”). When we talk about “you” or “User” in this Privacy Policy, we mean any individual or organization using and accessing our Services. Voxel51 has adopted this Privacy Policy to describe the type of data we collect, how we collect it, and how we use it. You should contact Voxel51 directly with any questions or concerns. By using our Services, you consent to our collection of your information and data. If we change our privacy policies and practices, we will update this Privacy Policy. VOXEL51 MAY CHANGE, MODIFY, AMEND, SUSPEND, TERMINATE, OR REPLACE THIS PRIVACY POLICY FROM TIME TO TIME AND WITHIN ITS SOLE AND ABSOLUTE DISCRETION. IN THE EVENT VOXEL51 MODIFIES, LIMITS, CHANGES, OR REPLACES THE SERVICES OR THIS AGREEMENT, WE WILL BRING IT TO YOUR ATTENTION BY PLACING A NOTICE ON A VOXEL51.COM WEBSITE, BY SENDING YOU AN EMAIL, AND/OR BY SOME OTHER MEANS. IF YOU DON’T AGREE WITH THE UPDATED TERMS, YOU ARE FREE TO REJECT THEM, BUT YOU WILL THEN NO LONGER BE ABLE TO USE THE SERVICES. IN THE EVENT VOXEL51 CHANGES, MODIFIES, AMENDS, OR REPLACES THIS PRIVACY POLICY, THE EFFECTIVE DATE, LOCATED ABOVE, WILL CHANGE. YOUR CONTINUED USE OF THE SERVICES AFTER A CHANGE IN THE EFFECTIVE DATE OF THIS PRIVACY POLICY CONSTITUTES YOUR MANIFESTATION OF ASSENT TO THE CHANGE, MODIFICATION, AMENDMENT, OR REPLACEMENT CONTAINED WITHIN. **Definitions** ● “General Data Protection Regulation (“GDPR”)” means the European Union (“EU”) law on data protection and privacy applicable to individuals within the EU. ● “Personal Data” means any information relating to a natural person who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier, or to one or more factors specific to the physical, physiological, genetic, mental, economic, cultural, or social identify of that natural person. **Types of Personal Data We Collect** The Personal Data that Voxel51 collects about you falls into two general categories: (i) Personal Data that is provided to us and (ii) Personal Data we collect automatically. **Personal Data that is provided to us:** You may provide Personal Data to us through the Services. This may be done, for example, when you use the Services or when you communicate with us in any way. When setting up an account for our Services, you will be asked to provide certain information about you such as: • First and last name; • Email address; • Mailing address; • Phone number; • Organization name; • Personal/organization website; • Billing/payment information; and • Any other information that you upload or submit to Voxel51 through the Services. **You may also provide Voxel51 with the following Personal Data:** • Videos; • Images; and • Any other additional data that you voluntarily submit to Voxel51 through the Services. Credit cards, debit cards, or other means may be used to pay for our Services. **Personal Data we collect automatically:** When you use the Services, we automatically collect certain Personal Data by using technologies, such as cookies. We do this to help us provide the Services and to ensure that we are providing Users with the best experiences with our Services. The Personal Data we automatically collect through the Services include the following: • Geolocation; • IP address; • Browser and search engine information; • Device information; • Usage of the Services, including, without limitation, any links or items clicked or pages viewed and statistics; • Information stored in cookies, pixel tags, or web beacons. The third-party service providers affiliated with Voxel51 have their own independent privacy policies governing the use of this information and we encourage you to read those privacy policies carefully. Voxel51 has no responsibility or liability for the activities of the third-party service providers used by Voxel51 to provide you with the Services. **Voxel51 uses your Personal Data to:** • provide you with the Services; • improve our Services; • facilitate your use of the Services and upgrades/replacements to the Services; • compile aggregate usage statistics about our Services; • conduct quality assurance, troubleshooting, testing, and validation; • comply with applicable legal requirements, agreements, and policies; • to perform other activities consistent with this Privacy Policy. When you interact with our Services, each request that you make (e.g. upload data, run an analytic, download output, etc.) is logged by our servers. **Voxel51 stores your Personal Data in the following manner:** Your Personal Data is stored and processed on computers and servers in the United States and, through your use of the Services, you unequivocally consent to the processing and storage of your Personal Data. All of your data and information is stored in separate, mutually exclusive locations in our cloud storage infrastructure. We use industry standard solutions from Google, Microsoft, and Amazon to store your data and information in the cloud. You should be aware that we may transfer or process your Personal Data in countries other than the country in which you are a resident. These countries may have data protection laws that are different than the laws of your country, and in some cases may not be as protective. Wherever your Personal Data is transferred, stored, or process, Voxel51 uses commercially reasonable safeguards to store and preserve the security of all information and data collected using the Services. For example, we take reasonable steps, such as requesting a unique password, to verify your identity before granting you access to your account. You are responsible for maintaining the secrecy of your unique password and account information at all times. Though we undertake commercially reasonable efforts to protect the information and data that you provide to us, Voxel51 cannot guarantee that the information and data used in connection with the Services will not be accessed, altered, or destroyed. Accordingly, you provide all such information and data at your own risk. In a further effort to protect the privacy of your Personal Data, Voxel51 uses industry standard Secure Sockets Layer (SSL) encryption in connection with the Services. Voxel51 does not store raw user passwords. Voxel51 may store billing/payment information either through secure third-party payment providers or internally via industry-standard PCI DSS compliant protocols. Instead, Voxel51 uses industry standard cryptographic tools to compute hashes of the passwords that are stored in our database and used for subsequent authentication. Also, we require that all user interaction with our cloud system be authenticated by a private access token tied uniquely to the user’s account. In the event that any of your data or information, under our control, is compromised as a result of a breach of security, Voxel51 will take reasonable steps to investigate the situation and where appropriate, notify the individuals whose data or information may have been compromised. Voxel51 will also take any other appropriate steps that are in accordance with any applicable laws and regulations. You understand and agree that Voxel51 may continue to store your information after you cease use of the Services or disable your account. **Cookies and similar technologies** When you use our Services, we use cookies and other similar tracking technologies like “web beacons” to collect and use Personal Data about you, including to allow us to maximize the performance of our Services and to personalize your experience. You have a variety of tools available to control the Personal Data collected by Cookies, web beacons, and similar technologies. For example, you can use controls in your Internet browser to limit how the websites you visit are able to use Cookies and to withdraw your consent by clearing or blocking Cookies. You can also stop Voxel51 from placing Cookies on your device by contacting support@voxel51.com to opt-out. **Voxel51 may share Personal Data with third parties in the following circumstances:** • Where Voxel51 has obtained your consent; • Where sharing or disclosure of your Personal Data with trusted third-party business partners is necessary to provide you with the Services (where these business partners will be given limited access to your Personal Data, as reasonably necessary, to provide you with the Services); • Where sharing or disclosure of your Personal Data is necessary to share Personal Data with Voxel51’s parents, subsidiaries, successors, assigns, licensees, affiliates, or business partners; • Where Voxel51 (or any combination of its product, services, assets, and/or businesses) has been purchased by, acquired by, or merged with a third party; • Where sharing or disclosure of your Personal Data is necessary to respond to requests by government authorities; • Where your Personal Data is demanded by a court order or subpoena; • Where sharing or disclosure of your Personal Data is needed to help prevent against fraud or the violation of any applicable law, statute, regulation, ordinance, or treaty; and • Where Voxel51 is otherwise legally obligated to share your Personal Data. **Users’ Rights Under the GDPR:** The GDPR provides Users located in the European Economic Area (EEA) under its protection certain rights with respect to their Personal Data collected by us on the Website. If you are a User from the EEA, where we are collecting your Personal Data as a controller, our legal basis for doing so will depend on the Personal Data concerned and the specific context in which we collect it. Accordingly, Voxel51 recognizes and will comply with the GDPR and those rights, except as limited by applicable law. The rights under the GDPR include: ● **Right of Access:** This includes the right to obtain from us your Personal Data and whether it is being processed, along with the purposes of the processing; categories of Personal Data concerned; recipients to whom your Personal Data has been disclosed; the period for which your Personal Data is being stored; and the right to lodge a complaint. ● **Right of Rectification:** This includes the right to correct inaccurate Personal Data collected and/or stored by us. ● **Right of Erasure (“Right to be Forgotten”):** This includes the right to have your Personal Data deleted. However, if applicable law requires us to comply with your request to delete information, fulfilment of your request may prevent you from using our Services and may result in closing your account. ● **Right to Restriction of Processing:** This includes the right to request restriction of how and why your Personal Data is used or processed by us. ● **Right to Data Portability:** This includes the right to receive your Personal Data in a structured, readable format and the right to have your Personal Data transferred, ● **Right to Object:** This includes the right to object to us processing your Personal Data for reasons such as direct marketing purposes and for historical research or statistical purposes. ● **Right to not be Subject to Automated Decision-Making:** This includes the right to not be subject to a decision based solely on automated processing, including profiling, that could have a legal effect on you from being made solely based on automated processes. **Purchase or sale of Services or other assets:** Voxel51 may purchase or acquire other businesses or sell components of its business. In the event Voxel51 purchases another business or sells any component of its business, your Personal Data will continue to be used consistent with the terms of this Privacy Policy. **Your rights, controls, and choices:** You can stop Voxel51 from collecting your Personal Data by contacting Voxel51 at support@voxel51.com and requesting that Voxel51 stop collecting your Personal Data. Additionally, you can adjust your web browser setting to limit or turn off Cookies or other tracking techniques, or you can cease use of the Services. You may contact Voxel51 with any requests regarding your Personal Data. Voxel51 does not honor Do Not Track requests and signals. You can access, review, change, update, or delete your Personal Data at any time. Please note that we may impose a small fee for access and disclosure of your Personal Data where permitted under applicable law, which will be communicated to you. We do not charge you to update or remove your Personal Data. **When using the Services, you are obligated to:** Inform Voxel51 of any changes to your Personal Data, and protect the security of your username, password, and your Personal Data. **California Residents:** Under California’s “Shine the Light Law,” California residents have the right to receive information that identifies any third-party companies or individuals that Voxel51 has shared your Personal Data with in the previous calendar year, as well as a description of the categories of Personal Data disclosed to that third party. You may obtain this information once a year and free of charge by contacting Voxel51 at the address below. **Children’s Online Privacy Protection Policy:** The Services are not intended for or directed to users under the age of 18, and Voxel51 does not knowingly or intentionally collect Personal Data from children under the age of 13 or other minors. Where appropriate, Voxel51 takes reasonable measures to determine that users are adults of legal age and to inform minors not to submit such information to the Services. If you are concerned that Personal Data may have been inadvertently provided to or collected by Voxel51, please contact us immediately so appropriate steps may be taken to remove such information from Voxel51’s database. **Third-Party Services:** Voxel51 may outsource or subcontract some of our technical support, tracking and reporting functions, database management functions, and other services to third parties. We may share information from or about you with them so that they can perform their services as long as they comply with this Privacy Policy. You understand and agree that Voxel51 will not be held responsible for any third-party communications sent by entities that Voxel51 does not own or control. You are encouraged to review any third-party privacy policies before utilizing any such third-party service. Voxel51 may include or offer third party products or services via the Services provided by Voxel51 and provide third party links to the same. These third-party websites have separate and independent privacy policies. Voxel51 has no responsibility or liability for the content and activities of such third parties and their websites. We encourage you to read carefully the privacy policies of all such third-party websites. We seek to protect the integrity of the Services and therefore welcome any feedback about any such third-party websites. **Contact and Notices:** All questions and concerns regarding this Privacy Policy may be submitted to Voxel51 at info@voxel51.com. Voxel51, Inc. 330 E. Liberty St. Ann Arbor, MI 48104 [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-16-lllmstxt|> ## Evaluating AI Models [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) Best Practices for Evaluating AI Models Accurately Dec 17, 2024 • 13 min read Article content In this article [Best Practices for Evaluating AI Models Accurately](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#60b61d0676e5) [Why Evaluation Matters](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#75b200115c14) [Risks of Poor AI Model Evaluation](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#a51cb6800050) [Foundational Principles of Model Evaluation](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#0defadbb4205) [Selecting the Right Evaluation Metrics Beyond Accuracy](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#93e4d7d2a99d) [Measuring Performance for Different Tasks](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#f898c1e54145) [Data Splitting Strategies for Robust Evaluation](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#208bade114b7) [Leveraging FiftyOne for Data Splitting](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#84b4c79d2631) [Mitigating Bias and Ensuring Fairness in Evaluation](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#a6702afa8767) [Leveraging FiftyOne for Streamlined Model Evaluation](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#de9e8f04fc84) [Powerful Visualization Tools for In-Depth Analysis](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#e299edbf72bf) [Use Cases: Evaluating Models for Success](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#20d2f99fc28f) [The Road Ahead: Continuous Learning and Improvement](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#1c4cd24ed7c5) [Conclusion](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#e244b9aaf626) [Next steps](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#69cb630b1b11) In this article [Best Practices for Evaluating AI Models Accurately](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#60b61d0676e5) [Why Evaluation Matters](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#75b200115c14) [Risks of Poor AI Model Evaluation](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#a51cb6800050) [Foundational Principles of Model Evaluation](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#0defadbb4205) [Selecting the Right Evaluation Metrics Beyond Accuracy](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#93e4d7d2a99d) [Measuring Performance for Different Tasks](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#f898c1e54145) [Data Splitting Strategies for Robust Evaluation](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#208bade114b7) [Leveraging FiftyOne for Data Splitting](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#84b4c79d2631) [Mitigating Bias and Ensuring Fairness in Evaluation](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#a6702afa8767) [Leveraging FiftyOne for Streamlined Model Evaluation](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#de9e8f04fc84) [Powerful Visualization Tools for In-Depth Analysis](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#e299edbf72bf) [Use Cases: Evaluating Models for Success](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#20d2f99fc28f) [The Road Ahead: Continuous Learning and Improvement](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#1c4cd24ed7c5) [Conclusion](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#e244b9aaf626) [Next steps](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately#69cb630b1b11) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) # Best Practices for Evaluating AI Models Accurately Accurately evaluating the performance of AI models is an important step during development. Model evaluation provides machine learning engineers insights into the strengths and weaknesses of their models. Through the lens of model evaluation metrics, development teams are able to refine and improve their models to meet the desired performance. When multiple models are in the equation, model evaluation metrics enable the systematic comparison of models and aid in choosing the best one for the use case. In this article, we’ll discuss the best practices for evaluating models. We’ll focus on: - Foundational principles of model evaluation - Choosing suitable metrics for your specific tasks and projects - Leveraging tools such as [**FiftyOne**](https://voxel51.com/fiftyone/) to streamline [**model evaluation**](https://docs.voxel51.com/user_guide/evaluation.html) At the end of this article, you’ll walk away with a better understanding of model evaluation best practices and how to use these in your ML work. Whether you’re building models for business applications, safety-critical systems, customer engagement, or more, these insights will empower you to create reliable, high-performing AI systems. ## Why Evaluation Matters AI models have become indispensable across industries such as healthcare, aerospace, retail, finance, and more, solving complex problems and driving data-driven decisions. However, their effectiveness depends entirely on accurate and thorough evaluation. For example: - [**Healthcare**](https://voxel51.com/computer-vision-use-cases/healthcare/): A misdiagnosis from a poorly evaluated model could lead to critical health risks. - [**Retail**](https://voxel51.com/computer-vision-use-cases/retail/using-computer-vision-to-enhance-customer-experience-in-retail/): An inaccurate recommendation system may fail to engage users, costing businesses opportunities and revenue. - [**Autonomous Vehicles**](https://voxel51.com/computer-vision-use-cases/driving/): Safety-critical systems demand rigorous testing to avoid accidents caused by misclassifications. ## Risks of Poor AI Model Evaluation When evaluating AI models, accuracy is paramount. Poor evaluation can lead to costly errors, compromised user safety, and reduced trust in AI systems. Let’s explore the significant risks of inadequate evaluation and why each deserves serious consideration. - **Unaddressed Model Biases:** Inadequate testing may allow biases to persist, leading to unfair outcomes and reputation/legal risks. - **Resource Wastage:** Rushed evaluations can result in ineffective models, wasting time, money, and effort on solutions that fail - **Safety Risks:** Poorly evaluated models can cause safety hazards in critical applications like self-driving cars or medical diagnostics. - **False Confidence in Metrics:** Over-reliance on simplified metrics like accuracy can mask real-world performance issues, delaying necessary fixes. ## Foundational Principles of Model Evaluation Building and deploying AI models that perform reliably in real-world settings isn’t easy, but following some core principles can go a long way in making it possible. These principles guide AI teams in designing evaluation processes that lead to dependable, high-quality models ready for real-world applications. Let’s take a closer look at what these principles are and why they matter. ## Selecting the Right Evaluation Metrics Beyond Accuracy Accuracy is a common evaluation metric, but relying on it alone can lead to misleading insights, especially in complex or high-stakes scenarios. For example, in a medical diagnosis task for a rare disease, a machine learning model that predicts “no disease” for every case might achieve 99% accuracy but completely fail to identify actual cases. To evaluate models effectively, it’s essential to consider metrics that reflect their performance more holistically. ### Common Evaluation Metrics Metric What it Measures Best Used For Precision Proportion of true positives among all positive predictions. Tasks where a false positive rate affects costs (e.g., spam detection). Recall Proportion of true positives identified out of all actual positives. Tasks where false negatives are critical (e.g., medical diagnosis). F1 Score Harmonic mean of precision and recall, balancing the two metrics. When both precision and recall are equally important (e.g., fraud detection). Intersection over Union (IoU) Degree of overlap between the predicted and actual values in spatial tasks. Object detection or segmentation tasks that require precise localization (e.g., self-driving cars). Mean Average Precision (mAP) How well a model detects objects by considering both Precision and Recall. Measuring overall performance of detection and segmentation tasks (e.g., facial recognition, visual search). Confusion Matrix Visual representation of predicted classification against known classifications with a breakdown of correct predictions as well as incorrect predictions To assess the effectiveness of any classification model by identifying where it makes mistakes and which classes are misclassified. Domain-Specific Metrics Custom metrics tailored to specific industries or applications. Examples include mean absolute error (MAE) for forecasting in finance. ## Measuring Performance for Different Tasks Using a single performance metric can obscure the model’s true performance. A multi-metric approach, depending on the task, is usually a more pragmatic approach. In this section, we’ll talk about the different classification metrics, object detection, and segmentation models. ### [Classification](https://docs.voxel51.com/user_guide/evaluation.html\#classifications) The objective of classification models is to find patterns within the data based on finite predefined categories or classes based on features and use that data to predict the class of new, unseen data points. Binary classification refers to predicting two classes, and multi-class classification is involved in predicting, as its name suggests, more than two classes. For example: detecting ‘spam’ or ‘non-spam’ calls is a binary classification, and predicting types of animals or plants in an image is a multi-class system. Identifying objects within an image and assigning them to specific categories falls into image classification. The typical metrics used for classification models are **accuracy**, **precision**, **recall**, **F1 score**, and **confusion matrix**. For example: In fraud detection, combining precision, recall, and F1 score ensures the model is effective at capturing true cases while minimizing false alarms. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ### [Detection](https://docs.voxel51.com/user_guide/evaluation.html\#detections) and [Segmentation](https://docs.voxel51.com/user_guide/evaluation.html\#semantic-segmentations) The goal of object detection is to identify and localize objects within an image or a video. Typically algorithms such as YOLO, R-CNN, F-CNN, and SSD are used for detection tasks, and bounding boxes are drawn around the objects. Segmentation, on the other hand, partitions the image into distinct and meaningful regions where each pixel is associated with a class label. Example: In this image, the bounding box identifies the object, and the segmentation identifies the exact region of the object. The typical metrics used for detection and segmentation models are **Intersection over Union (IoU)** and **Mean Average Precision (mAP)** ![](https://cdn.sanity.io/images/h6toihm1/production/a1d9018ace08f4c044ec0206d1a925f20bc30ee0-1049x590.png?auto=format&dpr=2&fit=max&q=75&w=1049) ## Data Splitting Strategies for Robust Evaluation Creating AI models that perform well in real-world settings requires a solid evaluation process, and data splitting is a critical part of it. How you split your data determines how well your model will generalize to new data. Let’s explore the main strategies for effective data splitting and how they can strengthen your evaluation process. ### Train-Test Split: The Basic Starting Point The **train-test split** is the simplest approach, where you divide your data into two parts: one for training the model and another for testing its performance. A typical split is **80-20** or **70-30**, with the larger portion for training. - **Pros**: Quick and easy to implement, good for a fast performance snapshot. - **Cons**: Sensitive to data variability; may not provide stable results with small or imbalanced datasets. This method is useful for initial experiments, but for more robust evaluation, advanced techniques are preferred. ### Cross-Validation: A More Reliable Alternative **Cross-validation** provides a more comprehensive view of the model’s performance by utilizing multiple train-test splits. One popular technique is **k-fold cross-validation**, where the dataset is divided into **k subsets (folds)**. The model is trained on k−1k-1k−1 folds and tested on the remaining fold, cycling through until every subset has been used for testing once. - **Pros**: Reduces bias by using all data for training and testing; provides stable metrics even with limited data. - **Cons**: Computationally intensive, especially with large datasets or high k-values. Here’s an example in Python demonstrating **k-fold cross-validation** using the scikit-learn library: ```python 1#install sklearn by writing pip install scikit-learn 2#import necessary packages 3 4from sklearn.model_selection import KFold, cross_val_score 5from sklearn.ensemble import RandomForestClassifier 6from sklearn.datasets import load_iris 7 8# Load a sample dataset 9data = load_iris() 10X, y = data.data, data.target 11 12# Initialize the model 13model = RandomForestClassifier(random_state=42) 14 15# Set up k-fold cross-validation (k=5) 16kfold = KFold(n_splits=5, shuffle=True, random_state=42) 17 18# Evaluate the model 19scores = cross_val_score(model, X, y, cv=kfold) 20 21# Display results 22print(f"Cross-Validation Scores: {scores}") 23print(f"Average Score: {scores.mean():.2f}") 24 ``` ### When to Use Cross-Validation - **Small Datasets**: Maximizes the utility of limited data. - **Unbalanced Datasets**: Ensures all data subsets contribute to evaluation. - **Model Benchmarking**: Provides more reliable metrics for comparing different algorithms. ## Leveraging FiftyOne for Data Splitting Splitting data correctly can get complicated, but the [FiftyOne](https://voxel51.com/) tool makes it much easier. FiftyOne offers the ability to manage data splitting, visualize your splits, and even apply custom strategies like stratified sampling. This becomes especially useful if you’re working with large or diverse datasets, where manually handling splits could lead to errors or inconsistencies. Creating dataset splits (test, train, validation) is a perfect use case for [tags in FiftyOne](https://docs.voxel51.com/user_guide/basics.html#tags) and can be implemented very easily: ```python 1import fiftyone as fo 2sample = fo.Sample(filepath="/path/to/image.png", tags=["train"]) 3sample.tags.append("my_favorite_samples") 4print(sample.tags) 5# ["train", "my_favorite_samples"] ``` ## Mitigating Bias and Ensuring Fairness in Evaluation Bias is a huge challenge in AI, and addressing it is crucial to building models that are fair, reliable, and genuinely beneficial. When bias is left unchecked, models can produce unfair outcomes or skewed results that impact real people, especially in areas like hiring, healthcare, and lending. Tackling bias during model evaluation helps ensure fairness and builds trust, especially in systems that affect diverse communities. Here’s why bias mitigation matters and some practical ways to do it. ### Recognizing Bias in Training Data and Test Sets AI models learn from data, so if there’s bias in the training or test data, the model will likely learn that bias too. For example, if a facial recognition model is trained mostly on images of lighter-skinned faces, it might struggle to accurately recognize darker-skinned faces, leading to biased outcomes. The first step in bias mitigation is recognizing where your data might lack diversity or represent one group more than another. This way, you can better understand where the model might struggle and start planning ways to improve its fairness. ### Using Stratified Sampling to Address Imbalances Imbalanced data is a common issue in AI where certain categories or demographics may dominate the dataset, especially if the data wasn’t collected with fairness in mind. To address this, stratified sampling is a helpful technique. Stratified sampling ensures that each subgroup is properly represented in both the training and testing sets, keeping the dataset balanced. ### Testing Across Diverse Demographic Groups Fairness isn’t achieved if a model performs well overall but fails for specific groups. To address this, it’s essential to evaluate model performance across different demographic groups, such as age, gender, race, and socioeconomic status. By doing this, you can identify if the model has any hidden performance gaps. ## Leveraging FiftyOne for Streamlined Model Evaluation Evaluating an AI model thoroughly can get complex, especially when you’re juggling different metrics, data splits, and analysis tools. FiftyOne steps in as a powerful ally for AI builders, helping simplify and streamline model evaluation. It’s designed to help teams save time, avoid errors, and make more informed decisions, all essential for developing models that truly perform well in real-world settings. Here’s an example of how to work with aggregate metrics in FiftyOne using the Python SDK: ```python 1# Get the 10 most common classes in the dataset 2counts = dataset.count_values("ground_truth.detections.label") 3classes = sorted(counts, key=counts.get, reverse=True)[:10] 4 5# Print a classification report for the top-10 classes 6results.print_report(classes=classes) 7 precision recall f1-score support 8 person 0.45 0.74 0.56 783 9 kite 0.55 0.72 0.62 156 10 car 0.12 0.54 0.20 61 11 bird 0.63 0.67 0.65 126 12 carrot 0.06 0.49 0.11 47 13 boat 0.05 0.24 0.08 37 14 surfboard 0.10 0.43 0.17 30 15traffic light 0.22 0.54 0.31 24 16 airplane 0.29 0.67 0.40 24 17 giraffe 0.26 0.65 0.37 23 18 micro avg 0.32 0.68 0.44 1311 19 macro avg 0.27 0.57 0.35 1311 20 weighted avg 0.42 0.68 0.51 1311 21 22 ``` Here’s what it looks like to get a side-by-side comparison of metrics on two models using the Model Evaluation Panel in the FiftyOne Application. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop FiftyOne provides a variety of builtin methods for evaluating your model predictions, including regressions, classifications, detections, polygons, instance, and semantic segmentation, on both image and video datasets. When you evaluate a model in FiftyOne, you get access to the standard aggregate metrics such as mAP, Precision, IOU, classification reports, confusion matrices, and PR curves. In addition, FiftyOne also provides fine-grained statistics like accuracy and false positive counts at the sample level, which you can interactively explore to diagnose the strengths and weaknesses of your models on individual data samples. Analyzing each metric and understanding which samples are causing poor scores and why helps uncover insights and find areas of gaps in your data that are helpful to iteratively improve model performance. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Check out the [model evaluation capabilities](https://docs.voxel51.com/user_guide/evaluation.html) in FiftyOne and try the [Model Evaluation panel](https://docs.voxel51.com/user_guide/evaluation.html#model-evaluation-panel-sub-new) that provides an out-of-the-box experience to visualize and interactively explore the evaluation results in the App. ## Powerful Visualization Tools for In-Depth Analysis One of the standout features of FiftyOne is its [visualization capabilities](https://docs.voxel51.com/user_guide/plots.html#interactive-plots), which allow AI teams to dig deep into model performance. With FiftyOne, you can visualize everything from confusion matrices to error distributions, gaining insights into where your model is excelling and struggling. For example, if a model produces a high rate of false positives, visualization tools can help you pinpoint specific patterns or data segments causing the issue. These insights are crucial for improving model accuracy, reducing bias, and addressing weaknesses that might impact performance. ## Use Cases: Evaluating Models for Success Evaluating AI models isn’t just a theoretical exercise, it has real, tangible consequences ensuring models perform effectively in the field. The next section provides detailed examples of how robust evaluation impacts success, with a brief mention of other potential use cases for those interested in exploring further. ### Ensuring Safety in Self-Driving Cars with Object Detection Models Self-driving cars are a prime example of how AI models must be rigorously evaluated to meet safety and reliability standards. An object detection model in these systems identifies pedestrians, vehicles, traffic signs, and obstacles, playing a critical role in enabling safe navigation. To ensure reliable performance, developers test the model on diverse datasets representing real-world driving conditions: - **Lighting Variations**: Daylight, dusk, and nighttime scenarios. - **Weather Conditions**: Rain, fog, snow, or bright sunlight. - **Environments**: Urban traffic, rural roads, and highways with varying levels of congestion. Evaluation practices often involve metrics like **Intersection over Union (IoU)** for assessing the accuracy of object localization and **precision/recall** for detecting critical objects (e.g., pedestrians). By exposing the model to such a wide range of scenarios, developers can identify blind spots and improve its robustness. These rigorous evaluations reduce the risk of accidents, enhance user trust, and accelerate the adoption of autonomous driving technologies. The ability to generalize across diverse conditions is key to achieving reliability and safety in such high-stakes applications. ### Broader Applications and Further Reading Robust evaluation isn’t just vital for self-driving cars, it applies to numerous fields. For instance: - **Healthcare**: Evaluating medical image analysis models to ensure accurate and equitable diagnoses. - **Security**: Refining facial recognition models to eliminate biases across demographic groups. To learn more about how evaluation shapes success in AI, explore additional resources on [medical](https://voxel51.com/computer-vision-use-cases/healthcare/) and [security](https://voxel51.com/computer-vision-use-cases/security/) use cases. ## The Road Ahead: Continuous Learning and Improvement AI models aren’t “one-and-done” solutions; they need ongoing evaluation to stay effective. As the world and data evolve, so must your models. Below are key strategies for ensuring long-term success ### Using A/B Testing to Compare Performance A/B testing is a powerful tool for continuous evaluation. By running two versions of a model and comparing their results, you can see which one performs better in real time. It’s a great way to test new features or changes, allowing you to make data-driven decisions without risking your entire model. ### Data Logging and Monitoring Real-time monitoring is essential for identifying issues as they arise. By implementing data logging and monitoring tools, you can track model predictions and detect any unexpected errors. If something goes wrong, like a sudden drop in accuracy, these tools will alert you right away, helping you address the problem quickly. ### Re-Evaluating with New Data As time goes on, the data you’re using may no longer represent the current landscape. Regularly re-evaluating your model with fresh data helps it stay relevant and accurate. For example, retraining a fraud detection system with new fraud patterns ensures it keeps up with evolving tactics. This ongoing process helps your model adapt and continue performing at its best. ## Conclusion Adopting best practices for AI model evaluation is essential for building reliable, high-performing models. By carefully selecting evaluation metrics, using diverse datasets, and continually assessing model performance, you can ensure that your AI applications succeed. This approach minimizes errors, reduces biases, and boosts trust in your AI systems. FiftyOne simplifies the evaluation process, providing AI developers with the tools needed to streamline testing, monitor performance, and visualize results. Whether you’re fine-tuning an object detection model or ensuring fairness in a classification system, FiftyOne helps you achieve impactful, reliable results. ## Next steps Voxel51 has made it easy to [get started](https://docs.voxel51.com/getting_started/install.html) evaluating AI models accurately with FiftyOne. Looking for a scalable solution for your ML team as you collaborate on visual AI projects? Check out [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) and [connect with an expert](https://voxel51.com/book-a-demo/) to see the collaborative features of FiftyOne in action. [Evaluation](https://voxel51.com/blog/tag/evaluation) [model evaluation](https://voxel51.com/blog/tag/model-evaluation) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/b83b51550d5f3f3fc2f98dbacb2996313896e147-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Quality Dataset Annotation Is Key to Machine Learning\\ \\ Learn\\ \\ • \\ \\ Feb 17, 2025](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning) [![](https://cdn.sanity.io/images/h6toihm1/production/d2e24d0a14de508f8ccd36eaffd3b909c9f193b9-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Image Embeddings Transform Computer Vision Capabilities\\ \\ Learn\\ \\ • \\ \\ Nov 25, 2024](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities) [![](https://cdn.sanity.io/images/h6toihm1/production/50ab5d62a585ad15e0d6c26c224c48e40f345266-2258x1264.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Implementing Mask R-CNN: Advanced Object Detection and Segmentation\\ \\ Learn\\ \\ • \\ \\ Apr 3, 2025](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-17-lllmstxt|> ## Image Preprocessing Best Practices [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) Image Preprocessing Best Practices To Optimize Your AI Workflows Apr 17, 2025 • 8 min read Article content In this article [The Underestimated Importance of Image Preprocessing](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#4f70ff5af1c1) [The Impact of Preprocessing on Performance](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#89bc6e03fbf6) [Advanced Image Preprocessing Techniques](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#14e4a1784dad) [The Impact of Preprocessing on Model Performance](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#5fab5608cd60) [Leveraging FiftyOne for Advanced Preprocessing](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#4373ac67e0ef) [First Impressions Matter](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#ebc17b3d5f70) In this article [The Underestimated Importance of Image Preprocessing](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#4f70ff5af1c1) [The Impact of Preprocessing on Performance](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#89bc6e03fbf6) [Advanced Image Preprocessing Techniques](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#14e4a1784dad) [The Impact of Preprocessing on Model Performance](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#5fab5608cd60) [Leveraging FiftyOne for Advanced Preprocessing](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#4373ac67e0ef) [First Impressions Matter](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows#ebc17b3d5f70) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/551f4c0ff851bce8ff3601a10107345619d138c3-612x612.gif?auto=format&dpr=2&fit=max&q=75&w=612) Modern computer vision systems depend on complex models and [well-curated datasets](https://voxel51.com/blog/data-quality-the-hidden-driver-of-ai-success/) for tasks like object detection and segmentation. However, even the most sophisticated models struggle when presented with noisy, inconsistent, or poorly formatted images. Introducing effective image preprocessing significantly improves performance, efficiency, and image quality. Though often overshadowed by model architecture and datasets, basic transformations such as resizing, normalization, and contrast enhancement profoundly impact training stability, feature representation, and convergence. This article explores essential preprocessing techniques, from fundamental adjustments to advanced domain-specific pipelines, highlighting how integration with the [FiftyOne](https://voxel51.com/fiftyone/) platform streamlines workflows and enables robust, high-performing visual AI solutions. ### **Definition of Image Preprocessing** _Image preprocessing_ refers to the set of techniques applied to raw images before they are used for training computer vision models. This process includes resizing, normalization, noise reduction, color correction, and other transformations aimed at enhancing image quality, consistency, and compatibility with model expectations. Effective preprocessing ensures that models can learn more efficiently, generalize better, and achieve higher performance across diverse visual conditions. ## **The Underestimated Importance of Image Preprocessing** ### **Preprocessing as a Foundational Step** ![](https://cdn.sanity.io/images/h6toihm1/production/9bd9fbb5bbc640e852998c0e259704903ffae648-1200x600.png?auto=format&dpr=2&fit=max&q=75&w=1200) When dealing with raw image data, cameras and sensors can produce images with noise, motion blur, or inconsistent lighting. A typical pipeline might only resize images and then feed them into a deep network, assuming advanced image processing “just works.” In reality, front-loading thoughtful preprocessing can prevent distorted image features and help the model learn more effectively. ```generic 1def basic_preprocess(img_pil): 2 # Random horizontal flip 3 if random.random() < 0.5: 4 img_pil = img_pil.transpose(Image.FLIP_LEFT_RIGHT) 5 # Increase contrast slightly 6 return ImageEnhance.Contrast(img_pil).enhance(1.2) 7 ``` ### **The Profound Impact of “Simple” Transformations** ![](https://cdn.sanity.io/images/h6toihm1/production/154090907c23ac9bf9c48bb90b8710f3ab0596b8-1800x600.png?auto=format&dpr=2&fit=max&q=75&w=1600) Thoughtful image preprocessing, including basic steps like aspect-ratio-preserving resizing and proper contrast/noise adjustments, significantly alters data distributions; employing these correctly, or using advanced adaptive strategies, optimizes images to support more accurate and efficient models by stabilizing training and highlighting crucial features. ### **Beyond Resizing and Normalization** While resizing and normalization are common, strategic preprocessing can reduce reliance on large augmentations or multi-stage fine-tuning. Techniques like domain adaptation, style transfer, and adaptive transforms let you tailor image preprocessing to evolving conditions, bridging gaps between synthetic and real scenes (or day vs. night) and improving generalization. Ultimately, image processing is a critical, proactive step for robust modeling. ## **The Impact of Preprocessing on Performance** Thoughtful image preprocessing enhances: - **Stability & Convergence**: Standardizing color or intensity values helps models train faster and avoid erratic gradients. - **Generalization**: Handling noise, distortions, or domain mismatches at the preprocessing stage reduces the risk of overfitting. - **Reduced Data Augmentation & Fine-Tuning**: When preprocessing handles domain shifts (like lighting changes), models require fewer specialized augmentations. Focusing on adaptive or advanced image processing can yield quicker training and more stable results. ## **Advanced Image Preprocessing Techniques** As digital images in real-world applications become increasingly diverse, going beyond basic augmentations is essential. Applying advanced image processing algorithms and techniques can significantly improve image analysis pipelines. Below are some examples: ### **Beyond Basic Augmentations** #### **CutMix and Mixup** ![](https://cdn.sanity.io/images/h6toihm1/production/2b5d4bcd8794e1dcbf8cc17f109734de63bed57e-1800x600.png?auto=format&dpr=2&fit=max&q=75&w=1600) - [CutMix](https://pytorch.org/vision/main/auto_examples/transforms/plot_cutmix_mixup.html): Pastes a rectangular patch from one image onto another, helping models handle partial occlusions and boundary ambiguities. - [Mixup](https://pytorch.org/vision/main/auto_examples/transforms/plot_cutmix_mixup.html): Interpolates two images (and labels), smoothing decision boundaries and lowering overfitting. ```python 1def cutmix(img1, img2): 2 w, h = img1.size 3 rx, ry = w//4, h//4 4 region = img2.crop((0, 0, rx, ry)) 5 img1.paste(region, (0, 0)) # Overwrite top-left corner 6 return img1 7 ``` Both methods strengthen model robustness by forcing the network to blend varying contexts. #### **Style Transfer Augmentation** ![](https://cdn.sanity.io/images/h6toihm1/production/05c72a1764659ad975a3ad066e178d99fa4959ca-1800x600.png?auto=format&dpr=2&fit=max&q=75&w=1600) Style transfer modifies surface details (e.g., color or texture) while preserving object shapes. This trains models to focus on core features rather than superficial differences like lighting or weather, making them more flexible in varied conditions. ```python 1def color_transfer_simple(src_bgr, ref_bgr): 2 src_lab = cv2.cvtColor(src_bgr, cv2.COLOR_BGR2LAB).astype(float) 3 ref_lab = cv2.cvtColor(ref_bgr, cv2.COLOR_BGR2LAB).astype(float) 4 # Match mean + std of LAB channels 5 src_lab = (src_lab - src_lab.mean()) * (ref_lab.std() / src_lab.std()) + ref_lab.mean() 6 return cv2.cvtColor(src_lab.clip(0,255).astype('uint8'), cv2.COLOR_LAB2BGR) 7 ``` By injecting a variety of stylistic cues, models become less domain-specific and more resilient to environmental changes. #### **Domain Adaptation Techniques** Domain adaptation addresses discrepancies between training and deployment environments (e.g., synthetic vs. real images). Preprocessing can reduce this domain shift by normalizing color, brightness, or geometry. In medical imaging, for instance, specialized transformations like histogram matching or intensity standardization help models better align with target scanning protocols. ```python 1def histogram_match(src_bgr, ref_bgr): 2 src_ycc = cv2.cvtColor(src_bgr, cv2.COLOR_BGR2YCrCb) 3 ref_ycc = cv2.cvtColor(ref_bgr, cv2.COLOR_BGR2YCrCb) 4 # Simple channel-by-channel equalization 5 for i in range(3): 6 src_ycc[..., i] = cv2.equalizeHist(src_ycc[..., i]) 7 return cv2.cvtColor(src_ycc, cv2.COLOR_YCrCb2BGR) 8 ``` #### **Self-Supervised Preprocessing** Self-supervised learning leverages unlabeled data through pretext tasks (e.g., predicting rotations or solving jigsaw puzzles) to learn meaningful representations of raw image data before formal training. By embedding self-supervised tasks in the preprocessing pipeline, you gain robust initial embeddings with less labeled data, potentially boosting object detection or image segmentation tasks. ## **The Impact of Preprocessing on Model Performance** ![](https://cdn.sanity.io/images/h6toihm1/production/378f8cef1e1800c081c29b0ab3c20681ee4e7221-1200x600.png?auto=format&dpr=2&fit=max&q=75&w=1200) ### **Beyond Accuracy: Evaluating the Impact of Preprocessing** #### **Calibration** Beyond improving accuracy and robustness, effective image preprocessing also influences other critical model evaluation criteria such as calibration, adversarial robustness, and fairness. These factors are essential for deploying models reliably in real-world scenarios. Model calibration measures how well predicted probabilities match real-world likelihoods. Excessive contrast enhancement or aggressive color jitter can induce overconfidence. While methods like temperature scaling can fix calibration post-hoc, well-planned image preprocessing ensures consistent distributions, reducing calibration issues. #### **Adversarial Robustness** Small noise patterns or perturbations can fool unprotected models. Defensive transformations (e.g., mild blurring or randomization) during preprocessing can disrupt adversarial attack vectors. Similarly, noise injection fosters more stable features and helps the model resist pixel-level manipulations. #### **Fairness and Bias** Preprocessing can reveal and mitigate biases in datasets, for example by balancing classes or normalizing conditions across demographic groups. Tools like outlier detection and domain-specific augmentations help ensure that no subset of images skews the model’s performance unfairly. ## **Leveraging FiftyOne for Advanced Preprocessing** FiftyOne is a powerful platform that unifies image processing experiments, data exploration, and performance tracking, critical for building strong preprocessing workflows. ### **Interactive Data Exploration and Analysis** ![](https://cdn.sanity.io/images/h6toihm1/production/9a5b8c009bbe17918b32a5f99f0fffe89f498f56-3826x1344.png?auto=format&dpr=2&fit=max&q=75&w=1600) #### **Identifying Potential Biases and Outliers** Skewed class distributions or unusual binary image aspect ratios can sabotage model performance. FiftyOne’s filtering helps you find subsets of data (e.g., overexposed images or underrepresented classes), guiding specialized image preprocessing solutions like brightness normalization or targeted domain augmentation. #### **Visualizing Preprocessing Impact** Side-by-side comparisons in FiftyOne let you confirm if a transformation preserves essential image features or introduces artifacts. This is especially useful for style transfer, domain adaptation, or mixing-based techniques like CutMix or Mixup, helping you gauge if your pipeline is too aggressive or just right. ### **Customizable Preprocessing Pipelines** #### **Defining Complex Pipelines** Real-world image processing techniques often involve multiple steps like denoising, edge detection, thresholding to create binary images, then an augmentation. With FiftyOne, you can build multi-step pipelines, integrate external libraries (OpenCV, albumentations, PyTorch, TensorFlow), and keep all outputs tracked in a central place. #### **Integration with External Libraries** No single library addresses all tasks. FiftyOne provides an open structure, letting you apply specialized transformations from various sources (e.g., scikit-image for classical filters, custom code for style transfer) and store the results for each image version, all in one consistent dataset. ### **Experimentation and Hyperparameter Optimization** #### **Rapid Experimentation** Finding the “best” image preprocessing pipeline usually requires iteration. FiftyOne allows you to clone datasets, tweak transformations, and quickly compare results. By labeling each branch of experimentation, you can precisely track which pipeline yields better performance or fewer errors. #### **Tracking and Comparing Preprocessing Performance** FiftyOne logs metrics (accuracy, IoU, mAP) per preprocessing strategy, making it straightforward to compare them. If CutMix and Mixup produce similar results but one is faster, you’ll see that difference and make a data-driven choice. This feedback loop refines your pipeline and ensures you continue to optimize. ## **First Impressions Matter** ### **Mastering Image Preprocessing for Better Visual AI Models** As image processing challenges grow, ranging from object detection in varying weather to specialized image segmentation in medical domains, thoughtful image preprocessing is the cornerstone of a robust, high-quality model. Techniques like style transfer, adaptive denoising, domain adaptation, and self-supervised representation learning can push your applications toward state-of-the-art performance without excessive complexity. ### **Encouraging Exploration and Experimentation** We invite you to explore advanced image enhancement strategies, experiment with image preprocessing techniques in FiftyOne, and refine your workflows. By combining cutting-edge methods with thorough analysis tools, your image data pipelines can produce more accurate, calibrated, and resilient models ready for real-world deployment. ### **Elevate Your Visual AI with Smarter Preprocessing** Try out FiftyOne for your image preprocessing workflows. Experiment with advanced transformations like Mixup or style transfer, track model performance, and discover how refined pipelines can elevate the success of your visual AI projects. ### **Explore the Jupyter Notebook** We’ve provided a companion [Jupyter notebook](https://colab.research.google.com/drive/1F0k_pfR9uaUnLVaPjNlnkoJ4mObl_hpJ?usp=sharing) demonstrating: 1. Applying various preprocessing pipelines (basic, adaptive, advanced). 2. Storing multiple processed image versions within FiftyOne. 3. Switching views in the FiftyOne App to compare transformations. 4. Computing image statistics (e.g., brightness) and using them to filter/sort in FiftyOne. 5. Evaluating preprocessing impact on object detection performance (YOLOv8). By working through this notebook, you can practice implementing, managing, and evaluating image preprocessing pipelines with FiftyOne. **Image Citations** - Lin, Tsung-Yi, et al. _"Microsoft COCO: Common Objects in Context."_ COCO Dataset 2017 Validation Split, cocodataset.org, 2017, [https://cocodataset.org/#home](https://cocodataset.org/#home). Accessed 25 Mar. 2025. [images](https://voxel51.com/blog/tag/images) [image restoration](https://voxel51.com/blog/tag/image-restoration) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/d2e24d0a14de508f8ccd36eaffd3b909c9f193b9-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Image Embeddings Transform Computer Vision Capabilities\\ \\ Learn\\ \\ • \\ \\ Nov 25, 2024](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities) [![](https://cdn.sanity.io/images/h6toihm1/production/9252e8bb5db5c4805f4a6f315b51527ee5c22072-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Guide to AI Image Segmentation\\ \\ Learn\\ \\ • \\ \\ Dec 19, 2024](https://voxel51.com/blog/a-guide-to-ai-image-segmentation) [![](https://cdn.sanity.io/images/h6toihm1/production/04eff439f21847f1068ea3f38b72f8bc7ec77f0e-2340x1308.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Image Similarity Search: Unlocking Pattern Detection in Visual Data\\ \\ Learn\\ \\ • \\ \\ Apr 16, 2025](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) Mastering Image Preprocessing: Optimizing Your Visual AI Workflow <|firecrawl-page-18-lllmstxt|> ## Mask R-CNN Implementation Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) Implementing Mask R-CNN: Advanced Object Detection and Segmentation Apr 3, 2025 • 10 min read Article content In this article [Understanding Mask R-CNN Architecture](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation#320cdaa99ff6) [Implementing Mask R-CNN](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation#aa619cf93a57) [Leveraging FiftyOne for Mask R-CNN](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation#5a2e95c58ba3) [Applications of Mask R-CNN](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation#ca8d8644f6a8) [Conclusion](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation#3dce76e24819) In this article [Understanding Mask R-CNN Architecture](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation#320cdaa99ff6) [Implementing Mask R-CNN](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation#aa619cf93a57) [Leveraging FiftyOne for Mask R-CNN](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation#5a2e95c58ba3) [Applications of Mask R-CNN](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation#ca8d8644f6a8) [Conclusion](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation#3dce76e24819) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/878ad4ffe165325bdff6478982b74c1ce9f9aa16-560x155.png?auto=format&dpr=2&fit=max&q=75&w=560) In modern computer vision, [object detection](https://en.wikipedia.org/wiki/Object_detection) and [image segmentation](https://en.wikipedia.org/wiki/Image_segmentation), particularly semantic segmentation, are foundational technologies, each with distinct capabilities and limitations. Object detection identifies objects and localizes them with bounding boxes, while semantic segmentation assigns each pixel a class label without distinguishing individual instances. These approaches serve many purposes, but both fall short when applications require both precise boundaries and the separation of individual objects. This is where [instance segmentation](https://voxel51.com/resources/learn/a-guide-to-ai-image-segmentation/) proves invaluable. Instance segmentation combines the strengths of detection and segmentation: each object is not only located by a bounding box but also represented at the pixel level with a precise object mask. When objects overlap or appear partially occluded, common scenarios in real-world applications, instance segmentation provides clarity that other methods cannot. [Mask R-CNN](https://arxiv.org/abs/1703.06870) stands as one of the most influential frameworks for instance segmentation. Building on the successes of Faster R-CNN, the Mask R-CNN framework extends traditional bounding box recognition with object instance segmentation, predicting segmentation masks alongside bounding boxes and class labels. By providing pixel-level precision and distinguishing individual instances, Mask R-CNN outperforms traditional object detection methods in complex real-world scenarios. ![](https://cdn.sanity.io/images/h6toihm1/production/878ad4ffe165325bdff6478982b74c1ce9f9aa16-560x155.png?auto=format&dpr=2&fit=max&q=75&w=560) In this article, we'll explore how Mask R-CNN works, demonstrate its implementation using [FiftyOne](https://voxel51.com/fiftyone/), and examine its practical applications across multiple domains. There’s also a companion [Jupyter notebook](https://colab.research.google.com/drive/1MasO18sQzkG2awKIzrQ6LNh1AgcqRNLX) demonstrating how to: - Set up a small Mask R-CNN instance segmentation dataset - Run inference using Detectron2 - Visualize results interactively in FiftyOne - Evaluate segmentation accuracy and explore failure cases - Consider strategies for fine-tuning Mask R-CNN Work through this notebook to replicate these methods on your data and gain insights into Mask R-CNN’s real-world performance. ## Understanding Mask R-CNN Architecture Mask R-CNN outperforms earlier models primarily due to its carefully refined architecture. While traditional methods focus primarily on bounding box recognition, Mask R-CNN introduces object instance segmentation, predicting detailed pixel-level masks alongside bounding boxes and class labels, making it particularly effective in complex scenes. Let’s break down the key architectural components of the Mask R-CNN framework: ### **Backbone Network (ResNet/FPN)** ![](https://cdn.sanity.io/images/h6toihm1/production/a70c92d151efabb9f708d75179cc605b0011ef2d-872x499.png?auto=format&dpr=2&fit=max&q=75&w=872) The foundation of Mask R-CNN is a deep convolutional backbone network, typically ResNet, which extracts feature maps from the input image. Earlier layers capture basic elements like edges and corners, while deeper layers recognize complex shapes and patterns. This backbone is enhanced with a [Feature Pyramid Network](https://en.wikipedia.org/wiki/Small_object_detection#Feature_Pyramid_Network_(FPN)) (FPN) that generates multi-scale feature representations. The FPN enables the model to detect objects at various sizes—a critical capability when scenes contain both large, prominent objects and small, distant ones. **Region Proposal Network (RPN)** \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop The Region Proposal Network generates candidate bounding boxes (called anchor boxes) that likely contain objects. This focuses computational resources on promising regions rather than exhaustively scanning every pixel. The RPN classifies proposed regions as either foreground (potentially containing objects) or background, efficiently filtering out unlikely areas. ### **ROI Align Layer** ![](https://cdn.sanity.io/images/h6toihm1/production/2e48837e2fa9ea579826cf2d7181abaf427d21dc-1815x963.png?auto=format&dpr=2&fit=max&q=75&w=1600) One of Mask R-CNN's key innovations is the ROI Align layer. Earlier R-CNN variants used ROI Pooling, which discretized bounding box features and lost spatial precision. ROI Align maintains exact spatial correspondence through bilinear interpolation, preserving the precise pixel-level details needed for accurate mask generation. This improvement is particularly important for small objects or those with intricate boundaries. ### **Head Networks** Mask R-CNN employs multiple specialized network "heads" that operate on the features extracted from each region: - **Classification Branch**: Identifies the object class (e.g., "person," "car," "dog") for each proposed region - **Bounding Box Regression Branch**: Fine-tunes the bounding box coordinates for more accurate localization - **Mask Branch**: Outputs a binary segmentation mask for each detected object, providing pixel-precise boundaries ![](https://cdn.sanity.io/images/h6toihm1/production/84624eda9f72725d4b82e05594803bf62a551021-797x621.png?auto=format&dpr=2&fit=max&q=75&w=797) ### **Multi-Task Loss Function** During training, Mask R-CNN optimizes a combined loss function that accounts for classification accuracy, bounding box precision, and mask quality. By simultaneously addressing all three objectives, the network learns to perform detection and segmentation in a unified and coherent manner. ## **Implementing Mask R-CNN** ### **Dataset Preparation** When building an instance segmentation model, your dataset must include pixel-level masks rather than just bounding boxes. Common choices include: - **COCO**: A large-scale [dataset](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation/) with 80 object categories and instance segmentation annotations - **Cityscapes**: Specialized [dataset](https://www.cityscapes-dataset.com/) for urban scenes with detailed annotations for traffic participants For smaller-scale projects or demonstrations, consider using a subset of these datasets. However, for production systems, comprehensive data that matches your target domain is essential. Instance segmentation demands precise labeling, as inaccurate boundaries will propagate through to your model's predictions. ### **Choosing a Framework** Several well-maintained libraries simplify Mask R-CNN implementation: - **Detectron2** (Facebook AI Research): Provides robust model implementations with various backbones (ResNet-50, ResNeXt, etc.) - **MMDetection** (OpenMMLab): Offers modular components and extensive configuration options For most applications, Detectron2 provides an excellent balance of performance and ease of use, with pre-trained models that can run inference with minimal setup. ### **Code Example: Pretrained Mask R-CNN Inference** Here's a concise example showing the core steps for running inference with a pre-trained Mask R-CNN model: ```python 1import cv2 2from detectron2.config import get_cfg 3from detectron2 import model_zoo 4from detectron2.engine import DefaultPredictor 5 6cfg = get_cfg() 7cfg.merge_from_file( 8 model_zoo.get_config_file("COCO-InstanceSegmentation/mask_rcnn_R_50_FPN_3x.yaml") 9) 10cfg.MODEL.WEIGHTS = model_zoo.get_checkpoint_url( 11 "COCO-InstanceSegmentation/mask_rcnn_R_50_FPN_3x.yaml" 12) 13cfg.MODEL.ROI_HEADS.SCORE_THRESH_TEST = 0.5 14cfg.MODEL.DEVICE = "cuda" 15predictor = DefaultPredictor(cfg) 16 17image_bgr = cv2.imread("example.jpg") 18image_rgb = cv2.cvtColor(image_bgr, cv2.COLOR_BGR2RGB) 19 20outputs = predictor(image_rgb) 21instances = outputs["instances"].to("cpu") 22 23boxes = instances.pred_boxes.tensor.numpy() 24scores = instances.scores.numpy() 25class_ids = instances.pred_classes.numpy() 26masks = instances.pred_masks.numpy() 27 ``` While using a pre-trained model works well for many applications, you can fine-tune Mask R-CNN on a custom dataset if your domain diverges significantly from standard benchmarks. This typically involves registering your dataset with the framework, adjusting hyperparameters, and potentially customizing the backbone architecture. ## **Leveraging FiftyOne for Mask R-CNN** Mask R-CNN provides accurate image segmentation, but FiftyOne takes a critical step further by enabling a data-centric approach to model development . FiftyOne is a tool that enables a data-centric approach to visual AI development, whether it’s fine-tuning Mask R-CNN or building and evaluating custom models. The App and Python library helps you visualize results, [evaluate performance](https://docs.voxel51.com/user_guide/evaluation.html), discover dataset issues, and iteratively refine your workflow. _The FiftyOne App displays ground truth segmentations and detections side by side, enabling data-centric exploration and iterative refinement._ ### **Dataset Integration** The following code shows how you can load a dataset into FiftyOne. This example loads a subset of the COCO dataset and its ground truth labels.FiftyOne seamlessly imports datasets in COCO format with a single command: ```python 1import fiftyone as fo 2 3dataset = fo.Dataset.from_dir( 4 dataset_dir="coco_small", 5 dataset_type=fo.types.COCODetectionDataset, 6 data_path="images", 7 labels_path="annotations/instances_val2017_50.json", 8 name="coco_val2017_50", 9 label_field="ground_truth_detections" 10) 11 ``` This automatically populates bounding boxes and instance segmentation data into a FiftyOne dataset. After running Mask R-CNN inference, predictions can be stored in a separate field, enabling direct comparison against ground truth. ![](https://cdn.sanity.io/images/h6toihm1/production/a70c92d151efabb9f708d75179cc605b0011ef2d-872x499.png?auto=format&dpr=2&fit=max&q=75&w=872) _The FiftyOne App displays ground truth segmentations and detections side by side, enabling data-centric exploration and iterative refinement._ ### **Visualizing and Exploring Data** FiftyOne's interactive App provides a powerful environment to: FiftyOne's interactive App provides a powerful environment to: - View individual images with toggleable label fields (e.g., switch between ground truth and predictions) ![](https://cdn.sanity.io/images/h6toihm1/production/b76977246106783a499307527aadb60901cbb606-874x721.png?auto=format&dpr=2&fit=max&q=75&w=874) _This view of the COCO dataset shows ground truth segmentations (purple) overlaid with Mask-RCNN model predictions (blue)._ - [Filter predictions by confidence](https://voxel51.com/blog/finding-the-optimal-confidence-threshold) to identify false positives ![](https://cdn.sanity.io/images/h6toihm1/production/50b31b3ef7228ebe74cd0f3f6a42ef0975419a23-874x721.png?auto=format&dpr=2&fit=max&q=75&w=874) _Here, the dataset is filtered to only show samples with a low prediction confidence threshold_ - Examine class distributions to check for imbalances ![](https://cdn.sanity.io/images/h6toihm1/production/ebbd18dff987a3d86e127c6ab9a7c7edb419f62e-797x621.png?auto=format&dpr=2&fit=max&q=75&w=797) _FiftyOne supports creating customized dashboards. Here, a categorical histogram shows the frequency of each ground truth class in the dataset._ - Zoom in on segmentation masks to inspect boundary precision ![](https://cdn.sanity.io/images/h6toihm1/production/a92a7366fc91c4f31c0883f77d0add2ff4abadb8-1841x963.png?auto=format&dpr=2&fit=max&q=75&w=1600) - Tag problematic samples for further review ```python ``` ![](https://cdn.sanity.io/images/h6toihm1/production/372c5b0f37e9a3b8d7fa3ffcdd197d3248c19ca8-1697x1091.png?auto=format&dpr=2&fit=max&q=75&w=1600) These capabilities transform model debugging from guesswork into systematic analysis. ### **Model Evaluation** For [instance segmentation](https://voxel51.com/resources/learn/a-guide-to-ai-image-segmentation/), rigorous evaluation is crucial. FiftyOne simplifies this process: ```python 1results = dataset.evaluate_detections( 2 "predictions", 3 gt_field="ground_truth", 4 eval_key="eval_masks", 5 use_masks=True, 6 compute_mAP=True 7) 8print("Mask mAP:", results.mAP()) 9 ``` Beyond aggregate metrics, FiftyOne stores per-sample evaluation results, enabling you to sort images by performance and focus on the most problematic cases. ### **Analyzing Results** Failure analysis is perhaps the most valuable component of a successful computer vision workflow. Through FiftyOne, you can: - Sort images by false positives or false negatives to immediately identify problem areas - Overlay ground truth masks and predicted masks to detect systematic errors - Group failures by object class, size, or occlusion levels to discover patterns - Perform error analysis on specific subsets to identify where your model struggles ![](https://cdn.sanity.io/images/h6toihm1/production/75220180a5718fccd0bf83e20775f01b5de5d514-600x338.gif?auto=format&dpr=2&fit=max&q=75&w=600)![](https://cdn.sanity.io/images/h6toihm1/production/6559404e68c08ffb57ec2cfd78ac39434aa2ea10-586x640.jpg?auto=format&dpr=2&fit=max&q=75&w=586) This targeted approach ensures that you invest your improvement efforts where they'll have the maximum impact. ## **Applications of Mask R-CNN** Mask R-CNN's ability to provide instance-level segmentation makes it valuable across numerous fields: ### **Autonomous Driving** In autonomous vehicles, detecting and precisely delineating other traffic participants is crucial for path planning and collision avoidance. Mask R-CNN excels at handling the complex and dynamic scenes encountered in urban environments, where pedestrians, vehicles, and obstacles frequently overlap in the vehicle's field of view. ![](https://cdn.sanity.io/images/h6toihm1/production/fec7b23838b23cb94088f2ba8618011646ca544c-1203x528.png?auto=format&dpr=2&fit=max&q=75&w=1203) ### **Robotics and Manufacturing** For robots operating in cluttered environments, distinguishing individual objects is essential for precise manipulation. Instance segmentation enables robots to identify specific items for picking, even when partially occluded by other objects. In manufacturing settings, Mask R-CNN can detect defects, verify component placement, and assess assembly quality. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ### **Medical Imaging** Medical applications demand extreme precision, making Mask R-CNN particularly valuable. The model can segment tumors, organs, or individual cells with high accuracy, supporting diagnosis, treatment planning, and research. Its ability to distinguish between multiple instances of the same class (such as individual cells) is especially relevant in histopathology. ### **Satellite and Aerial Imagery** When analyzing satellite imagery, separating individual buildings, vehicles, or land features is often necessary for tasks like urban planning, environmental monitoring, or traffic analysis. Mask R-CNN's instance segmentation capabilities provide the detailed delineation required for these applications. ## **Conclusion** Mask R-CNN represents a significant advancement in computer vision, bridging the gap between object detection and pixel-level segmentation. By leveraging region proposals, ROI Align, and multi-task learning, it achieves remarkable accuracy in delineating individual object instances, even in challenging scenarios with overlapping objects or complex boundaries. The combination of Mask R-CNN's sophisticated architecture with FiftyOne's data-centric workflow creates a powerful foundation for building robust instance segmentation solutions. Whether your application involves autonomous vehicles, medical imaging, robotics, or satellite imagery analysis, this approach allows you to not only implement state-of-the-art models but also understand their strengths and limitations in your specific domain. As computer vision continues to advance, instance segmentation will remain a cornerstone technology for applications requiring detailed scene understanding. By mastering Mask R-CNN and adopting data-centric practices with tools like FiftyOne, you'll be well-equipped to tackle these challenging visual perception tasks with confidence. **Image Citations** 1. Asaf antman. _Crowd at Noam Rotem concert._ Photograph. October 20, 2007. Wikimedia Commons. CC BY 2.0. [https://commons.wikimedia.org/wiki/File:Crowd\_at\_Noam\_Rotem\_concert.jpg](https://commons.wikimedia.org/wiki/File:Crowd_at_Noam_Rotem_concert.jpg). 2. Argenberg, Vyacheslav. _Kitchen, Tableware, Rostov-on-Don, Russia._ Photograph. January 19, 2014. Wikimedia Commons. CC BY 4.0. [https://commons.wikimedia.org/wiki/File:Kitchen,\_Tableware,\_Rostov-on-Don,\_Russia.jpg](https://commons.wikimedia.org/wiki/File:Kitchen,_Tableware,_Rostov-on-Don,_Russia.jpg). 3. Croasdell, Victoria Lee. _MRISAR Hand Crafted Three Finger Robotic Arm-2._ Photograph. October 24, 2018. Wikimedia Commons. CC BY-SA 4.0. [https://commons.wikimedia.org/wiki/File:MRISAR\_hand\_crafted\_three\_finger\_robotic\_arm-2.jpg](https://commons.wikimedia.org/wiki/File:MRISAR_hand_crafted_three_finger_robotic_arm-2.jpg). 4. NOMAD. _Trafficjamdelhi._ Photograph. (Uploaded January 1, 2008). Wikimedia Commons. CC BY 2.0. [https://commons.wikimedia.org/wiki/File:Trafficjamdelhi.jpg](https://commons.wikimedia.org/wiki/File:Trafficjamdelhi.jpg). 5. Halicki, Jacek. _2023 Pluszowy miś_. Photograph. June 5, 2023. Wikimedia Commons. CC BY-SA 4.0. [https://commons.wikimedia.org/wiki/File:2023\_Pluszowy\_mi%C5%9B.jpg](https://commons.wikimedia.org/wiki/File:2023_Pluszowy_mi%C5%9B.jpg). 6. Miguel Chevalier. _Body Voxels – The Walker._ 2013\. Photograph. Wikimedia Commons. CC BY-SA 4.0. [https://commons.wikimedia.org/wiki/File:Body\_Voxels\_-\_The\_Walker,\_Miguel\_Chevalier,\_2013.jpg](https://commons.wikimedia.org/wiki/File:Body_Voxels_-_The_Walker,_Miguel_Chevalier,_2013.jpg). [object detection](https://voxel51.com/blog/tag/object-detection) [object segmentation](https://voxel51.com/blog/tag/object-segmentation) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/45b6b66f2f7c3ba6e83db74c270be16d8087ee94-2344x1306.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Comprehensive Guide to Keypoint Detection for Object Recognition\\ \\ Learn\\ \\ • \\ \\ May 5, 2025](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition) [![](https://cdn.sanity.io/images/h6toihm1/production/ca4f84addae3f8e97daad00cf856754301cbb5e0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Are Image Segmentation Maps Superior to Bounding Boxes?\\ \\ Learn\\ \\ • \\ \\ Feb 26, 2025](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes) [![](https://cdn.sanity.io/images/h6toihm1/production/53c3b307e202d573a6fccf19494ba35c43a3dc7e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Best Practices for Evaluating AI Models Accurately\\ \\ Learn\\ \\ • \\ \\ Dec 17, 2024](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-19-lllmstxt|> ## Visual AI in Healthcare [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) Visual AI in Healthcare: 2025 Landscape May 12, 2025 • 9 min read Article content In this article [State-of-the-Art Models & Datasets](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#0cd730b82958) [Unsolved Problems in Visual AI for Healthcare](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#366595b598fb) [Leading Companies](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#b67a4b262bfb) [Influential Researchers](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#32fbd9b6509e) [The Future of AI-Assisted Diagnosis](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#a7dccdce5496) [Frequently Asked Questions](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#03ffd35d6dce) [Just wrapping up!](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#741a8979df90) [Stay Connected:](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#e805f80c6365) [What is next? Join Us: Meetups & Workshops](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#875526b0f759) [Author’s note](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#096748612c71) In this article [State-of-the-Art Models & Datasets](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#0cd730b82958) [Unsolved Problems in Visual AI for Healthcare](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#366595b598fb) [Leading Companies](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#b67a4b262bfb) [Influential Researchers](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#32fbd9b6509e) [The Future of AI-Assisted Diagnosis](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#a7dccdce5496) [Frequently Asked Questions](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#03ffd35d6dce) [Just wrapping up!](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#741a8979df90) [Stay Connected:](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#e805f80c6365) [What is next? Join Us: Meetups & Workshops](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#875526b0f759) [Author’s note](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape#096748612c71) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) A radiologist once told me that what keeps her up at night isn’t missing a tumor, it’s missing the story behind it. The patient who waited too long, the early sign hidden in a sea of images, the delays that cost precious time. This is where Visual AI can help, not by replacing people, but by supporting them with tools that bring speed, clarity, and focus. Visual AI is changing healthcare in real and powerful ways. It’s helping doctors find problems sooner, plan treatments more precisely, and use their time better. By combining machine learning with medical images, we’re solving problems that have challenged healthcare for years. But as we build these tools, we need to ask an important question: The real challenge isn’t building more innovative tools, it’s making sure they bring us closer to people, not further from what makes care human. This work is not just about building innovative systems. It’s about keeping compassion, trust, and experience at the center of it all. ## State-of-the-Art Models & Datasets AI in medical imaging is advancing quickly and is powered by high-quality datasets and cutting-edge models. From vision-language tools to specialized segmentation networks, these models are pushing the limits of diagnosis, prediction, and clinical decision support. Remember that Data is not just code and pixels; it’s people’s lives, diagnoses, and treatments. How do we steward that responsibility? Below is a curated list of standout contributions shaping the next generation of intelligent, patient-centered healthcare. ### Leading Models - **GMAI-VL**: The GMAI-VL (General Medical AI Vision-Language) model is a cutting-edge multimodal AI system that integrates medical imaging with natural language understanding. It was introduced alongside the GMAI-VL-5.5M dataset, which comprises 5.5 million meticulously curated image-text pairs derived from various medical datasets. _\[ [Paper](https://arxiv.org/abs/2411.14522)\]\[ [Repo](https://github.com/uni-medical/GMAI-VL)\]_ - **LlaVa-Med**: A multimodal AI model that combines computer vision and natural language processing (NLP), designed specifically for medical imaging interpretation and clinical reasoning. Built on the LLaVA framework (originally for general vision-language tasks), LLaVA-Med integrates: 1) **Visual encoders** (like CLIP) to process medical images (e.g., X-rays, CT scans, pathology slides), and 2) **Language models** (e.g., LLaMA or similar) to generate medical explanations, diagnoses, or reports. _\[ [Paper](https://arxiv.org/abs/2306.00890)\]\[ [Repo](https://github.com/microsoft/LLaVA-Med?tab=readme-ov-file)\]\[ [Model Weights](https://huggingface.co/microsoft/llava-med-v1.5-mistral-7b)\]_ - **CHIEF**: Clinical Histopathology Imaging Evaluation Foundation (CHIEF) is a groundbreaking AI system that researchers at Harvard Medical School developed. Designed to revolutionize cancer diagnostics, CHIEF integrates advanced machine learning techniques to analyze histopathological images, enabling accurate cancer detection, prognosis prediction, and treatment guidance across multiple cancer types. _\[ [Paper](https://www.nature.com/articles/s41586-024-07894-z)\]\[ [Repo](https://github.com/hms-dbmi/CHIEF?tab=readme-ov-file)\]\[ [Model Weights](https://hub.docker.com/r/chiefcontainer/chief/)\]_ - **BioMedCLIP**: A domain-specific adaptation of OpenAI’s CLIP architecture, tailored for biomedical applications. It leverages a vast dataset of image-text pairs to bridge the gap between visual and textual biomedical data, facilitating tasks like image classification, retrieval, and visual question answering. _\[ [Paper](https://arxiv.org/abs/2303.00915)\]\[ [Model Weights](https://huggingface.co/microsoft/BiomedCLIP-PubMedBERT_256-vit_base_patch16_224)\]_ - **SAM-VMNet**: A hybrid architecture that combines the Segment Anything Model (SAM) with VM-UNet, a vision-based medical network. This integration leverages SAM’s robust feature extraction and VM-UNet’s efficient processing capabilities to enhance segmentation accuracy and speed in coronary angiography images. The model achieved a segmentation accuracy of up to 98.32% and sensitivity of 99.33%, outperforming existing models in this domain. _\[ [Paper](https://arxiv.org/abs/2406.00492)\]\[ [Repo](https://github.com/qimingfan10/SAM-VMNet)\]_ - **MedSAM2:** A promptable foundation model for 3D medical image and video segmentation. Built upon the Segment Anything Model 2 (SAM2), it was fine-tuned on a large medical dataset comprising over 455,000 3D image-mask pairs and 76,000 annotated video frames. MedSAM2 introduces a memory attention mechanism to handle temporal information, enabling efficient and accurate segmentation across various organs, lesions, and imaging modalities. It also significantly reduces manual annotation efforts by over 85%. _\[ [Paper](https://arxiv.org/abs/2504.03600)\]\[ [Repo](https://github.com/SuperMedIntel/Medical-SAM2)\]\[ [Model Weights](https://huggingface.co/wanglab/MedSAM2)\]_ ### Key Datasets - **MedTrinity-25M**: 25M images over 10 modalities, 65+ diseases. UC Santa Cruz, Stanford, and Harvard. This dataset introduces a novel automated pipeline that generates multigranular annotations, encompassing both global textual information (e.g., disease type, modality, region-specific descriptions) and detailed local annotations for regions of interest (ROIs), such as bounding boxes and segmentation masks. _\[ [Paper](https://arxiv.org/abs/2408.02900)\]\[ [Repo](https://github.com/UCSC-VLAA/MedTrinity-25M)\]\[ [Project](https://yunfeixie233.github.io/MedTrinity-25M/)\]_ - **GMAI-VL-5.5M**: Rich VLM training set of 5.5M image-text pairs. Developed to address the limitations of general AI models in medical applications, this dataset facilitates the training of vision-language models (VLMs) capable of accurate diagnoses and clinical decision-making. _\[ [Paper](https://arxiv.org/abs/2411.14522)\]\[ [Repo](https://github.com/uni-medical/GMAI-VL)\]_ - **MedPix 2.0**: A comprehensive multimodal biomedical dataset designed to advance AI applications in the medical domain. It builds upon the original MedPix® archive, widely used for medical education, by introducing structured data suitable for training and evaluating multimodal AI models. _\[ [Paper](https://arxiv.org/abs/2407.02994)\]\[ [Repo](https://github.com/CHILab1/MedPix-2.0)\]_ - **ARCADE**: The ARCADE dataset is a publicly available benchmark designed to facilitate the development and evaluation of automated methods for coronary artery disease (CAD) diagnostics. Introduced as part of the ARCADE challenge at the 26th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), it provides expert-labeled X-ray coronary angiography (XCA) images, enabling researchers to develop and assess deep learning models for vessel segmentation and stenosis detection. _\[ [Related Paper](https://arxiv.org/html/2503.01601v1)\]\[ [ARCADE Challenge](https://arcade.grand-challenge.org/)\]\[ [Repo](https://github.com/cmctec/ARCADE)\]_ - **DeepLesion**: A large-scale, publicly available dataset comprising over 32,000 annotated lesions identified on CT images. Developed by the National Institutes of Health (NIH) Clinical Center, it aims to facilitate the development of computer-aided detection (CADe) and diagnosis (CADx) systems by providing a diverse set of lesion annotations across various body parts. _\[ [Paper](https://pubmed.ncbi.nlm.nih.gov/30035154/)\]\[ [Dataset](https://nihcc.app.box.com/v/DeepLesion/folder/50715173939)\]\[ [Repo](https://github.com/ComputationalImageAnalysisLab/DeepLesionData)\]_ ## **Unsolved Problems in Visual AI for Healthcare** But hold your horses. Despite rapid progress, key challenges remain, from data privacy and trust to making AI fit naturally into clinical workflows. Solving these isn’t just technical; it’s about building tools that earn their place in human care. AI promises precision, but its success depends on trust. Transparency isn’t optional; it’s critical. Here is the list of the unsolved problems in Visual AI for Healthcare. - **Data Privacy**: Securely handling sensitive medical images. - **Model Interpretability**: Ensuring clinicians trust and understand AI insights. - **Generalization**: Avoiding performance drops across devices and demographics. - **Workflow Integration**: Embedding AI without disrupting care routines. ## Leading Companies Behind every breakthrough model is the question: Who brings it to life in the real world? The shift from research to real impact depends on those who turn algorithms into tools, embed them into clinical settings, and prove their value at scale. This is where AI meets care. - **Aidoc**: Radiology triage and anomaly detection. _\[ [Webpage](https://www.aidoc.com/)\]_ - **Ascertain**: AI agents for hospital operations. _\[ [Webpage](https://www.ascertain.com/)\]_ - **Tempus**: AI-powered clinical + genomic data analysis. _\[ [Webpage](https://www.tempus.com/)\]_ - **GE Healthcare**: AI-enhanced imaging systems. _\[ [Webpage](https://www.gehealthcare.com/products/magnetic-resonance-imaging/effortless-imaging-ai-mri-solutions?srsltid=AfmBOorc180nzqNMP8dZ2WbFeCDYwd8iFvlalKKzVprIGJ0BnwXe2PU-)\]_ - **HeartFlow**: Non-invasive CAD analysis using AI + CCTA. [_\[Webpage\]_](https://www.heartflow.com/heartflow-one/ffrct-analysis/?ss=google&st=ct-ffr&utm_campaign=1463060622&utm_source=google&utm_medium=cpc&utm_term=ct-ffr&utm_content=adgroupid%3A57063354216+creative%3A426162832000+matchtype%3Ab+network%3Ag+device%3Ac+position%3A+placement%3A&hsa_acc=1403132650&hsa_cam=1463060622&hsa_grp=57063354216&hsa_ad=426162832000&hsa_src=g&hsa_tgt=kwd-612990568538&hsa_kw=ct-ffr&hsa_mt=b&hsa_net=adwords&hsa_ver=3&gad_source=1&gad_campaignid=1463060622&gbraid=0AAAAAC_bAMcIcvI-xb7wyLDl_Jc1imYMv&gclid=CjwKCAjwz_bABhAGEiwAm-P8YaH2-xc2CJztL1MFxZVsbYwSHT0jRG8tOSqLnNZGkrRFe2zLArySnhoC6NwQAvD_BwE) - **IQVIA**: Big data + AI for research and outcomes. [_\[Webpage\]_](https://www.iqvia.com/) ## Influential Researchers An AI influencer is a thought leader who drives awareness, understanding, and responsible use of artificial intelligence. Innovation in healthcare AI is driven by people who ask bold questions, build new methods, and lead with purpose. These researchers are advancing the science and shaping how it reaches and serves patients. - **Dr. Mihaela van der Schaar**: ML for healthcare decision-making. _\[ [LinkedIn](https://www.linkedin.com/in/mihaela-van-der-schaar/)\]\[ [Lab](https://www.vanderschaar-lab.com/)\]_. - **Dr. Mark Michalski**: CEO, Ascertain. \[ [_LinkedIn_](https://www.linkedin.com/in/mark-michalski-b3b23411/)\] - **Dr. Taha Kass-Hout**: GE Healthcare, Chief Medical Officer. \[ [_LinkedIn_](https://www.linkedin.com/in/tahak/)\] - **Dr. Hugo Aerts**: Professor @ Harvard \| Director, Artificial Intelligence in Medicine. \[ [_LinkedIn_](https://www.linkedin.com/in/hugoaerts/)\] - **Dr. Sanjay Rajagopalan**: Cardiovascular imaging leader. \[ [_LinkedIn_](https://www.linkedin.com/in/sanjay-rajagopalan-md-mba-810983205/)\] - **Dr. Heather Couture:** Consultant, Researcher, Writer & Host of Impact AI Podcast. _\[ [LinkedIn](https://www.linkedin.com/in/hdcouture/)\]_ ## The Future of AI-Assisted Diagnosis In the future of healthcare, doctors won’t be working alone. **Visual AI** will stand beside them, spotting patterns, flagging concerns, and offering real-time insights. It’s an intelligent companion, not a replacement. The judgment, empathy, and final decisions will always rest with the people who care. Together, humans and machines will deliver more accurate, personalized, and human care. ## **Frequently Asked Questions** **What are the ethical considerations in deploying AI in healthcare diagnostics?** AI in healthcare raises serious ethical questions. Patient data must be protected with strict privacy measures and clear consent. Bias is another concern; models trained on limited data can lead to unfair results. Doctors also need to understand how AI makes decisions, which means transparency tools are key. Finally, it’s about shared responsibility: AI can support decisions, but humans must stay in charge. **How does AI integration affect the workflow of radiologists?** AI is making radiologists’ jobs more efficient. It can quickly screen images, flag issues, and suggest next steps, helping prioritize urgent cases and reduce burnout. But success depends on smooth integration. If AI tools disrupt workflows or add friction, they won’t be used. The best systems support, not slow down, clinical work. **What measures ensure the accuracy of AI-generated medical diagnoses?** Accuracy comes from testing AI models on outside datasets, not just the training data. Many go through regulatory checks from the FDA or similar bodies. Clinicians still review AI outputs — it’s a team effort. After deployment, systems are monitored for performance shifts, and explainability tools help doctors understand what the AI sees and why. ## **Just wrapping up!** With Visual AI in healthcare, we can run faster diagnostics or better models, but we always need to remind ourselves how we care for people in a world shaped by data, complexity, and urgent needs. We require more than code and computation; we demand collaboration between engineers and doctors, with a solid commitment to making technology work for everyone. The future is already unfolding. **The real question is:** How do we shape it with care? Please Share Your Thoughts, Ask Questions, and Provide Testimonials. Your insights might help others in our next posts. Don’t forget to participate in the challenge and try out the notebook I have created for you all. _Together, we can innovate in action recognition and make meaningful contributions to AI for Good. Let’s build something impactful!_ ## Stay Connected: - **Follow me on Medium:** [https://medium.com/@paularamos\_phd](https://medium.com/@paularamos_phd) - **Follow Me on LinkedIn:** [https://www.linkedin.com/in/paula-ramos-phd/](https://www.linkedin.com/in/paula-ramos-phd/) - **Join the Conversation:** [Discord Fiftyone-community](https://discord.gg/fiftyone-community) # What is next? Join Us: Meetups & Workshops We invite you to attend our upcoming [**Meetups**](https://voxel51.com/events), where we will discuss real-world AI in healthcare, beyond noise trends. Hear from experts, share your perspective, and connect with a growing community. And don’t miss our **“Getting Started with Visual AI in Healthcare” Workshop**! , July 17th — Link TBD, follow this [meetup.com](https://www.meetup.com/pro/ai-machine-learning-data-science-network/) for updates - Use [**FiftyOne**](https://docs.voxel51.com/) to explore datasets like **ARCADE** and **DeepLesion.** - Work hands-on with models like **BioMedCLIP**, **MedSAM2**, and **SAM-VMNet.** - Learn to calculate embeddings, visualize results, and surface key medical insights. ## Author’s note _I want to acknowledge that I am not a healthcare professional. I write this piece as an observer and researcher, drawn to the powerful intersection of artificial intelligence and healthcare. The technologies discussed here represent significant progress and_ _**c** omplex, high-stakes challenges._ _In my view, the most essential principle is simple: we must remain responsible. These tools are not just lines of code, they interact with real lives, real diagnoses, and real decisions. We need to ensure humans stay in the loop, bringing context, compassion, and judgment to every AI-assisted step. The goal isn’t to replace clinicians, but to empower them._ _Let’s move forward with curiosity, courage, and care._ [artificial intelligence](https://voxel51.com/blog/tag/artificial-intelligence) [healthcare](https://voxel51.com/blog/tag/healthcare) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [machine learning](https://voxel51.com/blog/tag/machine-learning) ![](https://cdn.sanity.io/images/h6toihm1/production/e926c07c7d1426c0fde8fdefa637c528d47b16f4-512x512.webp?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/ebd118c3d8181dcba0532200d024759e23c7ec1e-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Computer vision in healthcare: 12 breakthrough case studies\\ \\ Industry Solutions\\ \\ • \\ \\ Jul 15, 2025](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies) [![](https://cdn.sanity.io/images/h6toihm1/production/6445eab4dfdba3548381f187231e576c97c3c5a5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Computer Vision Is Changing Healthcare\\ \\ Industry Solutions, Product & News\\ \\ • \\ \\ Aug 31, 2023](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare) [![](https://cdn.sanity.io/images/h6toihm1/production/3a661345dfbc596f7118a7b8ec8375ea1f73d138-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 3\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-20-lllmstxt|> ## Voxel51 at CVPR 2024 [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Voxel51 at CVPR 2024! Jun 10, 2024 • 5 min read Article content In this article [Visit Voxel51 booth #1519 to …](https://voxel51.com/blog/voxel51-at-cvpr-2024#4a274017c965) [Check out these two workshops with Voxel51 Chief Scientist Jason Corso](https://voxel51.com/blog/voxel51-at-cvpr-2024#4b14d9e71f5b) [Learn about embeddings and 3D in a flash](https://voxel51.com/blog/voxel51-at-cvpr-2024#5212f7dcf3da) In this article [Visit Voxel51 booth #1519 to …](https://voxel51.com/blog/voxel51-at-cvpr-2024#4a274017c965) [Check out these two workshops with Voxel51 Chief Scientist Jason Corso](https://voxel51.com/blog/voxel51-at-cvpr-2024#4b14d9e71f5b) [Learn about embeddings and 3D in a flash](https://voxel51.com/blog/voxel51-at-cvpr-2024#5212f7dcf3da) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) CVPR 2024 is next week! Here’s when and where you can find members of the Voxel51 team. We encourage you to put a visit to the Voxel51 team on your CVPR 2024 agenda – we'd love to meet you! But first, you’re in for a special treat this week. It’s the CVPR Preshow! Hacker-in-Residence Harpreet Sahota and Chief Scientist Jason Corso interviewed authors who presented research in vision-language models, 3D computer vision, and diffusion models. Watch these interviews to get a deeper insight into what type of bleeding-edge research is being presented at CVPR 2024 Follow [Harpreet on LinkedIn](https://www.linkedin.com/in/harpreetsahota204/) to be among the first to see preshow interviews as they are posted. ## Visit Voxel51 booth \#1519 to … ### See the mind-blowing open source FiftyOne project in action If you're new to Voxel51, we're developers of the [open source FiftyOne](https://github.com/voxel51/fiftyone) project that helps visual AI builders improve data quality and boost model performance. We love showing the mind-blowing workflows made easy with FiftyOne. Stop by booth #1519 anytime and see, for example, how to use FiftyOne to: - Easily curate and manage your data - Evaluate models - Visualize embeddings - Find annotation mistakes - Find image quality issues - Reverse image search - Zero-shot prediction - Mine hard samples - Work with multi-modal data (3D, images, videos, more!) Or, tell us what you’re working on and we’ll show you how FiftyOne can help! If you've already visited us at CVPR, no fear -- there are lots of new goodies we think you’ll like. In fact, since last year's CVPR, there have been [20 new releases](https://docs.voxel51.com/release-notes.html) of open source FiftyOne, including the [latest 0.24.0 release](https://voxel51.com/blog/announcing-fiftyone-0-24-with-3d-meshes-and-custom-workspaces/) with 3D meshes, custom workspaces, and more! ### Score some coveted swag We're confident you'll walk away knowing how FiftyOne can instantly alleviate some of the pain points you're facing in your visual AI projects, but we're also giving away some sweet swag that can be yours while supplies last. ### Learn about our open roles Voxel51 is growing quickly and we are looking for people to grow with us. Visit our recruiting desk in our booth to learn more about our open roles (below) and meet members of the team to find out what it’s like working here. - [Machine Learning Customer Success Engineer](https://voxel51.com/jd/?4367259005?gh_jid=4367259005) - [Machine Learning Developer Evangelist](https://voxel51.com/jd/?4067392005?gh_jid=4067392005) - [Machine Learning Engineer](https://voxel51.com/jd/?4404907005?gh_jid=4404907005) - [Machine Learning Scientist](https://voxel51.com/jd/?4405009005?gh_jid=4405009005) - [Product Manager](https://voxel51.com/jd/?4004582005?gh_jid=4004582005) - [Product Marketing Manager](https://voxel51.com/jd/?4409782005?gh_jid=4409782005) - [Software Engineer](https://voxel51.com/jd/?4400468005?gh_jid=4400468005) ### Say hi and let us know what you’re working on! We’re AI/ML builders and enthusiasts, too, and we love hearing what fellow members of the community are working on. Stop by and let us know what you’re up to – your latest research, AI apps you’re building, datasets you’re building, models you’re fine-tuning, or whatever’s on your mind. ## Check out these two workshops with Voxel51 Chief Scientist Jason Corso [Jason Corso](https://web.eecs.umich.edu/~jjcorso/pubs/), Professor of Robotics, Electrical Engineering, and Computer Science at the University of Michigan and Co-Founder / Chief Science Officer of Voxel51, will speak in two workshops this year at CVPR. Have a look at the details and we hope to see you there! **Computer Vision with Humans in the Loop (CVHL)** - Date: Tuesday, June 18, 2024 - Time: 8:30 AM **–** 5:45 PM - Room: Room Summit 329 - More Details: [https://cvhl.org/](https://cvhl.org/) Summary: This workshop aims to explore the pivotal role of human interaction in advancing computer vision technologies. Despite significant progress over the past two decades, computer vision systems often fall short of human capabilities, especially in specialized or complex scenarios. Featuring a series of invited talks and panels, this workshop will highlight innovations like interactive segmentation, advancements in language and visual prompt integration, and the development of self-aware systems capable of recognizing and compensating for their limitations. Distinguished speakers from both academia and industry will share their insights, offering a rich dialogue on the integration of human cognitive skills with machine precision to tackle the challenges of computer vision. Participants will gain an understanding of the evolving landscape of human-in-the-loop methodologies and their impact on practical applications and foundational models in the field. Join us to discuss historical perspectives, current research, and future directions in this dynamic area. **Visual Odometry and Computer Vision Applications Based on Location Clues** - Date: Tuesday, June 18, 2024 - Time: 8:30 AM–5:30 PM - Location: Summit 330 - More Details: [https://sites.google.com/view/vocvalc2024/](https://sites.google.com/view/vocvalc2024/) Summary: Visual odometry and localization have maintained an increasing interest in recent years, especially with the extensive applications for autonomous driving, augmented reality, and mobile computing. With the location information obtained through odometry, services based on location clues are also rapidly emerging. Particularly, in this workshop, we focus on mobile and robot platform applications. This workshop invites papers in the areas including advances in visual odometry and computer vision applications based on location context. ## Learn about embeddings and 3D in a flash We signed up to present two Flash Sessions! Flash sessions are 15-minute TED-style talks where you can learn a lot about computer vision techniques in a short amount of time. Check out our two sessions, in the Expo Hall on stage at booth 1841. **Thursday, June 20, 3:00-3:15 pm: 5 Handy Ways to Use Embeddings, the Swiss Army Knife of AI** Presented by Hacker-in-Residence Harpreet Sahota Discover the incredible potential of vector search engines beyond RAG for large language models! Explore 5 handy embeddings applications: robust OCR document search, cross-modal retrieval, probing perceptual similarity, comparing model representations, concept interpolation, and a bonus—concept space traversal. Sharpen your data understanding and interaction with embeddings and open source FiftyOne. **Friday, June 21, 12:00-12:15 pm:Build Your Own Virtual World: 3D Reconstruction in FiftyOne** Presented by ML Engineer Daniel Gural 3D is one of the fastest-growing spaces in ML, and new models are coming out that can achieve incredible results. In this talk, you'll learn some of the methods used to create 3D reconstructions, the drawbacks of today's models, and what there is to be excited about on the horizon. [CVPR](https://voxel51.com/blog/tag/cvpr) [CVPR 2024](https://voxel51.com/blog/tag/cvpr-2024) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/5d9fe483cd6c19e4ee6246bac487e4716fe99fe7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ 5 Reasons to Visit Voxel51 at CVPR\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jun 9, 2023](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr) [![](https://cdn.sanity.io/images/h6toihm1/production/e60eea36edcff16d65c62e2c3fea99a66d36a8d5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 @CVPR 2025: Smarter, Faster Visual AI\\ \\ Event Recaps, Product & News\\ \\ • \\ \\ Jun 3, 2025](https://voxel51.com/blog/cvpr-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/22ddad88a69b3782a530e73d0ec66f5debd8d516-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Forsight Finds a Centralized Dataset Management Solution in FiftyOne Teams\\ \\ Product & News\\ \\ • \\ \\ Nov 2, 2022](https://voxel51.com/blog/forsight-finds-a-centralized-dataset-management-solution-in-fiftyone-teams) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-21-lllmstxt|> ## Top Data Annotation Services [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) 15 Best Data Annotation Companies & Data Labeling Services May 19, 2025 • 10 min read Article content In this article [15 Best Data Annotation Companies & Data Labeling Services](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#af8a9830dbac) [What do data annotation companies do?](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#a3a306a46041) [What do data annotation services offer?](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#e236bf677735) [Methodology](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#987e9ed16491) [How to choose the right AI data annotation company for you](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#9781bea071c8) [Top 15 data annotation and labeling vendors](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#94a840c885ff) [Getting Started — What Matters Most](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#bea0e31dc2b0) In this article [15 Best Data Annotation Companies & Data Labeling Services](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#af8a9830dbac) [What do data annotation companies do?](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#a3a306a46041) [What do data annotation services offer?](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#e236bf677735) [Methodology](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#987e9ed16491) [How to choose the right AI data annotation company for you](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#9781bea071c8) [Top 15 data annotation and labeling vendors](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#94a840c885ff) [Getting Started — What Matters Most](https://voxel51.com/blog/15-best-data-annotation-companies-data-labeling-services#bea0e31dc2b0) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) # 15 Best Data Annotation Companies & Data Labeling Services Accurate labels on data are critical for training performant models. Yet building large, high-quality training sets in-house can drain budgets and delay releases. This guide compares fifteen of the best AI data-annotation companies and data-labeling services to help you shortlist both classic human-in-the-loop vendors and modern automated platforms. [Image source](https://en.wikipedia.org/wiki/Object_detection#/media/File:Detected-with-YOLO--Schreibtisch-mit-Objekten.jpg) ## **What do data annotation companies do?** Annotation vendors recruit trained workforces or build specialized ML pipelines to add structure to unstructured data. Tasks include drawing bounding boxes on images, segmenting objects in video, transcribing audio, and classifying text. Modern providers pair human expertise with model-assisted pre-labels and automated quality checks to deliver faster without sacrificing accuracy. ## **What do data annotation services offer?** Annotation vendors typically offer the following mix of products and services. - **Tooling**: web or API interfaces for annotators and reviewers - **Workforce management**: vetting, training, and routing tasks to labelers - **Quality control**: multi-tier review, consensus checks, or statistical audits of labeled data - **Security and compliance**: ISO, SOC 2, HIPAA, or FedRAMP where needed - **ML-assisted workflows**: pre-labels from foundation models, active-learning loops, or automated QA ## Methodology This article summarizes publicly available documentation of tooling depth, turnaround time, workforce scale, and security posture. Preference went to platforms with ML-assisted workflows, transparent pricing, and proven compliance support. ## How to choose the right AI data annotation company for you Before choosing an annotation product or service, consider your organization's needs around data modalities, velocity, and security/regulatory requirements. - **Supported data types**: image, video, 3-D, text, audio, or mixed? - **Volume and turnaround**: can the service deliver hours of video per week, or thousands of X-ray frames overnight? - **Quality targets**: are precision/recall or inter-annotator agreement metrics transparent? - **Security requirements**: are PII considerations, geo-fencing, or on-prem processing needed? - **Pricing model**: per-label, per-hour, or usage-based? Match your answers to each provider’s strengths and service-level commitments. ## Top 15 data annotation and labeling vendors ### Voxel51 **Voxel51** centers its offerings on its [FiftyOne](https://voxel51.com/fiftyone/) toolkit, a Python-based workbench that lets machine-learning engineers build AI applications from massive image and video datasets. Its auto-labeling workflows pair zero-shot model predictions with metrics like Reconstruction Error Ratios (RERs) to surface only the handful of low-confidence samples that still need a human touch, trimming annotation spend by up to 75 percent. Because the platform is code-first, teams can integrate directly into their MLOps stack, automate QA in CI pipelines, and export labels in any schema they choose. However, users must bring their own foundation-model checkpoint (or rely on off-the-shelf models) so non-technical teams may face a steeper learning curve. Support today is focused on visual data and audio; text and other modalities will need a second tool. For companies that value data-centric debugging and want to keep IP on-prem or in their own cloud, Voxel51 delivers a flexible, engineering-friendly solution. ### CloudFactory **CloudFactory** blends a managed workforce with active-learning tooling to label visual data at scale. Clients get a named account team and SLA-backed turnarounds, which are key for safety-critical verticals such as autonomous driving and precision agriculture. The company’s auditor model means at least two sets of eyes review every label. Per-hour pricing simplifies budget forecasts for long-running projects, yet it can exceed usage-based competitors once workflows stabilize and automation kicks in. Onboarding niche ontologies may require bespoke guideline creation, which lengthens ramp-up. ### Hive **Hive** taps global crowdsourcing and large proprietary models to deliver bounding boxes, polygons, OCR, and LiDAR annotations. Its API-first platform lets customers submit data and receive labels, often within hours, making it popular among social-media firms, ad platforms, and short-form-video apps that ingest millions of images per day. Built-in pre-labeling bootstraps accuracy, and a consensus algorithm flags disagreements for secondary review. The trade-off is opacity: Hive largely remains a black box, and offers limited visibility into annotator training or error-analysis pipelines. Custom label schemas and complex 3-D use cases force teams to adapt requirements to Hive’s templates. ### AWS SageMaker Ground Truth Plus **AWS SageMaker Ground Truth Plus** delivers labeling as a managed service that plugs into S3 buckets and SageMaker pipelines, so data never leaves the Amazon ecosystem. Active-learning loops route high-confidence predictions past humans, and AWS claims up to 40% lower costs for common vision workflows like object detection and image classification. The downside is vendor lock-in: JSON manifests, IAM roles, and CloudWatch metrics are all AWS-centric, making migration painful. Ground Truth Plus also masks its human workforce behind service endpoints, limiting transparency into annotator expertise. For organizations already “all-in” on AWS, however, the convenience and native security controls are compelling. ### Appen **Appen** boasts one of the world’s largest multilingual crowds, spanning 235+ dialects and specialized in speech, text, and conversational AI. Customers can purchase turnkey datasets or engage the workforce for custom annotation, sentiment scoring, and search-relevance tasks. The company backs projects with ISO 9001 quality processes and a suite of reviewer dashboards that measure inter-annotator agreement. Yet quality can drift on large jobs unless you pay for Appen’s premium multi-stage QC, and smaller engagements often find pricing opaque. Ramp-up times stretch if task guidelines or language expertise are niche. ### SuperAnnotate **SuperAnnotate** offers a browser-based IDE with dataset versioning, branching, and diff views for teams iterating rapidly on vision models. Built-in automation suggests polygon masks and tracking boxes for video frames, and an optional marketplace of vetted vendors can absorb overflow work. The platform’s analytics panel shows accuracy and reviewer throughput, which help managers calibrate effort versus impact. However, per-frame pricing climbs quickly for hi-res footage, and large jobs still require external workforce contracts that SuperAnnotate merely brokers. Text and audio support lag behind vision features. ### Cogito **Cogito** specializes in regulated verticals like healthcare, insurance, and finance. The company maintains secure facilities, geo-fenced data centers, and a vetted medical workforce that can annotate DICOM images. Every label passes through double-blind review plus automated rule checks to maximize precision for diagnostic models. That rigor, however, extends turnaround times and elevates per-image costs to sometimes double a generalist vendor. Cogito’s tooling is purpose-built for visual medical data, so teams operating in other verticals may need another partner. ### Encord **Encord** couples a timeline-based labeling UI with a Python SDK that lets engineers run active-learning loops and push predictions back to annotators. Its ontology versioning is designed to keep teams operating in domains like large robotics or autonomous-driving in sync. The platform is vision-centric: text, tabular, or audio tasks feel tacked on, and the marketplace of third-party labelers is smaller than those of longer-established rivals. Pricing is usage-based but skews higher for sensor-fusion modalities like LiDAR. ### Labelbox **Labelbox** positions itself as an end-to-end “data engine" that blends labeling, curation, and model-error analysis. Model-in-the-loop pre-labels are designed to speed up bounding-box and segmentation tasks, while also surfacing sparse classes or failure slices for re-training. A talent marketplace gives customers on-demand access to vetted labeling partners, but enterprise-grade seats and advanced QA workflows sit behind premium SKUs. Teams get granular analytics, yet long-term storage fees can surprise if you park raw media indefinitely. Because Labelbox is cloud-only, on-prem or air-gapped deployments aren’t supported. ### Kili Technology **Kili Technology** prioritizes security, holding ISO 27001 and SOC 2 Type II certifications and offering on-prem or single-tenant cloud deployments for classified projects. Its UI supports image, video, text, and PDF annotations. Collaborative dashboards display reviewer consensus in real time to spot drift. Public pricing is scarce, and API support is thinner than that of bigger rivals, so integration work often falls on internal engineers. Kili’s talent pool is smaller, which can slow ramp-up on very large, multilingual jobs. ### SuperbAI **SuperbAI** differentiates with its few-shot _Custom Auto-Label_ feature: you upload a couple thousand seed annotations and the system retrains a lightweight detector that improves with each review cycle. A Bayesian uncertainty engine highlights low-confidence predictions for targeted cleanup. I features a modern WebUI along with granular role-based access control. Because the engine learns per project, users must allocate GPU hours, which increases cost if classes evolve frequently. Reasonable accuracy still demands roughly 2,000 seed labels per class, so the “few” in few-shot isn’t free. ### TrainingData.pro **TrainingData.pro** offers fully managed data-collection and annotation services for image, video, text, audio, DICOM, LiDAR. Custom quotes start at a $550 minimum invoice and 100 % post-payment terms. The concierge model handles security (GDPR, NDAs) and iBeta-certified workflows for biometrics, but lacks a self-serve IDE and publishes no per-label pricing, so smaller teams may face longer sales cycles than with other tools. ### Keymakr **Keymakr** focuses on retail, smart-home footage, and security CCTV. It offers pixel-level segmentation and attribute tagging with claimed accuracy above 95%. Its annotators receive domain-specific training to enable faster ramp-up and fewer guideline iterations. The firm provides a custom QA dashboard that visualizes class distribution across store layouts. Yet this narrow focus limits relevance to other verticals like autonomous vehicles, robotics, or NLP. Keymakr’s pricing is competitive within retail but less so for general vision work. ### Playment (GT Studio) **Playment’s GT Studio**—acquired by TELUS International in 2021—specializes in autonomous-vehicle data, and lets teams label video, LiDAR, and radar concurrently in a synchronized 3-D viewer. Enterprise clients benefit from SLAs and geo-fenced labeling centers. Since the TELUS acquisition, new contracts are enterprise-only and pricing is available on request, creating a barrier for startups. Some customers note slower feature rollouts post-merger. ### Scale AI **Scale AI** pioneered hybrid human-plus-model pipelines that have labeled billions of frames for autonomous driving, defense, and industrial robotics. Reviewers praise Scale’s throughput, helped by a global on-demand workforce and integrated automation. Premium SLAs, FedRAMP Moderate authorization, and dedicated annotation taxonomies make it popular among government and Fortune 500 customers. The flip side is cost (often 1.5-2× generalist vendors) and tight vendor lock-in, because many tooling components are proprietary. ## Getting Started — What Matters Most Before signing an annotation contract, align on the _outcomes_ that will make your models production-ready. That means looking beyond headline price and asking how well each vendor fits your data, workflows, and risk profile. Use the checklist below to focus discussions and structure a low-risk pilot: - **Domain & data fit:** Does the provider have proven success with your [modality](https://docs.voxel51.com/user_guide/groups.html) (e.g., LiDAR, medical DICOM) and ontology complexity? - **Quality guarantees:** Are precision/recall targets, disagreement thresholds, and escalation paths clearly spelled out in the SLA? - **Workflow integration:** Will labels flow directly into your MLOps pipeline (S3, GCS, DVC, etc.) and can you automate model-in-the-loop pre-labels or QA? - **Security & compliance:** Do you need HIPAA, FedRAMP, [on-prem deployment](https://docs.voxel51.com/enterprise/installation.html)? - **Pricing transparency:** How do costs scale after automation kicks in—per label, per hour, or usage tiers tied to active-learning loops? ### Run a Data-Backed Pilot 1. **Sample ≈1 % of your dataset.** Choose a representative slice that covers common classes and edge-cases. 2. **Instrument baseline metrics.** Track QA pass rates, review latency, disagreement rate, and Reconstruction Error Ratios (RERs). 3. **Automate the feedback loop.** Pass low-confidence or high-RER samples back to annotators while auto-accepting high-confidence labels to keep humans focused on the hardest cases. 4. **Stress-test collaboration.** Invite reviewers, engineers, and project managers: ensure review threads and [versioning](https://docs.voxel51.com/enterprise/dataset_versioning.html) works as expected. 5. **Score the pilot.** Compare cost per labeled item, turnaround, and rework effort against your internal KPIs. Shortlist only vendors that meet or beat your thresholds. When the pilot clears your bar, lock in SLAs, define automated QA hooks, and scale from 1% to 100% of the data in staggered batches. By keeping metrics front and center, you shorten time-to-model, control spend, and free your team to iterate on model innovation instead of annotation. ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-22-lllmstxt|> ## Multimodal AI Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps), [Industry Solutions](https://voxel51.com/blog/category/industry-solutions) The Multimodal Frontier in Computer Vision, Medicine, and Agriculture— CVPR 2025 Reflections Jun 24, 2025 • 6 min read Article content In this article [Why Multimodality Matters](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#5d1809bf75d0) [Multimodal Computer Vision — Oral Session](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#4a6a471a4b30) [1\. SegEarth-OV: Training-Free Open-Vocabulary Segmentation\\ for Remote Sensing](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#f4f8df043bd0) [2\. IceDiff: High-Resolution Arctic Sea Ice Forecasting via Generative Diffusion](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#e70fb9a409b2) [3\. Efficient Test-Time Adaptive Detection via Sensitivity-Guided Pruning](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#eff96d57967d) [4\. Keep the Balance: Parameter-Efficient RGB+X Semantic\\ Segmentation](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#2ab4b6ef3095) [M&M Workshop: Multimodal Models and Medicine](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#ea12ad87c7ed) [Multimodal Computer Vision and Foundation Models in Agriculture](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#043cb08bd5e2) [Reflections: The Maturation of Multimodal AI](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#93d9afe4d0f3) [What is next?](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#831e47d70380) In this article [Why Multimodality Matters](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#5d1809bf75d0) [Multimodal Computer Vision — Oral Session](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#4a6a471a4b30) [1\. SegEarth-OV: Training-Free Open-Vocabulary Segmentation\\ for Remote Sensing](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#f4f8df043bd0) [2\. IceDiff: High-Resolution Arctic Sea Ice Forecasting via Generative Diffusion](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#e70fb9a409b2) [3\. Efficient Test-Time Adaptive Detection via Sensitivity-Guided Pruning](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#eff96d57967d) [4\. Keep the Balance: Parameter-Efficient RGB+X Semantic\\ Segmentation](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#2ab4b6ef3095) [M&M Workshop: Multimodal Models and Medicine](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#ea12ad87c7ed) [Multimodal Computer Vision and Foundation Models in Agriculture](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#043cb08bd5e2) [Reflections: The Maturation of Multimodal AI](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#93d9afe4d0f3) [What is next?](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections#831e47d70380) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Why Multimodality Matters CVPR 2025 was an exceptional experience for me. I had the chance to learn from outstanding researchers and see the latest ideas in computer vision. How many people are now working on multimodal AI caught my attention. This means building systems that don’t just see images, but can also understand other kinds of data like text, sound, depth, or even temperature. I’ve always believed that the real world is not just something we look at. We hear it, feel it, and understand it through many senses. When I walk into a room, I don’t just see it — I notice how warm it is, how it sounds, if it’s calm or busy. Cameras can’t do that alone. However, with multimodal systems, we can start teaching machines to understand the world a little more like we do. Even more exciting, multimodality can help us go beyond our human limits. These systems can use sensors to detect things we can’t see, hear, or feel — like invisible gases, vibrations, or radiation. This means we’re not just making machines more human-like; we’re also extending what’s possible, improving how we capture and understand the world around us. In this blog, I’m sharing highlights from the sessions I attended and helped organize at CVPR 2025. I focus on how multimodal systems push the boundaries in computer vision and medicine, and what that means for our future. ## Multimodal Computer Vision — Oral Session CVPR2025 Oral 3B session brought together groundbreaking research exploring how to fuse, adapt, and condition models across multiple data modalities, from satellite imagery to RGB+X segmentation and climate modeling. ## [1\. SegEarth-OV: Training-Free Open-Vocabulary Segmentation\ \ for Remote Sensing](https://openaccess.thecvf.com/content/CVPR2025/papers/Li_SegEarth-OV_Towards_Training-Free_Open-Vocabulary_Segmentation_for_Remote_Sensing_Images_CVPR_2025_paper.pdf) This work proposes a novel method for semantic segmentation of remote sensing images without task-specific training. By adapting CLIP-based features for dense prediction, the authors introduced the SynUp module to upsample low-resolution patch tokens and mitigate global bias in CLIP’s class token representations. A content retention module preserves spatial details, crucial for segmenting delicate structures like buildings or roads. The method showed strong generalization across 17 datasets. ![](https://cdn.sanity.io/images/h6toihm1/production/88ebbc8e14de16ad50f5c0a977f589acaa881aac-1290x730.webp?auto=format&dpr=2&fit=max&q=75&w=1290) ## [2\. IceDiff: High-Resolution Arctic Sea Ice Forecasting via Generative Diffusion](https://openaccess.thecvf.com/content/CVPR2025/papers/Xu_IceDiff_High_Resolution_and_High-Quality_Arctic_Sea_Ice_Forecasting_with_CVPR_2025_paper.pdf) Targeting climate forecasting, IceDiff combines a U-Net predictor with a guided diffusion-based super-resolution module. The system improves spatial fidelity and temporal consistency by downscaling coarse 25km forecasts to fine-grained predictions. Its patch-based inference strategy and dynamic noise guidance mechanism ensure robustness to extreme climate events. This model exemplifies how temporal-spatial multimodality can be leveraged to improve environmental monitoring systems. ## [3\. Efficient Test-Time Adaptive Detection via Sensitivity-Guided Pruning](https://openaccess.thecvf.com/content/CVPR2025/papers/Wang_Efficient_Test-time_Adaptive_Object_Detection_via_Sensitivity-Guided_Pruning_CVPR_2025_paper.pdf) This work confronts the challenge of online domain adaptation for object detectors under shifting environments (e.g., day-to-night, foggy-to-clear). The proposed method prunes sensitive feature channels that are unstable across domains and focuses adaptation only on domain-stable channels. Through global and object-specific sensitivity scores, the model reduces computation by 12% and improves robustness without needing source data access. A strong step toward resource-aware multimodal adaptation. ![](https://cdn.sanity.io/images/h6toihm1/production/d7584bb7e575a7f6c2b40ef2e102058a29a251c4-1279x844.webp?auto=format&dpr=2&fit=max&q=75&w=1279) ## [4\. Keep the Balance: Parameter-Efficient RGB+X Semantic\ \ Segmentation](https://openaccess.thecvf.com/content/CVPR2025/papers/Cai_Keep_the_Balance_A_Parameter-Efficient_Symmetrical_Framework_for_RGBX_Semantic_CVPR_2025_paper.pdf) This paper addresses the growing use of RGB+X data (e.g., RGB+depth, RGB+thermal) in vision tasks, introducing a lightweight framework with modality-specific adapters and a dynamic fusion strategy. Instead of relying on heavy dual-stream models, the proposed architecture uses modality-aware prompting, spatially adaptive fusion, and a self-teaching mechanism to handle cases where one modality is missing or corrupted. It achieved competitive performance using just 4.4% of the parameters of full fine-tuning, highlighting progress toward scalable, real-world multimodal deployment. ![](https://cdn.sanity.io/images/h6toihm1/production/21d062e44dd0cac062bb2afbcf3358606d594312-1290x977.webp?auto=format&dpr=2&fit=max&q=75&w=1290) ## M&M Workshop: Multimodal Models and Medicine One of the most impactful sessions I attended at CVPR 2025 was the [M&M: Multi-modal Models and Medicine workshop](https://mandm2025.github.io/). It demonstrated the challenge of integrating fragmented healthcare data into coherent, actionable intelligence, from clinical text and imaging to signals and structured records. - **Lena Maier-Hein** opened the workshop with the concept of xeno-learning, drawing parallels to xeno-transplantation in medicine. The idea is to leverage cross-species data for spectral image analysis in pathology. Her call for benchmark standardization and international data-sharing collaborations emphasized the need for collaborative multimodal infrastructure in medical AI. ![](https://cdn.sanity.io/images/h6toihm1/production/e48df6e82ccbd0941e3ebac648835002d7e62805-1400x777.webp?auto=format&dpr=2&fit=max&q=75&w=1400) - **Vivek Natarajan** showcased Gemini for Biomedicine, an AI system capable of reasoning over multimodal inputs like CT scans and patient dialogue. The demonstration highlighted interactive diagnosis, medical conversation agents, and systems like AMIE that emulate clinical decision-making through structured dialogue and visual grounding. - **Akshay Chaudhari** presented some of the most exciting vision-language foundation models in radiology: 1) **RoentGen**, a generative model that creates chest X-rays from textual prompts; 2) **Merlin**, a 3D CT foundation model trained on 15.5K scans and evaluated across 10,000 internal and external studies; 3) **CheXAgent**, a powerful 8B parameter model for open-ended clinical QA on X-ray data. What was most striking was that Merlin is nearing FDA clearance for bone density estimation, underscoring how multimodal models are no longer confined to research labs but are entering clinical pipelines with real diagnostic utility. ![](https://cdn.sanity.io/images/h6toihm1/production/04cf6beaef1d2c0080fc3a03cf3915bf76c86f77-1228x665.webp?auto=format&dpr=2&fit=max&q=75&w=1228) ## Multimodal Computer Vision and Foundation Models in Agriculture This year, I also had the privilege of co-organizing the CVPR 2025 tutorial on [Multi-Modal Computer Vision and Foundation Models in Agriculture](https://www.agriculture-vision.com/agriculture-vision-2025/tutorial-2025) on June 12. It was an incredible experience to bring together researchers and practitioners working at the intersection of AI, agriculture, and foundation models. We opened the morning with Dr. **Melba Crawford**, who presented a powerful overview of sensor fusion in yield prediction. Her case studies, from multispectral to LiDAR, illustrated how diverse sensing platforms can be combined to inform critical agricultural decisions, especially when fused through vision-based modeling. Dr. **Alex Schwing** followed with a clear and concise breakdown of foundation models like CLIP, SAM, and DINO. His session highlighted the architectural innovations behind these models and how they’ve unlocked capabilities such as zero-shot learning, segmentation, and generative modeling, all through multimodal fusion. Finally, Dr. **Soumik Sarkar** presented compelling case studies applying foundation models to agriculture. From pest monitoring to weather-aware yield prediction, he emphasized how fusing textual, visual, and environmental data is key to creating robust agricultural AI systems. His reflections on data curation, computing, and evaluation brought practical depth to the discussion. Being part of this tutorial reminded me that agriculture is one of the most multimodal challenges we face: soil, weather, plant health, satellite imagery, farmer knowledge, everything matters. And we’re just starting to scratch the surface. ## Reflections: The Maturation of Multimodal AI After attending both research and applied sessions, I felt something shift. Multimodal AI is no longer just an academic curiosity. It’s becoming the standard, especially in areas where real-world complexity demands it. Models are now learning how to combine images and text and when to trust each modality, compensate for missing signals, and remain efficient and explainable. This is a fusion, and it is reasoning. The stakes are high in healthcare, agriculture, and autonomous systems. We’re seeing the rise of AI collaborators that can assist, explain, adapt, and even innovate alongside domain experts. The systems I saw at CVPR2025 are learning to work with us, not just for us. ## What is next? If you’re interested in following along as I dive deeper into the world of AI and continue to grow professionally, feel free to connect or follow me on [LinkedIn](https://www.linkedin.com/in/paula-ramos-phd/). Let’s inspire each other to embrace change and reach new heights! You can find me at some Voxel51 events ( [https://voxel51.com/computer-vision-events/](https://voxel51.com/events)), or if you want to join this fantastic team, it’s worth taking a look at this page: [https://voxel51.com/careers/](https://voxel51.com/careers) [CVPR](https://voxel51.com/blog/tag/cvpr) [medical imaging](https://voxel51.com/blog/tag/medical-imaging) [multimodal](https://voxel51.com/blog/tag/multimodal) [agriculture](https://voxel51.com/blog/tag/agriculture) ![](https://cdn.sanity.io/images/h6toihm1/production/e926c07c7d1426c0fde8fdefa637c528d47b16f4-512x512.webp?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/8bf3c50a5edd9f83e1013ad5df86ae159d519dae-1200x626.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Best of CVPR 2025: Conversations at the Cutting Edge of AI\\ \\ Event Recaps\\ \\ • \\ \\ Jul 3, 2025](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/97cf3d887735ab9574b2b3e2d3825146016d875e-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Embodied Computer Vision at CVPR 2025: The Next AI Frontier\\ \\ Event Recaps\\ \\ • \\ \\ Jun 30, 2025](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier) [![](https://cdn.sanity.io/images/h6toihm1/production/7ec61c80f387b16f24b4b2fe33f804264a38b486-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Rethinking How We Evaluate Multimodal AI\\ \\ Event Recaps\\ \\ • \\ \\ Jun 12, 2025](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-23-lllmstxt|> ## Computer Vision in Sports [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Industry Solutions](https://voxel51.com/blog/category/industry-solutions), [Product & News](https://voxel51.com/blog/category/product-news) How Computer Vision Is Changing Sports Jan 16, 2024 • 17 min read Article content In this article [Industry Overview](https://voxel51.com/blog/how-computer-vision-is-changing-sports#8ddc79f52a15) [Key Industry Challenges in Sports](https://voxel51.com/blog/how-computer-vision-is-changing-sports#e31b159106df) [Computer Vision Applications in Sports](https://voxel51.com/blog/how-computer-vision-is-changing-sports#f0d73c1949a7) [Organizations at the Cutting Edge of Computer Vision in Sports](https://voxel51.com/blog/how-computer-vision-is-changing-sports#d644cf266bd9) [Sports Datasets](https://voxel51.com/blog/how-computer-vision-is-changing-sports#5104f0fe679b) In this article [Industry Overview](https://voxel51.com/blog/how-computer-vision-is-changing-sports#8ddc79f52a15) [Key Industry Challenges in Sports](https://voxel51.com/blog/how-computer-vision-is-changing-sports#e31b159106df) [Computer Vision Applications in Sports](https://voxel51.com/blog/how-computer-vision-is-changing-sports#f0d73c1949a7) [Organizations at the Cutting Edge of Computer Vision in Sports](https://voxel51.com/blog/how-computer-vision-is-changing-sports#d644cf266bd9) [Sports Datasets](https://voxel51.com/blog/how-computer-vision-is-changing-sports#5104f0fe679b) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the fourth installment of [Voxel51](https://voxel51.com/)’s computer vision industry spotlight blog series. In this series, we highlight how different industries — from construction to climate tech, from retail to robotics, and more — are using computer vision, machine learning, and artificial intelligence to drive innovation. We’ll dive deep into the main computer vision tasks being put to use, current and future challenges, and companies at the forefront. In this edition, we’ll focus on _sports_! Read on to learn about computer vision in the sports industry. ## Industry Overview Sports brings people around the globe together in fitness, fun, and the spirit of competition. In addition to being an entertainment source, the sports industry boosts economies by standing behind new technical innovations and opening up job opportunities. Key facts and figures: - The global sports market is growing, increasing from [$486.61 billion in 2022 to $512.14 billion in 2023](https://www.thebusinessresearchcompany.com/report/sports-global-market-report) - Millions of people worldwide are employed in the sports sector, including more than 456,000 people in the United States with [a projected 7 percent job growth over the next decade](https://www.workandmoney.com/s/sports-industry-job-stats-807f31c91e6642fa#:~:text=In%20the%20United%20States%20alone%2C,growth%20over%20the%20next%20decade) - Narrowing in on technology, the global sports technology market was valued at USD 13.14 billion in 2022 and is expected to [grow at a CAGR of 20.8% from 2023 to 2030​](https://www.grandviewresearch.com/industry-analysis/sports-technology-market) - The global AI in sports market is projected [to reach $19.2 billion by 2030](https://www.alliedmarketresearch.com/artificial-intelligence-in-sports-market-A12905), growing at a CAGR of 30.3% from 2021 to 2030​ Applying computer vision and artificial intelligence (AI) to sports opens up a multitude of helpful new tools for teams, coaches, sports analytics professionals, players, scouts, and fans, while also presenting a multi-billion dollar opportunity for tech companies. Computer vision enables organizations to build capabilities like real-time video analysis, fitness and health tracking, sports predictions, and improving the overall fan experience. These tech advancements can help improve how things like player performance analysis, fan engagement, and marketing strategies are handled. Before we dive into various popular applications of computer vision-based AI technologies in sports, it’s important to highlight the key challenges facing the industry, presenting areas of opportunity for AI innovations. ## Key Industry Challenges in Sports - **Player health and performance:** Ensuring player well-being and implementing [strategies to optimize player availability](https://barcainnovationhub.fcbarcelona.com/blog/factors-affecting-the-availability-and-performance-of-elite-athletes-in-team-sports-2/) is crucial for sustained performance in elite team sports​. - **Fan engagement in the digital age:** While digital platforms have expanded reach and revenue, striking the right balance between meaningful interaction and oversaturation is crucial. The top 25 leagues in the world had a combined audience of over four billion and [generated more than €2.8 billion and 676 billion impressions](https://horizm.com/digital-value-of-fans-2023/) through their digital inventory. Keeping fans engaged and fostering a sense of community requires innovative strategies to ensure that the essence of sports isn't lost in the digital noise. - **Venue evolution:** [According to industry experts](https://www.sportsbusinessjournal.com/Journal/Issues/2016/04/18/Power-Players/Challenges.aspx), work must be done to make the venues themselves (stadiums, arena, and ballparks) more attractive, affordable, comfortable, safe, and technology-equipped to satisfy the needs of today’s fans and as they evolve over time. Continue reading to learn about several exciting and useful ways in which computer vision applications are helping organizations in the sports industry. ## Computer Vision Applications in Sports ### Sports Analytics & Strategy ![](https://cdn.sanity.io/images/h6toihm1/production/0777deaa4e4dd3cdb8fe768c0cbd395d2e203fb6-1152x648.png?auto=format&dpr=2&fit=max&q=75&w=1152) Sports are exhilarating for all involved — teams, players, coaches, and fans. Whether it’s a come-from-behind victory or a record-breaking play, the thrill of the game keeps people coming back for more. And when one match ends, sports pros on the field and in the back office are already thinking of ways to continue to improve performance and safety, and make the game even more exciting for fans. Data is the key to unlocking winning strategies. The use of cameras, equipment sensors, wearables, and even radar and LiDAR scans like in the MLB, makes a variety of visual information available. Now, every jump, sprint, shot, throw, and maneuver can be captured so that the information can be organized and analyzed. Tracking players' movements, positions, speeds, and trajectories offers a rich data source. This wealth of information enables rich analytics for coaches, athletes, and sports professionals to gain valuable insights into performance — not only to advance individual and team performance at game time, but also to refine training plans, scout new talent, and for competitive analysis and strategies. Pose estimation and object tracking are on the forefront of computer vision in sports. For example, coaches and analysts use pose estimation to determine ideal swing and pitch patterns in MLB to squeeze every ounce of performance out of their players. Soccer is also quick to adopt both of these technologies, tracking players on the pitch and seeing how they react to sudden changes in the ball's positions. Through thorough analysis of penalty kick pose estimations, goalies could get a leg up on the competition and look for tells on where the shot is likely to go, using insights only possible through computer vision. For further reading, check out these articles on computer vision and AI in popular sports franchises: - [A Deeper Look at MLB's New Statcast](https://builtin.com/consumer-tech/mlb-statcast-tech-update-hawk-eye-integration) - [Fueling the future of sports: How the NFL is using data to change the game–on the field, in the stands, and in your home](https://techcrunch.com/sponsor/aws/fueling-the-future-of-sports-how-the-nfl-is-using-data-to-change-the-game-on-the-field-in-the-stands-and-in-your-home/) - [The NBA’s Game-Changing Approach to Data](https://www.wired.com/sponsored/story/the-nbas-game-changing-approach-to-data/) - [NFL’s Next Gen Stats](https://nextgenstats.nfl.com/news/all) Also check out these papers related to using computer vision for athletic motion tracking: - [Motion Capture for Sporting Events Based on Graph Convolutional Neural Networks and Single Target Pose Estimation Algorithms](https://www.mdpi.com/2076-3417/13/13/7611) - [A survey on location and motion tracking technologies, methodologies, and applications in precision sports](https://www.sciencedirect.com/science/article/abs/pii/S0957417423009946) - [All Keypoints You Need: Detecting Arbitrary Keypoints on the Body of Triple, High, and Long Jump Athletes](https://arxiv.org/pdf/2304.02939.pdf) - [VIRD: Immersive Match Video Analysis for High-Performance Badminton Coaching](https://arxiv.org/pdf/2307.12539.pdf) - [A Comprehensive Review of Computer Vision in Sports: Open Issues, Future Trends and Research Directions](https://arxiv.org/pdf/2203.02281.pdf) ### Injury Prevention and Rehabilitation ![](https://cdn.sanity.io/images/h6toihm1/production/9fe69153a904033336d9d2229515f71fe6c63e29-1570x822.png?auto=format&dpr=2&fit=max&q=75&w=1570) Using computer vision in sports has led to the development of cutting-edge ways to prevent injuries and heal from them. By analyzing athletes' movements during exercises and game-time matches, computer vision techniques like feature extraction, [pose estimation](https://mobidev.biz/blog/human-pose-estimation-technology-guide) and [motion detection](https://link.springer.com/referenceworkentry/10.1007/978-3-030-58080-3_339-1) can detect improper actions that might lead to injuries. Analyzed data comes in handy for coaches and medical teams when they are putting together personalized training and conditioning programs to prevent potential injuries. This information is also important for improving the design and construction of protective gear and equipment. When it comes to rehabilitation, computer vision is an essential asset. The healing journey of athletes can be monitored to make sure that rehabilitation exercises are done correctly to lower the risk of reinjury. Techniques such as [marker-less human pose estimation](https://journals.sagepub.com/doi/full/10.1177/11795727211022330) are particularly promising, allowing for cost-effective and reliable telerehabilitation services without additional equipment​. [Digitizing rehabilitation not only improves accuracy but also holds the potential to speed up recovery](https://pubmed.ncbi.nlm.nih.gov/35960507/), making the process more organized and data-driven. Here is an example to check out on vision-based technologies being used in a popular sports franchise: - [NFL + AWS Partner to Create the Digital Athlete to Revolutionize Player Health & Safety](https://www.nfl.com/playerhealthandsafety/equipment-and-innovation/aws-partnership/digital-athlete-spot) Here are some papers related to using computer vision for injury prevention and rehabilitation: - [Hybridized Hierarchical Deep Convolutional Neural Network for Sports Rehabilitation Exercises](https://ieeexplore.ieee.org/abstract/document/9126809) - [A Beta Version of an Application Based on Computer Vision for the Assessment of Knee Valgus Angle: A Validity and Reliability Study](https://www.mdpi.com/2227-9032/11/9/1258) ### AI Referee Assistance ![](https://cdn.sanity.io/images/h6toihm1/production/fd0cc38a23042bb8722777cf268c93e7a02d12b4-1200x800.png?auto=format&dpr=2&fit=max&q=75&w=1200) AI referee assistance aids human referees in overseeing sports games. Through computer vision and machine learning, AI referees can make precise calls in real time, reducing human errors. They can detect goals, misconduct between players, and other rule infringements swiftly and accurately. This not only enhances the credibility of the game but also alleviates the pressure on human referees. In soccer for example, an AI refereeing system can use multiple cameras to capture the field from different angles. These images are then analyzed in realtime to detect events like fouls or offside situations. For example, in an offside scenario the system can calculate the positions of players, the ball, and the last defender at the moment the ball is played. If a player is found in an offside position, the system instantly alerts the human referee and can provide a visual representation of the scene on a sideline monitor for verification, ensuring that the call is accurate and fair. Here are a few resources on vision-based referee assistance technologies in use in popular sports franchises: - [NBA Replay Center](https://official.nba.com/replay/) - [MLS’s Video Review FAQ](https://www.mlssoccer.com/news/video-review-answering-all-your-frequently-asked-questions-337958) - [FIFA Video Assistant Referee](https://www.fifa.com/technical/football-technology/football-technologies-and-innovations-at-the-fifa-world-cup-2022/video-assistant-referee-var) Here are a couple of papers related to using computer vision for refereeing: - [VARS: Video Assistant Referee System for Automated Soccer Decision-Making from Multiple Views](https://arxiv.org/abs/2304.04617) - [Vision Based Dynamic Offside Line Marker for Soccer Games](https://arxiv.org/pdf/1804.06438.pdf) ### Fan Experience Enhancement ![](https://cdn.sanity.io/images/h6toihm1/production/7fa3f9da7f73688a05e2f0e67ed8231f152ed875-660x444.png?auto=format&dpr=2&fit=max&q=75&w=660) Enhancing fan experiences and fueling new ones are exciting applications of computer vision in sports. While there are more ways unfolding to inform and engage fans further than ever before, in this article we’ll focus on two prominent use cases: augmenting the broadcast experience and new ways to engage via Augmented Reality (AR) and Virtual Reality (VR). #### Augmenting the Broadcast Experience Computer vision is paving the way for enhancements to sports broadcasting to create a richer viewing experience. It’s now possible to analyze live and recorded content in real time in order to extract insights to pair with the broadcast. For example, vision-based AI can provide real-time statistics and player information, real-time ball tracking, the ability to generate on-screen graphics, as well as identify key moments in matches. All of these new possibilities help fans develop deeper connections to the game. Check out this resource on AI-powered sports features: - [7 AI-powered features you’ll find on Prime Video's ‘Thursday Night Football‘ this season](https://www.aboutamazon.com/news/aws/prime-video-thursday-night-football-next-gen-stats-ai-features) #### Immersive AR and VR Experiences With the combination of AR and VR, the traditional stadium experience is evolving, making sports events more engaging and interactive for fans. Fans can virtually explore stadiums, get the latest player profiles, enjoy 3D game rewinds, and live chat with other fans. Computer vision technology captures and analyzes the game as it happens, turning complex on-field actions into digital data. This data then powers the AR and VR applications, making spectators feel like they’re part of the game. [Moreover, renowned football teams and leagues have already begun experimenting with AR/VR technologies.](https://www.altmansolon.com/insights/ar-vr-game-changers-for-increasing-sports-fan-engagement-and-monetization/) Through interactive apps and other digital platforms, [fans can virtually immerse themselves in live games](https://www.archdaily.com/941030/how-ar-and-vr-will-enhance-the-future-of-the-sports-arena-experience), accessing features like panoramic camera angles, real-time stats, and on-demand replays. With the continued advancement in AR, VR, and computer vision technologies, the boundaries of how fans experience sports are set to expand further, potentially leading to features like holographic player projections and live streams of remote audiences, offering a new level of engagement and excitement for sports enthusiasts worldwide. Here are some resources related to enhancing the fan experience using AR and VR: - [The impact of virtual reality (VR) technology on sport spectators' flow experience and satisfaction](https://www.sciencedirect.com/science/article/abs/pii/S0747563218306265) - [The Metaverse and Sports: How VR and AR Could Transform the Fan Experience](https://medium.com/blockchain-smart-solutions/the-metaverse-and-sports-how-vr-and-ar-could-transform-the-fan-experience-d15e73c459a2) ### Personalized Fitness & Training ![](https://cdn.sanity.io/images/h6toihm1/production/4423c834b3b819994da4487eb0381bab8bde1335-2119x1415.jpg?auto=format&dpr=2&fit=max&q=75&w=1600) Computer vision is revolutionizing the realm of personalized fitness and training. Advancements in AI and ML make it possible to analyze human movement with remarkable accuracy, providing real-time feedback and tailored guidance to people pursuing their fitness goals. Popular use cases include: - **Posture and form tracking and correction:** By monitoring body positions, angles, and movements, computer vision systems can detect deviations from proper form, providing immediate feedback to help people correct their technique and prevent injuries. - **Personalized workout recommendations:** Data-driven approaches enable users to receive workout recommendations tailored to their specific needs and abilities, helping them achieve their fitness goals. - **Virtual coaching:** With the ability to remotely monitor and assess user movements, fitness professionals and coaches can provide personalized guidance and support in virtual settings. Check out this paper on using computer vision for fitness and training: - [Computer Simulation Based Parameter Selection for Resistance Exercise](https://arxiv.org/pdf/1306.4724.pdf) ## Organizations at the Cutting Edge of Computer Vision in Sports ### The United States Tennis Association (USTA) ![](https://cdn.sanity.io/images/h6toihm1/production/046110671119aa2cfaa117a574a77207fe0244db-1999x1334.jpg?auto=format&dpr=2&fit=max&q=75&w=1600) The [USTA](https://www.usta.com/en/home.html) is using AI to level up player performance. With an increasing amount of data available from court cameras, video recordings, and wearables during practice, AI plays a significant role in bringing performance insights to the forefront for the entire performance team of athletes, coaches, mental skills staff, and strength and conditioning professionals. At the heart of USTA’s AI-driven performance management system is data: the x-y coordinate of the player on the court, every shot of the rally, the speed of the shots, the spin, the number of changes in direction, and more. Being able to access and analyze critical data is important in evolving player strategies and winning more matches. AI-generated insights enable athletes to compare their technique with top players, for example those with outstanding backhands, to understand where exactly to tune up their stroke to improve performance. Athletes can also analyze their performance over time, and narrow in on perfecting the one or two techniques that will make a sizable impact on match performance. Tennis tournaments generate a wealth of data, while highlighting the importance of AI-powered systems for the sport: athletes can analyze the performance of players they will compete against, as well as analyze their own matches once they’re over, so they can swiftly and continuously tune and improve. ### AiSport ![](https://cdn.sanity.io/images/h6toihm1/production/d2b0e422505c9f793299702971c2dc62c413cbf8-1999x968.png?auto=format&dpr=2&fit=max&q=75&w=1600) [AiSport](https://ai-prosport.com/) is building an AI fitness platform to give users real-time feedback on their workout techniques, all from the convenience of their smartphones. AiSport’s platform uses AI and computer vision to analyze and correct people’s posture while they are exercising to maximize effectiveness and prevent injuries. With AiSport, fitness clubs can provide their members with not only equipment and a location, but also with a personal AI fitness trainer. The AiSport’s tech analyzes and provides real-time recommendations to maximize performance while keeping workouts injury-free. Computer vision techniques at the center of AiSport’s platform include: 3D body pose and shape recognition, biomechanical analysis, pattern matching, and deviation estimation. AiSport was co-founded by two Ukrainian women and long-time sports enthusiasts, Anna Stepura and Dariia Hordiiuk. Development of the AiSports platform continues today in Silicon Valley. ### Hawk-Eye Innovations ![](https://cdn.sanity.io/images/h6toihm1/production/43a0cacbe36b3e7100b0361196e583147af635d3-425x523.png?auto=format&dpr=2&fit=max&q=75&w=425) [Hawk-Eye Innovations](https://www.hawkeyeinnovations.com/), a pioneering UK-based company and part of the Sony group, focuses on applying computer vision to sports. With a team of dedicated professionals, Hawk-Eye has become a household name in the sporting world, delivering precise and real-time tracking, analytics, and officiating assistance in a wide variety of sports, including tennis, cricket, and soccer. Their [key technologies](https://www.hawkeyeinnovations.com/our-technology) include the Synchronized Multi-Angle Replay Technology (SMART) for enhanced video capture, review, clipping, and distribution, the TRACK systems for Performance Tracking, Ball Tracking, and Object Tracking, and the INSIGHT suite, which provides data collation, storage, aggregation, delivery, and visualization capabilities. Founded in 2001, Hawk-Eye has grown into a thriving company with a global presence. Not only have they partnered with [Major League Baseball (MLB)](https://www.prnewswire.com/news-releases/hawk-eye-innovations-and-mlb-introduce-next-gen-baseball-tracking-and-analytics-platform-301115828.html) for optical tracking and vision-processing technology and the [National Basketball Association (NBA)](https://pr.nba.com/nba-sony-hawk-eye-innovations-partnership/) for deploying 3D optical tracking technology, they have achieved international recognition and are an integral part of sports events in more than 90 countries worldwide​. ### Sportlogiq ![](https://cdn.sanity.io/images/h6toihm1/production/1f9b69e286cccf0886039366fb8a5924888b7aec-1024x574.png?auto=format&dpr=2&fit=max&q=75&w=1024) Based in Montreal, Quebec, [Sportlogiq](https://www.sportlogiq.com/) initially focused on professional hockey and then expanded its reach to collaborate with major sports teams and data providers around the world. Numerous NHL clubs, more than 150 professional and amateur hockey teams worldwide, media outlets, content producers, top performance research firms for soccer and football, as well as amateur sports and video companies, all rely on their data and insights. The company is [supported by prominent investor Mark Cuban](https://www.prnewswire.com/news-releases/mark-cuban-a-part-of-17m-financing-round-for-sportlogiq-incs-player-tracking-and-analytics-platform-515414781.html) and the [TandemLaunch incubator](https://www.tandemlaunch.com/en/portfolio/sportlogiq/). Sportlogiq is made up of a [team of 13 AI researchers](https://www.sportlogiq.com/about-us/). They have been granted 180 patents and publications and have 75 full-time professionals onboard. Sportlogiq is helping to shape the future of AI in sports by providing creative solutions for improving athletic performance and training. ### Ludimos ![](https://cdn.sanity.io/images/h6toihm1/production/8c6103f6621a26e07f8c1e0a4beaa36923a2f46f-374x600.png?auto=format&dpr=2&fit=max&q=75&w=374) [Ludimos](https://www.ludimos.com/), a smartphone-based cricket training app, is transforming the world of cricket coaching. Founded by Madan Rajagopal, an Indian cricket enthusiast living in the Netherlands, Ludimos was born out of his frustration with inconsistent coaching advice and a lack of tools to track players' progress. As a data scientist and AI engineer, Rajagopal developed Ludimos to address these challenges. The app has gained widespread popularity, with over 19,000 users across 15 countries, including national cricket associations and teams like Royal Challengers Bangalore. What sets Ludimos apart are its specialized features aimed at cricket training. It offers multi-angle video analysis, allowing a thorough look at player techniques from different viewpoints. The app excels in ball and bat tracking, giving a data-driven insight into player performance. Additionally, it provides a communication platform for coaches to assign drills and give feedback, making the coaching process more interactive and efficient. While ball tracking is its strong suit for now, Ludimos has plans to expand into bat tracking and biomechanics analysis, showing a promising trajectory for evolving cricket coaching and player analysis. ### Track160 ![](https://cdn.sanity.io/images/h6toihm1/production/f9d19fdf64fa790961fdb5c8b36f558d5a1a9583-625x513.png?auto=format&dpr=2&fit=max&q=75&w=625) [Track160](https://www.track160.com/) is changing the way soccer is coached and analyzed. Founded and chaired by Miky Tamir, a pioneer in sports computer vision, Track160 combines cutting-edge technology with a deep understanding of the game to provide valuable insights to soccer clubs and academies. At the core of Track160's offerings is an AI-based solution that utilizes multiple cameras tethered to a single base. These cameras, equipped with computer vision and deep learning algorithms, capture a wealth of data related to player performance and team tactics. What sets Track160 apart is its commitment to data accuracy, earning FIFA certification for its data quality from a single installation point. ### Tonal ![](https://cdn.sanity.io/images/h6toihm1/production/1be306a241efdc21d58ccf59495ac477e143bcce-1999x1125.jpg?auto=format&dpr=2&fit=max&q=75&w=1600) [Tonal](https://www.tonal.com/) is an AI-powered home gym system that combines strength training equipment, personalized fitness coaching, and live and on-demand classes. It was founded by Aly Orady, a supercomputer engineer who wanted to create a more effective and convenient way to strength train at home. Tonal's main piece of equipment is a wall-mounted device with two electromagnetic pulleys that can provide up to 200 pounds of resistance. The device also has a touchscreen display that shows you how to perform each exercise and tracks your progress. Tonal uses AI to dynamically adjust the resistance for each exercise based on your individual strength and fitness level in order to provide you with your most effective workout. Tonal is [trusted by](https://www.tonal.com/reviews/) an impressive number of world-class athletes. ## Sports Datasets If you are interested in exploring applications of computer vision in sports, check out these datasets: ![](https://cdn.sanity.io/images/h6toihm1/production/763122e5c71de2c7b2a1cecaff282460a9a17f95-805x404.png?auto=format&dpr=2&fit=max&q=75&w=805) - [SoccerNet-V3](https://try.fiftyone.ai/datasets/soccernet-v3/samples): A fantastic collection of annotated soccer games capturing the key moments of the league. Spanning 3 years, 33 teams, and almost 1800 matches, the dataset features both bounding box and polyline detection. This dataset combines the best of both worlds allowing for high level sports analysis with high grain filters, as well as annotation to empower even the most advanced computer vision techniques. - [NFL-Impact-Detection](https://try.fiftyone.ai/datasets/nfl-impact-detection-dataset/samples): Helping the NFL usher in a new age for analytics and player safety, the NFL Impact Detection dataset serves as an open community challenge to help detect concussions in videos from past NFL games. In a different approach than most datasets, the dataset has annotated bounding boxes on only the helmets of the players. This dataset is just one of the challenges the NFL has issued to help advance safety in football by using machine learning. - [Football Player Segmentation](https://www.kaggle.com/datasets/ihelon/football-player-segmentation): This dataset is specifically designed for computer vision tasks related to player detection and segmentation in soccer (football) matches. The dataset contains images of players in different playing positions, such as goalkeepers, defenders, midfielders, and forwards, captured from various angles and distances. The images are annotated with pixel-level masks that indicate the player's location and segmentation boundaries, making it ideal for training deep learning models for player segmentation. [Explore this dataset with FiftyOne in your browser.](https://try.fiftyone.ai/datasets/football-player-segmentation/samples) - [SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes](https://github.com/MCG-NJU/SportsMOT): A large-scale multi-object tracking dataset consisting of 240 video clips from 3 categories (i.e., basketball, football, and volleyball). The objective is to only track players on the playground (i.e., except for a number of spectators, referees and coaches) in various sports scenes. [Explore this dataset with FiftyOne in your browser.](https://try.fiftyone.ai/datasets/sportsmot/samples) - [Sports Videos in the Wild (SVW): A Video Dataset for Sports Analysis](http://cvlab.cse.msu.edu/project-svw.html): This dataset comprises 4,200 videos captured via smartphones, covering 30 categories of sports and 44 different actions. The dataset can be used for genre categorization, action recognition, and other computer vision applications. - [DeepSportRadar-v1](https://paperswithcode.com/dataset/deepsportradar-v1): This dataset can be used to solve four challenging tasks related to basketball—ball 3D localization, camera calibration, player instance segmentation, and player re-identification. - [UCF Sports Action Data Set](https://www.crcv.ucf.edu/data/UCF_Sports_Action.php): This dataset consists of actions collected from various sports typically featured on broadcast television channels such as the BBC and ESPN. The video sequences were obtained from a wide range of stock footage websites, including BBC Motion Gallery and GettyImages. The dataset includes a total of 150 sequences. - [Olympic Sports Dataset](http://vision.stanford.edu/Datasets/OlympicSports/): This dataset consists of videos of athletes practicing 16 different sports. - [Sports-1M](https://paperswithcode.com/dataset/sports-1m): This dataset consists of over a million videos with 487 sports-related categories, with 1,000 to 3,000 videos per category. If you would like to see any of these or other computer vision sports datasets added to the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html), get in touch, and we can work together to make this happen! [ai referee](https://voxel51.com/blog/tag/ai-referee) [fitness](https://voxel51.com/blog/tag/fitness) [industry use cases](https://voxel51.com/blog/tag/industry-use-cases) [injury prevention](https://voxel51.com/blog/tag/injury-prevention) [rehabilitation](https://voxel51.com/blog/tag/rehabilitation) [sports](https://voxel51.com/blog/tag/sports) [sports analytics](https://voxel51.com/blog/tag/sports-analytics) [sports fan enhancement](https://voxel51.com/blog/tag/sports-fan-enhancement) ![](https://cdn.sanity.io/images/h6toihm1/production/3b39056326e925c10b46da1324bc3c5840a1629c-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Dan Gural Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/6aafb2b5fa699824c252fabfe2607eaeb820616a-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Enabling the AV Datasets of the Future with NVIDIA NuRec and FiftyOne\\ \\ Product & News\\ \\ • \\ \\ Aug 11, 2025](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/663dd6a3e6f3e57a03425f932440b5d242133451-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Search and curate video data with FiftyOne, Twelve Labs, and Databricks Vector Search\\ \\ Product & News, Integrations\\ \\ • \\ \\ Jun 5, 2025](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs) [![](https://cdn.sanity.io/images/h6toihm1/production/35e10bff3f7e49854806cbbd163022932bf07548-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How Voxel51 is Powering Physical AI with Databricks\\ \\ Product & News\\ \\ • \\ \\ Aug 20, 2025](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-24-lllmstxt|> ## CVPR 2023 Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) CVPR 2023 and the State of Computer Vision May 18, 2023 • 9 min read Article content In this article [Patterns and trends in computer vision from CVPR papers](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#9447893e9a3e) [Digging into the data](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#4d81bff8893d) [Heating up](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#b62e71351d02) [How abstract was your abstract? Creativity in computer vision](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#fb511f1c65f6) [Visit Voxel51 at CVPR!](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#cc4fb3841eea) [Join the FiftyOne community!](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#bd6de6555782) In this article [Patterns and trends in computer vision from CVPR papers](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#9447893e9a3e) [Digging into the data](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#4d81bff8893d) [Heating up](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#b62e71351d02) [How abstract was your abstract? Creativity in computer vision](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#fb511f1c65f6) [Visit Voxel51 at CVPR!](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#cc4fb3841eea) [Join the FiftyOne community!](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#bd6de6555782) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Patterns and trends in computer vision from CVPR papers \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop The annual IEEE/CVF [Conference on Computer Vision and Pattern Recognition](https://cvpr2023.thecvf.com/) (CVPR) is just around the corner. Every year, thousands of computer vision researchers and engineers from across the globe come together to take part in this monumental event. The prestigious conference, which can trace its origin back to 1983, represents the pinnacle of progress in computer vision. With CVPR playing host to some of the field’s most pioneering projects and painstakingly crafted papers, it's no wonder that the conference has the fourth highest h5-index of any conference or publication, trailing only _Nature_, _Science_, and _The New England Journal of Medicine._ This year, the conference will take place in Vancouver, Canada, from June 18th - June 22nd. With [2359 accepted papers](https://cvpr2023.thecvf.com/Conferences/2023/AcceptedPapers), [100 workshops](https://cvpr2023.thecvf.com/Conferences/2023/workshop-list), [33 tutorials](https://cvpr2023.thecvf.com/Conferences/2023/tutorial-list), and Flash Sessions happening in the Expo (including two by Voxel51!), CVPR will have something for everyone. Such a large volume of high quality content can be overwhelming, but it also provides us with invaluable insight into the current state of affairs in computer vision, what’s hot, and where the field is going. To help you navigate the rapidly changing world of computer vision in 2023, we gathered, scraped, cleaned, and dug into the data. We even used AI to quantify how creative authors were! In this blog post, we’ll share what we learned. And stay tuned for our upcoming CVPR 2023 Survival Guide, including the top ten papers you won’t want to miss! - [Digging into the CVPR paper data](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#digging-into-the-data) - [What’s trending in computer vision](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#whats-trending) - [Analyzing creativity in computer vision with AI](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#creativity-ai) - [Where to find Voxel51 at CVPR](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision#voxel51-at-cvpr) ## Digging into the data The starting point for all of our analyses was the [list of all accepted papers](https://cvpr2023.thecvf.com/Conferences/2023/AcceptedPapers). We scraped the web to find the actual papers corresponding to the titles in the initial list. If there was a version of the paper on Arxiv, we grabbed the author list, title (for many papers, the title on Arxiv differed slightly from the title posted on the CVPR website), and abstract. If the paper was not on Arxiv, but was available elsewhere on the web, we attempted to capture that information instead. Note: this data was scraped on April 20th, 2023. Any papers that have been uploaded to the Arxiv in the intervening weeks may be excluded from this analysis. ### Quick facts - 2359 papers accepted (out of 9155 submissions) - 1724 papers with versions on Arxiv - 68 additional papers found elsewhere ### Authors per paper - The average CVPR paper had approximately 5.4 authors. - The paper with the most authors was [_Why is the winner the best?_](https://arxiv.org/abs/2303.17719), which had 125 authors. All 125 of them can share this ignominious honor. - 13 papers had only one author. ### Primary Arxiv category Of the 1724 papers with versions on Arxiv, 1545, or just under 90%, had cs.CV listed as their primary category. cs.LG was second, with 101. eess.IV (26) and cs.RO (16) also claiming a piece of the pie. Other categories with CVPR papers include: cs.HC, cs.CV, cs.AR, cs.DC, cs.NE, cs.SD, cs.CL, cs.IT, cs.CR, cs.AI, cs.MM, cs.GR, eess.SP, eess.AS, math.OC, math.NT, physics.data-an, and stat.ML. ### “Meta” data - The words “dataset” and “model” appeared jointly in 567 abstracts. “Dataset” appeared on its own in 265 abstracts, whereas “model” appeared independently 613 times. Only 16.2% of CVPR accepted paper abstracts contained neither word. - According to CVPR paper abstracts, the most popular datasets this year were ImageNet (105), COCO (94), KITTI (55), and CIFAR (36). - 28 papers introduce a new “benchmark”. ### ACRONYMS abound It seems that you can’t have machine learning projects without acronyms. Out of the 2359 papers, the titles of 1487 had acronyms or compound words with multiple capital letters in them. That’s 63%! Some of these acronyms are catchy - they are easy to read and roll off the tongue: - [CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal Pose](https://arxiv.org/abs/2206.11752) - [PATS: Patch Area Transportation with Subdivision for Local Feature Matching](https://arxiv.org/abs/2303.07700) - [CIRCLE: Capture In Rich Contextual Environments](https://arxiv.org/abs/2303.17912) Some are more difficult to speak aloud than others: - [SIEDOB: Semantic Image Editing by Disentangling Object and Background](https://arxiv.org/abs/2303.13062#:~:text=SIEDOB%3A%20Semantic%20Image%20Editing%20by%20Disentangling%20Object%20and%20Background,-Wuyang%20Luo%2C%20Su&text=Abstract%3A%20Semantic%20image%20editing%20provides,the%20backgrounds%20are%20quite%20different.) - [FJMP: Factorized Joint Multi-Agent Motion Prediction over Learned Directed Acyclic Interaction Graphs](https://arxiv.org/abs/2211.16197) Some of them seemed to use creative license on acronym construction: - [SCOTCH and SODA: A Transformer Video Shadow Detection Framework](https://arxiv.org/abs/2211.06885) - [EXCALIBUR: Encouraging and Evaluating Embodied Exploration](https://github.com/facebookresearch/exploring_exploration) ## Heating up In addition to the 2023 paper titles, we scraped the titles for all 2022 accepted CVPR papers. From these two lists, we calculated the relative frequency of various keywords, giving high-level insight into what is trending up and what is trending down. Here are the highlights, presented in terms of percent change relative to 2022 frequencies. ### Models #### Diffusion models It should be no surprise to see diffusion models trending upward, with image generation models like stable diffusion and [Midjourney](https://www.midjourney.com/) going viral. Diffusion models are also finding applications in denoising, [image editing](https://arxiv.org/abs/2210.09276), and style transfer. Add all of this up, and you get by far the biggest winner across all categories, with a 573% increase year-over-year. #### Radiance Fields Neural Radiance Fields, or NeRFs, have also grown in popularity, as exhibited by an 80% increase in usage of the word _radiance_, and a 39% increase for _NeRF_. NeRFs have moved beyond proof of concept to [editing](https://arxiv.org/abs/2303.13277), [applications](https://arxiv.org/abs/2212.14710), and training process [optimization](https://arxiv.org/abs/2303.15951). #### Transformers The dips for “Transformer” and “ViT” are less indicative of transformer models going out of style, but rather a reflection of how dominant these models were in 2022. [In 2021](https://openaccess.thecvf.com/CVPR2021?day=all), the word “transformer” only appeared in 37 paper titles. In 2022, that number skyrocketed to 201. Transformers are not going away any time soon. #### Changing of the guard On the other hand, CNNs, which fell off by 68%, appear to be falling out of favor. Once the darling of computer vision, in 2023 it seems CNNs have lost their edge. Many titles that mention CNNs also mention other models. For instance, these papers mention both CNNs and transformers: - [Lite-Mono: A Lightweight CNN and Transformer Architecture for Self-Supervised Monocular Depth Estimation](https://arxiv.org/abs/2211.13202) - [Learned Image Compression with Mixed Transformer-CNN Architectures](https://arxiv.org/abs/2303.14978) ### Tasks #### Generative Traditional discriminative tasks like detection, classification, and segmentation are not falling out of favor, but their share of the computer vision mind-space is shrinking due to a flurry of advances in generative CV applications, as evidenced by upticks for “Editing”, “Synthesis,” and of course “Generation”. #### Masks The keyword “mask” saw a 263% increase year-over-year, appearing 92 times in the 2023 accepted paper titles - sometimes twice in a single title. Some of these occurrences are still in the context of segmentation: - [SIM: Semantic-aware Instance Mask Generation for Box-Supervised Instance Segmentation](https://arxiv.org/abs/2303.08578) - [DynaMask: Dynamic Mask Selection for Instance Segmentation](https://arxiv.org/abs/2303.07868) But the majority (64%) actually refer to “masked” tasks, including 8 instances of “masked image modeling”, and 15 “masked autoencoder” tasks. Additionally, there are 8 occurrences of “masking”. It is also worth noting that 3 paper titles with the word “mask” actually refer to “mask-free” tasks. #### Zero vs Few “Zero-shot” learning is gaining traction, with the rise of transfer learning, generative approaches, prompting, and general purpose models. At the same time, “few-shot” learning is down from last year. In terms of raw numbers, however, “few-shot” (45) maintains a slight edge over “zero-shot” (35), at least for now. ### Modalities #### The boundaries blur While the frequency of traditional computer vision keywords like “image” and “video” is relatively unchanged, “text”/”language” and “audio” are appearing more often. Even if the word “multi-modal” itself is not cropping up in paper titles, it is hard to deny that computer vision is trending towards a multi-modal future. This is especially pronounced for vision-language tasks, as evidenced by the sharp upticks for “Open”, “Prompt,” and “Vocabulary”. The most extreme example of this is the compound term “Open-vocabulary”, which occurred just 3 times in 2022, but shows up 18 times in 2023. #### Point cloud 9 Three-dimensional computer vision applications are moving away from inferring 3D information from 2D images (“Depth” and “Stereo”). Instead, computer vision systems are being trained to work directly on 3D point cloud data. ## How abstract was your abstract? Creativity in computer vision No attempt at comprehensive ML-related coverage in 2023 would be complete without bringing ChatGPT into the mix. We decided to make things interesting and use ChatGPT to find the _most creative_ titles from CVPR 2023. For each paper that had a draft uploaded to Arxiv, we scraped the abstract, and asked ChatGPT (GPT-3.5 API) to generate a title for the corresponding CVPR paper. We then took these ChatGPT generated titles and the actual paper titles, and used OpenAI’s \`text-embedding-ada-002\` model to generate embedding vectors, and computed the cosine similarity between ChatGPT-generated title and author-generated title. What can this tell us? The closer ChatGPT can get to the actual paper title, the more predictable it was. In other words, the further ChatGPT’s prediction was, the more “creative” the authors were in naming their paper. Embeddings plus cosine similarity gives us an interesting, albeit far from perfect, way to quantify this. We sorted the papers according to this metric. Without further ado, here are the most creative titles: **Actual**: Tracking Every Thing in the Wild **Predicted**: Disentangling Classification from Tracking: Introducing TETA for Comprehensive Benchmarking of Multi-Category Multiple Object Tracking **Actual** _:_ Learning to Bootstrap for Combating Label Noise **Predicted** _:_ Learnable Loss Objective for Joint Instance and Label Reweighting in Deep Neural Networks **Actual** _:_ Seeing a Rose in Five Thousand Ways **Predicted** _:_ Learning Object Intrinsics from Single Internet Images for Superior Visual Rendering and Synthesis **Actual** _:_ Why is the winner the best? **Predicted** _:_ Analyzing Winning Strategies in International Benchmarking Competitions for Image Analysis: Insights from a Multi-Center Study of IEEE ISBI and MICCAI 2021 As for the least creative titles, around 40 title pairs were either exact matches, or differed by only punctuation! Want to see which titles? Try this analysis out and share what you find in the FiftyOne community Slack :) ## Visit Voxel51 at CVPR! Want to discuss the state of computer vision, talk about data-centric AI, or learn how the open source computer vision toolkit [FiftyOne](https://github.com/voxel51/fiftyone) can help you overcome data quality issues and build higher quality models? Come by our booth #1618 at CVPR. Not only will we be available to discuss all things computer vision with you, we’d also simply love to meet fellow members of the CV community, and swag you up with some of our latest and greatest threads. Want to make your dataset easily accessible to the fastest growing community in machine learning and computer vision? Add your dataset to the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html) so anyone can load it with a single line of code. Reach out to me on [Linkedin](https://www.linkedin.com/in/jacob-marks/)! I'd love to discuss how we can work together to help bring your research to a wider audience :) ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,600+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 3,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 4,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 266+ repositories - 58+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [ChatGPT](https://voxel51.com/blog/tag/chatgpt) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [CVF](https://voxel51.com/blog/tag/cvf) [CVPR](https://voxel51.com/blog/tag/cvpr) [IEEE](https://voxel51.com/blog/tag/ieee) [Midjourney](https://voxel51.com/blog/tag/midjourney) [NeRF](https://voxel51.com/blog/tag/nerf) [point clouds](https://voxel51.com/blog/tag/point-clouds) [transformers](https://voxel51.com/blog/tag/transformers) [ViT](https://voxel51.com/blog/tag/vit) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/ff87d65e5b4e5ef50c5732e905f3f16aff8b0a4e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 Survival Guide\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 25, 2023](https://voxel51.com/blog/cvpr-2023-survival-guide) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ec2410f8c8844aa682fea044da0a14e3d9c5c7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ VoxelGPT: Your AI Assistant for Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jun 7, 2023](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/0aa3f8dad8ae1464d05d81ac4a301bd92aea55e3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2024 Survival Guide: Five Vision-Language Papers You Don’t Want to Miss\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Apr 15, 2024](https://voxel51.com/blog/cvpr-2024-survival-guide-five-vision-language-papers-you-dont-want-to-miss) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-25-lllmstxt|> ## Visual AI in Manufacturing [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Industry Solutions](https://voxel51.com/blog/category/industry-solutions) Visual AI in Manufacturing: 2025 Landscape Jul 16, 2025 • 13 min read Article content In this article [Introduction: Back to the Future](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#fb9dc35f2450) [Visual AI’s Impact in Manufacturing — 2025](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#44487720c946) [Real-World Deployments & UseCases](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#ffcc0a4ad31a) [Accessible Visual AI Resources](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#4884b78e64ad) [Unsolved Problems in Visual AIfor Manufacturing](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#6ba8971b3a05) [LeadingCompanies & Voices Driving Innovation](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#ab9a72bbcc54) [The Future of Visual AI in Manufacturing](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#dee5571654c5) [Frequently Asked Questions](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#cf63310fffe8) [Just Wrapping Up!](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#2d83a82ec8f5) [Let’s Connect and Innovate Together!](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#0d33a99bbc4d) [What’s Next?](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#6a4df050e03b) [Author’s Note](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#c449d63e3b2d) In this article [Introduction: Back to the Future](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#fb9dc35f2450) [Visual AI’s Impact in Manufacturing — 2025](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#44487720c946) [Real-World Deployments & UseCases](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#ffcc0a4ad31a) [Accessible Visual AI Resources](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#4884b78e64ad) [Unsolved Problems in Visual AIfor Manufacturing](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#6ba8971b3a05) [LeadingCompanies & Voices Driving Innovation](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#ab9a72bbcc54) [The Future of Visual AI in Manufacturing](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#dee5571654c5) [Frequently Asked Questions](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#cf63310fffe8) [Just Wrapping Up!](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#2d83a82ec8f5) [Let’s Connect and Innovate Together!](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#0d33a99bbc4d) [What’s Next?](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#6a4df050e03b) [Author’s Note](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape#c449d63e3b2d) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Introduction: Back to the Future Back to the Future promised us hoverboards, flying cars, and weather that changed with a click. While we didn’t get flying cars by 2015, something else arrived that truly changed our industrial world: machines that can see, analyze, and learn, powering a new era of manufacturing. I still remember taking a control systems course during my master's program, where we studied backpropagation and predictive maintenance. I was hooked, sitting in front of graphs and signal patterns, trying to catch early signs of system failure. It was like solving a mystery with data. Back then, we were mostly analyzing numbers and waves. Now, that same idea has evolved into something even more powerful, Visual AI. Today, those fault predictions are no longer signal-based; they come from AI models that monitor live video feeds, spotting defects with pixel-level precision, and enabling machines to take action immediately. What I once did with vibration data and math models, factories now do with cameras and AI, faster, smarter, and with more impact. In this blog, I’ll take you through how Visual AI is changing manufacturing in 2025, from the models and datasets driving it to the real-world challenges and the innovators building the future of intelligent production. ![](https://cdn.sanity.io/images/h6toihm1/production/4f67fe7d09a1f24ce7da190faa4e454321a34694-1024x1024.webp?auto=format&dpr=2&fit=max&q=75&w=1024) ## Visual AI’s Impact in Manufacturing — 2025 Artificial Intelligence (AI) is rapidly redefining manufacturing, propelled by the fusion of visual systems, physical AI agents, and cloud-to-edge multimodal intelligence. What was once intractable, like minimizing unplanned downtime, addressing labor shortages, and achieving zero-defect production, is now being addressed through intelligent automation. Manufacturers are deploying and not just experimenting. From predictive maintenance that slashes downtime by up to 50%, to AI-based quality inspection systems that catch defects in milliseconds, modern production floors are powered by data-driven precision. Companies such as Amazon, Siemens, and Foxconn are embedding AI into every layer, leveraging visual data for everything from real-time defect detection and quality assurance in robotic tasks to dynamic route optimization in supply chain orchestration. We’ll explore these and other groundbreaking applications in more detail shortly. ### Sustainable, Scalable, Smart Next-generation manufacturing is faster, more innovative, and more sustainable. AI supports Material efficiency by reducing scrap and overproduction, energy optimization through sensor-informed decision systems, and smart logistics via demand forecasting and dynamic routing. Platforms also increasingly integrate 3D printing and generative design, enabling rapid prototyping with minimal waste. ### Streamlining AI-Driven Production Modern AI manufacturing platforms simplify the journey from design to deployment, enabling: On-demand access to powerful AI models, automated feedback loops between inspection and corrective action, and support for additive manufacturing and sustainability-driven design. As a result, even small and mid-sized factories can deploy advanced intelligence without requiring massive infrastructure. ### Edge AI & IoT: The real-time backbone Edge computing is indispensable for AI to react in real time, to halt a robotic arm during a defect, or recalibrate a welding pattern mid-process. By processing data locally on smart cameras or embedded devices, these systems avoid cloud latency and deliver sub-second response times. Why it matters: Enables low-latency inference at the point of action, preserves data privacy within factory walls, and scales effortlessly across dispersed production sites. Compact AI models and efficient neural networks enable computationally lightweight yet powerful decisions to be made, right where the work occurs, closer to the data source. The foundation of immediate, on-site data processing laid by Edge AI and IoT is pivotal, propelling Visual AI from concept to mission-critical infrastructure across a broad spectrum of manufacturing tasks. ## Real-World Deployments & UseCases Visual AI in manufacturing has evolved into a mission-critical infrastructure across a broad spectrum of tasks. Visual AI in manufacturing has evolved into a mission-critical infrastructure across a broad spectrum of tasks: - **Predictive Maintenance:** [Vision-powered robots](https://www.agilityrobotics.com/industries/manufacturing) and AI systems are increasingly used to identify wear, cracks, and structural anomalies in real time, enabling early intervention and preventing costly breakdowns. Research and industry reports consistently show that AI-based predictive maintenance can reduce unplanned downtime by up to 50% and lower maintenance costs by 20–30% across diverse manufacturing settings ( [MDPI, 2023](https://www.mdpi.com/2227-7390/13/6/981)). - **Smart Quality Assurance:** [Visual AI systems](https://tupl.com/automated-defect-detection-for-manufacturing/) can detect assembly or soldering defects in under 200 milliseconds, enabling real-time corrections that minimize error propagation and reduce rework. Both industry deployments and [academic research](https://arxiv.org/abs/2211.10274) confirm sub-second inspection capabilities in electronics and high-precision manufacturing environments. - **Generative Design Feedback:** Visual AI systems, such as OpenECAD, are becoming increasingly capable of real-time CAD optimization through feedback loops that analyze rendered outputs to inform and refine design decisions. The [CADFusion framework](https://arxiv.org/abs/2501.19054) alternates between code generation and visual rendering to optimize structures, while [Springer research](https://link.springer.com/article/10.1007/s00170-025-15830-2) supports the use of visual data in guiding engineering models. Industry publications, such as Design Engineering and Forbes, highlight how generative AI is reshaping CAD workflows, enabling faster iteration and higher design precision. Additionally, insights from [Novedge](https://novedge.com/blogs/design-news/driving-the-future-ai-enhanced-cad-for-automated-design-optimization) demonstrate AI’s growing role in automated, sustainable design improvements. EVA enables real-time CAD optimization by learning from visual outputs. - **Sustainable Manufacturing & Intelligent Inventory Optimization:** Tesla integrates AI and advanced data analytics across its Gigafactories to optimize battery production, reducing energy consumption, minimizing material waste, and lowering emissions. These systems support Tesla’s broader sustainability goals by improving throughput while reducing environmental impact ( [SupplyChain360](https://supplychain360.io/tesla-supply-chain-big-data-and-ai-in-action/), [arXiv](https://arxiv.org/abs/2307.05521), [ScienceDirect](https://www.sciencedirect.com/science/article/pii/S2950264025000401)). In parallel, Amazon leverages visual and machine learning systems to manage inventory placement, demand forecasting, and logistics routing. Their Supply Chain Optimization Technologies (SCOT) team uses AI to automate product ordering and distribution decisions, resulting in faster delivery and leaner inventory management ( [Amazon Science](https://www.amazon.science/latest-news/the-evolution-of-amazons-inventory-planning-system), [Logistics Viewpoints](https://logisticsviewpoints.com/2025/03/26/amazon-and-the-shift-to-ai-driven-supply-chain-planning/), [Silicon Review](https://thesiliconreview.com/2025/03/amazon-ai-supply-chain-optimization)) - **Additive Manufacturing:** BMW leverages AI-driven additive manufacturing to optimize part geometry, reduce material usage, and streamline prototyping across its production lines. At its dedicated Additive Manufacturing Campus, BMW integrates over 100 industrial 3D printers guided by AI-powered workflows to enhance quality and precision ( [Design Engineering](https://www.design-engineering.com/bmw-launches-additive-manufacturing-campus-to-drive-3d-printing-innovation-1004045462/)). Academic research further indicates that BMW’s approach can achieve a material reduction of up to 60% in automotive applications through topology optimization and simulation-driven design ( [Scientific Publications, 2024](https://thescipub.com/pdf/ajeassp.2024.116.125.pdf)). Meanwhile, the aerospace sector is pushing AI-integrated additive manufacturing even further — applying AI to monitor melt pool stability, detect build defects, and produce high-performance components with real-time feedback ( [Inside Metal AM](https://insidemetaladditivemanufacturing.com/2025/03/21/ai-driven-am-in-aerospace/)). - **Worker Safety Examples Using Visual AI:** Visual AI systems in manufacturing play an essential role in monitoring and enhancing worker safety. From enforcing PPE compliance to real-time fall detection and proximity alerts, these systems utilize smart cameras and AI to identify hazards and mitigate risks. Industry players, such as Toyota Material Handling, use AI to reduce forklift collisions, while startups like [Protex AI](https://voxel51.com/customers/protex-ai) enable the detection of unsafe zones and behaviors. Academically, researchers highlight that AI can anticipate dangers and reduce accidents by up to 30% ( [ResearchGate](https://www.researchgate.net/publication/391988272_THE_GLOBAL_IMPACT_OF_AI_ON_WORKPLACE_SAFETY_OPPORTUNITIES_AND_CHALLENGES_FOR_THE_FUTURE_OF_WORK_The_Global_Impact_of_AI_on_Workplace_Safety_Opportunities_and_Challenges_for_the_Future_of_Work), [MDPI](https://www.mdpi.com/2227-9717/13/5/1312)). ### Critical use cases - **[Foxconn x NVIDIA](https://www.reuters.com/world/china/nvidia-foxconn-talks-deploy-humanoid-robots-houston-ai-server-making-plant-2025-06-20/):** In their Houston AI server factory, humanoid robots trained on vision tasks perform repetitive operations such as cable insertion. - **[Schaeffler & Microsoft](https://www.wired.com/story/ai-swaps-desk-work-for-the-factory-floor/):** Their Factory Operations Agent uses LLMs and vision data to diagnose energy inefficiencies and root-cause anomalies in real-time. - **[Gecko Robotics & Aquant](https://www.businessinsider.com/artificial-intelligence-robotics-predictive-maintenance-manufacturing-factory-solutions-2025-5):** Combine vision and sensor AI to predict failures and conduct preventive maintenance — already serving companies like Coca-Cola, HP, and Siemens. - **[Cobots in Manufacturing](https://www.dobot-robots.com/insights/news/collaborative-robots-in-manufacturing.html):** AI-powered collaborative robots assist with welding, assembly, and materials handling alongside human operators. ## Accessible Visual AI Resources In 2025, one of the most empowering shifts in manufacturing is the growing availability of open-source Visual AI models and datasets. These tools make it easier than ever for engineers, researchers, and practitioners to prototype, test, and deploy innovative vision systems on standard hardware, even a laptop. ![](https://cdn.sanity.io/images/h6toihm1/production/348ee81d6fbd2b3f50b5c0076a64aa9bbae14089-1400x1593.webp?auto=format&dpr=2&fit=max&q=75&w=1400) The table showcases industry-grade solutions and research-backed benchmarks across key domains such as anomaly detection, digital twins, additive manufacturing, and worker safety: - **Anomaly Detection:** Tools like [FADE](https://github.com/BMVC-FADE/BMVC-FADE) enable the detection of surface defects or operational anomalies in a few shots and at a scalable level. These models integrate with leading datasets such as [MVTec AD](https://www.mvtec.com/company/research/datasets/mvtec-ad), [ISP-AD](https://zenodo.org/records/14911043), and [3D-ADAM](https://huggingface.co/datasets/pmchard/3D-ADAM), which offer diverse industrial images with defect annotations. - **Cobots & Humanoids:** Datasets like [RoboMIND](https://huggingface.co/datasets/x-humanoid-robomind/RoboMIND) and [NVIDIA’s GR00T-X](https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim) capture thousands of robot manipulation tasks — from dual-arm coordination to object handling — supporting the development of vision- guided, autonomous cobots for safe and efficient factory work - **Additive Manufacturing:** The growing field of AI-enhanced 3D printing benefits from structured mesh datasets, such as those used by PartCrafter, which are trained on sources like Objaverse and ShapeNet. These datasets support vision-to-CAD workflows, which enhance design precision and reduce material waste. - **Digital Twins:** Platforms like Meta’s [Digital Twin Catalog](https://github.com/facebookresearch/DigitalTwinCatalog) and structured video annotation sets (e.g., RECAST) allow simulation and vision models to interact with virtual manufacturing environments. This enables testing AI in synthetic settings before deploying on the factory floor. - **Worker Safety:** Vision systems for occupational health now leverage datasets like [SH17 PPE](https://github.com/ahmadmughees/SH17dataset), which contain annotated images for detecting personal protective equipment in real-time, as well as [forklift-object\\ \\ detection](https://huggingface.co/datasets/keremberke/forklift-object-detection), which supports collision prevention and safe behavior monitoring. With these resources, teams can move quickly from research to implementation, closing the gap between concept and impact in AI-powered manufacturing. ## Unsolved Problems in Visual AIfor Manufacturing Visual AI holds promise, but challenges persist: - **Robustness in Harsh Environments:** Dust, lighting, and occlusion degrade accuracy. - **Data Scarcity:** Rare failure examples or edge defects limit training. - **Integration with Human Workflows:** Avoiding disruption and promoting collaboration. - **Model Explainability:** Factories require traceable, auditable decisions. - **Environmental Adaptability:** AI must respond to real-world variation in temperature, humidity, and operator behavior. ## LeadingCompanies & Voices Driving Innovation - **Microsoft:** LLM-powered factory agents and analytics. - **Rockwell Automation:** End-to-end operational AI in smart factories. - **Gecko Robotics:** AI vision robots for industrial inspection. - **Aquant:** Predictive service analytics using historical and real-time vision data. - **Tupl:** Visual AI for defect detection and maintenance in manufacturing. ### Voices Driving Innovation **[Stefano Soatto](https://scholar.google.com/citations?hl=en&user=lH1PdF8AAAAJ) (VP of Applied Science at AWS, Professor at UCLA):** A pioneer in dynamic vision and sensor fusion, Soatto’s work underpins perception systems that self-calibrate and map real-world environments — core technologies for visual inspection, robot guidance, and digital twins. He has been featured in Robotics & Automation News. **[Sanja Fidler](https://scholar.google.com/citations?hl=en&user=CUlqK5EAAAAJ) (Director of AI at NVIDIA, Professor at University of Toronto):** Fidler leads NVIDIA’s work on 3D object detection, scene understanding, and multimodal perception — essential components for deploying Visual AI in assembly-line automation, quality inspection, and safety workflows. Her work is widely cited in the literature on visual reasoning and perception. **Mattia Nardon and [Fabio Poiesi](https://www.linkedin.com/in/fabio-poiesi-49101053/) (Researcher, Fondazione Bruno Kessler):** Lead developers of the [ViMAT project](https://arxiv.org/abs/2506.15285), which fuses multi-view video with symbolic task reasoning to track assembly operations in real time. The system verifies procedural correctness without physical markers, enabling flexible, adaptive inspection workflows in factories. [https://tev-fbk.github.io/ViMAT/](https://tev-fbk.github.io/ViMAT/) **[Christos Margadji](https://www.linkedin.com/in/christos-margadji/?originalSubdomain=uk) & [Sebastian W. Pattinson](https://www.sebastianpattinson.com/) (University of Cambridge):** In the [CIPHER framework](https://arxiv.org/abs/2506.08462?utm_source=chatgpt.com), a hybrid vision-language-action system is developed that empowers machines to understand context, perform complex assembly tasks, and explain their decisions, enabling transparent and trusted industrial automation. **[Sassine Ghazi](https://www.linkedin.com/in/sassine-ghazi/) (President & CEO, Synopsys):** Driving chip-to-factory AI vision pipelines through AI-native design automation and edge deployment. Ghazi advocates for silicon-integrated intelligence across the whole manufacturing stack. **[Simon Floyd](https://www.linkedin.com/in/simon-a-floyd/) (Director of Manufacturing & Mobility, Microsoft):** A thought leader in digital twin ecosystems, Floyd promotes the industrial metaverse vision — combining simulation, real-time data, and visual AI to transform factory operations. **[Dayan Rodriguez](https://www.linkedin.com/in/dayanr/) (Principal Specialist, Robotics & AI, AWS):** Focused on factory vision, robotics orchestration, and multi-agent AI deployment, Rodriguez supports the rollout of real-time visual monitoring and automation in smart manufacturing. ## The Future of Visual AI in Manufacturing In the factories of tomorrow, Visual AI won’t just monitor, it will mentor. Systems will alert, diagnose, and suggest, not just act. AI will be the co-pilot, not the pilot. The operator, the engineer, and the technician remain at the center. Together, we’ll build factories that not only produce more, but also produce better. ## Frequently Asked Questions - **How is AI transforming visual inspection on the production line?** AI enables the real-time detection of defects, such as scratches, misprints, and misalignments, replacing slow, manual inspection with 24/7 precision vision systems. This improves quality, reduces rework, and speeds throughput. - **What are the most significant risks in deploying Visual AI at scale?** Common hazards include poor generalization to new environments, lack of high-quality failure data, and over-reliance on AI predictions without human validation. Careful integration with existing workflows and robust validation is key. - **How do companies measure ROI from AI-powered maintenance and QA?** Firms like Aquant and Gecko report reduced service costs (up to 23%), fewer breakdowns, and millions saved in uptime. Metrics include downtime reduction, defect catch rates, and operator trust in AI insights. ## Just Wrapping Up! Visual AI in manufacturing represents a significant leap in technology, and it is also a profound transformation in how we perceive, analyze, and continually improve production. As we embed intelligence into machines, we’re creating a collaborative ecosystem where humans and AI work hand in hand. The operator, the engineer, and the technician remain at the heart of this evolution, their expertise augmented by powerful visual insights. The future of intelligent production is already unfolding. The real question is: _How do we collectively shape it with vision, responsibility, and a shared commitment to building better factories for tomorrow?_ ## Let’s Connect and Innovate Together! Please share your thoughts, ask questions, and provide testimonials. Your insights might help others in our next posts. Don’t forget to participate in the challenge and try out the notebook I’ve created for you all. **Stay Connected:** - Follow me on Medium: [https://medium.com/@paularamos\_phd](https://medium.com/@paularamos_phd) - Follow me on LinkedIn: [https://www.linkedin.com/in/paula-ramos-phd/](https://www.linkedin.com/in/paula-ramos-phd/) - Join the Conversation: [Discord Fiftyone-community](https://voxel51.com/community) ## What’s Next? **Join Us: Meetups & Workshops** We invite you to attend our upcoming Meetups, where we will discuss real-world AI in manufacturing, focusing on actionable insights and practical deployments. Hear from experts, share your perspective, and connect with a growing community of passionate individuals dedicated to intelligent production. And don’t miss our “ **Getting Started with Visual AI in Manufacturing” Workshop!** September 30 — Link TBD, follow this [meetup.com](https://www.meetup.com/pro/ai-machine-learning-data-science-network/) for updates. - Work hands-on with models for quality inspection and predictive maintenance. - Use FiftyOne to explore manufacturing datasets for anomaly detection and assembly verification. - Learn to calculate embeddings, visualize results, and surface key operational insights. ## Author’s Note _While I am not a manufacturing floor operator, I write this piece as an observer and researcher, drawn to the powerful intersection of artificial intelligence and industrial production. The technologies discussed here represent significant progress and complex, high-stakes challenges within the manufacturing sector._ _The most essential principle is simple: we must remain responsible. These tools are not just lines of code; they interact with real processes, real production lines, and real-time decisions. We need to ensure humans stay in the loop, bringing context, expertise, and judgment to every AI-assisted step. The goal isn’t to replace skilled workers, but to empower them to achieve unprecedented levels of ef iciency and innovation._ _Let’s move forward with curiosity, courage, and care._ [manufacturing](https://voxel51.com/blog/tag/manufacturing) [anomaly detection](https://voxel51.com/blog/tag/anomaly-detection) [robotics](https://voxel51.com/blog/tag/robotics) ![](https://cdn.sanity.io/images/h6toihm1/production/e926c07c7d1426c0fde8fdefa637c528d47b16f4-512x512.webp?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/01272b85091efcfd9e9dda5b9044b0f080831781-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Computer Vision Is Changing Manufacturing\\ \\ Industry Solutions, Product & News\\ \\ • \\ \\ Mar 9, 2023](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing) [![](https://cdn.sanity.io/images/h6toihm1/production/a558b86370f2f17212fb2f2c894d590101458a85-5760x3241.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Multimodal Frontier in Computer Vision, Medicine, and Agriculture— CVPR 2025 Reflections\\ \\ Event Recaps, Industry Solutions\\ \\ • \\ \\ Jun 24, 2025](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections) [![](https://cdn.sanity.io/images/h6toihm1/production/8bf3c50a5edd9f83e1013ad5df86ae159d519dae-1200x626.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Best of CVPR 2025: Conversations at the Cutting Edge of AI\\ \\ Event Recaps\\ \\ • \\ \\ Jul 3, 2025](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-26-lllmstxt|> ## Filtered Views Newsletter [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) Voxel51 Filtered Views Newsletter – July 12, 2024 Jul 12, 2024 • 11 min read Article content In this article [📰 The Industry Pulse](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#c566d8054eca) [OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#bb523fffdbb9) [The Robots are Coming for Our Jobs!](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#be71f75ac367) [The Robots Are Going to Help Us with ADHD!](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#b21e590bca32) [💎 GitHub Gems](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#d0a0bbeaf4e1) [📙 Good Reads](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#09c05a951c34) [🎙️ Good Listens](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#ad72f6dba869) [👨🏽‍🔬 Good Research: Is Tokenization the Key to Truly Multimodal Models?](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#1bc7fc63493a) [🗓️. Upcoming Events](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#b3f62c09a97d) In this article [📰 The Industry Pulse](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#c566d8054eca) [OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#bb523fffdbb9) [The Robots are Coming for Our Jobs!](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#be71f75ac367) [The Robots Are Going to Help Us with ADHD!](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#b21e590bca32) [💎 GitHub Gems](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#d0a0bbeaf4e1) [📙 Good Reads](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#09c05a951c34) [🎙️ Good Listens](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#ad72f6dba869) [👨🏽‍🔬 Good Research: Is Tokenization the Key to Truly Multimodal Models?](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#1bc7fc63493a) [🗓️. Upcoming Events](https://voxel51.com/blog/voxel51-filtered-views-newsletter-july-12-2024#b3f62c09a97d) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Welcome to Voxel51's bi-weekly digest of the latest trending AI, machine learning and computer vision news, events and resources! [Subscribe to the email version.](https://voxel51.com/filtered-views-newsletter/) # 📰 The Industry Pulse ## OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding ![](https://cdn.sanity.io/images/h6toihm1/production/8d5fc8d7e52d6611e3d5157bab98116e341ac146-512x341.jpg?auto=format&dpr=2&fit=max&q=75&w=512) [OMG-LLaVA](https://lxtgh.github.io/project/omg_llava/) combines robust pixel-level vision understanding with reasoning abilities in a single end-to-end trained model. It uses a universal segmentation method as the visual encoder to integrate image information, perception priors, and visual prompts into visual tokens provided to a large language model (LLM). The LLM is responsible for understanding the user's text instructions and providing text responses and pixel-level segmentation results based on the visual information. This allows OMG-LLaVA to achieve image-level, object-level, and pixel-level reasoning and understanding, matching or surpassing the performance of specialized methods on multiple benchmarks. What you need to know: - Elegant end-to-end training of one encoder, one decoder, and one LLM rather than using an LLM to connect multiple specialist models - Ability to accept various visual and text prompts for flexible user interaction - Strong performance on image, object, and pixel-level reasoning tasks compared to specialized models Instructions on running the model are [here](https://github.com/lxtGH/OMG-Seg/tree/main/omg_llava), and you can run the demo on Hugging Face Spaces [here](https://huggingface.co/spaces/LXT/OMG_Seg). ## The Robots are Coming for Our Jobs! ![](https://cdn.sanity.io/images/h6toihm1/production/4337345c67ae94e238ff52a009f25258b0b521e1-400x400.jpg?auto=format&dpr=2&fit=max&q=75&w=400) [Source](https://www.therobotreport.com/agility-robotics-digit-humanoid-lands-first-official-job/) Agility Robotics has signed a multi-year deal with GXO Logistics to deploy its Digit humanoid robots in various logistics operations. The first official deployment is already underway at a Spanx facility in Connecticut, where a small fleet of Digit robots is being used under a robotics-as-a-service (RaaS) model. Digit robots pick up totes from 6 River Systems' Chuck autonomous mobile robots (AMRs) and place them onto conveyors at the Spanx facility. Digit can handle empty and full totes and pick them up from an AMR's bottom or top shelf. The robots are orchestrated through Agility Arc, the company's cloud automation platform. GXO Logistics also tests other humanoid robots, such as Apollo from Apptronik. There are currently no safety standards specifically for humanoids. Most manufacturers and integrators are leveraging existing industrial robot standards as a baseline, and Digit is not working with or near humans at the Spanx facility. ## The Robots Are Going to Help Us with ADHD! ![](https://cdn.sanity.io/images/h6toihm1/production/810731f9e7e69d3454c8a7064af94931fd4f272f-512x288.jpg?auto=format&dpr=2&fit=max&q=75&w=512) [Source](https://today.ucsd.edu/story/meet-carmen-a-robot-that-helps-people-with-mild-cognitive-impairment) [CARMEN (Cognitively Assistive Robot for Motivation and Neurorehabilitation)](https://cseweb.ucsd.edu/~lriek/papers/hri2024-bouzida.pdf) is a small, tabletop robot designed to help people with mild cognitive impairment (MCI) learn skills to improve memory, attention, and executive functioning at home. Developed by researchers at UC San Diego in collaboration with clinicians, people with MCI, and their care partners, CARMEN is the only robot that teaches compensatory cognitive strategies to help improve memory and executive function. Here’s what CARMEN is currently capable of: - Delivers simple cognitive training exercises through interactive games and activities - Designed to be used independently without clinician or researcher supervision - Plug and play with limited moving parts and able to function with limited internet access - Communicates clearly with users, expresses compassion and empathy, and provides breaks after challenging tasks In a study, CARMEN was deployed for a week in the homes of several people with MCI and clinicians experienced in working with MCI patients. After using CARMEN, participants with MCI reported trying strategies they previously thought were impossible and finding the robot easy to use. The next steps include: - Deploying CARMEN in more homes. - Enabling conversational abilities while preserving privacy. - Exploring how the robot could assist users with other conditions like ADHD. Many elements of the CARMEN project are [open-source and available on GitHub](https://github.com/UCSD-RHC-Lab/CARMEN). # 💎 GitHub Gems ![](https://cdn.sanity.io/images/h6toihm1/production/66349e6fbce9f562907bfc41b3fa3bc69843d67f-1024x1429.png?auto=format&dpr=2&fit=max&q=75&w=1024) You didn’t think the FiftyOne team would sleep on the Florence2 release, did you? Jacob Marks, OG DevRel at FiftyOne, created the \` [fiftyone\_florence2\_plugin](https://github.com/jacobmarks/fiftyone_florence2_plugin) \` repository on GitHub. This repository is a plugin for integrating the Florence2 model into the FiftyOne open-source computer vision tool. The key components of the plugin include: - Code to load the Florence2 model and generate embeddings and predictions on image data - Integration with FiftyOne to visualize the Florence2 model outputs alongside the image dataset [Here’s a notebook that shows you how to use the plugin!](https://colab.research.google.com/drive/1QGskMcqbbR1hRAtWiTY4Hka7PbmpyQLm?usp=sharing) # **📙** Good Reads This week’s good read is a massive collaborative [three](https://www.oreilly.com/radar/what-we-learned-from-a-year-of-building-with-llms-part-i/) [part](https://www.oreilly.com/radar/what-we-learned-from-a-year-of-building-with-llms-part-ii/) [series](https://www.oreilly.com/radar/what-we-learned-from-a-year-of-building-with-llms-part-iii-strategy/) by some popular AI/ML folks on Twitter titled “What We Learned from a Year of Building with LLMs.” It’s a solid read with some down-to-earth, practical, no-nonsense advice. If you’ve been building with AI/ML for a while, you’ll find that what they say about building with LLMs isn’t too different from what you already know. I feel kinda smart reading this and having many of my thoughts and experiences validated by several people in this space that I admire and consider virtual mentors. Here’s what I think are the best pieces of advice from the series: - Retrieval-augmented generation (RAG) will remain important even with long-context LLMs. Effective retrieval is still needed to select the most relevant information to feed the model. A hybrid approach combining keyword search and vector embeddings tends to work best. - Break complex tasks into step-by-step, multi-turn flows executed in a deterministic way. This can significantly boost performance and reliability compared to a single prompt or non-deterministic AI agent. - Rigorous, continuous evaluation using real data is critical. Have LLMs evaluate each other's outputs, but don't rely on that alone. Regularly review model inputs and outputs yourself to identify failure modes. Design the UX to enable human-in-loop feedback. - Building LLM apps requires diverse roles beyond AI engineers. It is key to hire the right people at the right time, like product managers, UX designers, and domain experts. Focus on the end-to-end process, not just the core LLM. - Center humans in the workflow and use LLMs to enhance productivity rather than replace people entirely. Build AI that supports human capabilities. - Use LLM APIs to validate ideas quickly, but consider self-hosting for more control and savings at scale. Avoid generic LLM features; differentiate your core product. - LLM capabilities are rapidly increasing while costs decrease. Plan for what's infeasible now to become economical soon. Move beyond demos to reliable, scalable products, which takes significant engineering. - The technology is less durable than the system and data flywheel you build around it. Start simple, specialize in memorable UX, and adapt as the tech evolves. A thoughtful, human-centred strategy is essential. Many of the authors recently appeared [on a podcast](https://www.youtube.com/watch?v=c0gcsprsFig) (which I haven’t listened to yet) to discuss the piece and answer questions from the audience. # **🎙️** Good Listens [Aravind Srinivas: Perplexity CEO on the Lex Fridman Podcast](https://lexfridman.com/aravind-srinivas/) ![](https://cdn.sanity.io/images/h6toihm1/production/b94b51ddd28b0f38473e9ad091f71ab735e8182e-512x300.png?auto=format&dpr=2&fit=max&q=75&w=512) I’m a huge fan of [Perplexity.ai](http://perplexity.ai/). Perplexity AI is an "answer engine" that provides direct answers to questions by retrieving relevant information from the web and synthesizing it into a concise response using large language models. Every sentence in the answer includes a citation. It uses a retrieval augmented generation (RAG) approach - retrieving relevant documents for a query, extracting key snippets, and using those to generate an answer. The LLM is constrained only to use information from the retrieved sources. I first heard about it at the beginning of the year, and after using the free tier for two weeks, I realized that it’s a tool worth investing in. I quickly signed up for their “Pro” tier, and it accelerated the pace at which I could conduct research and access knowledge. I was so excited when I saw Perplexity CEO [Aravin Srinivas](https://twitter.com/AravSrinivas) on the Lex Fridman podcast. I’ve only heard him on short-form podcasts, which always left me wanting to hear more from him. In a three-hour conversation (which I’ve listened to twice), Aravind and Lex discussed Perplexity's technical approach, product vision, competitive landscape, and the future of AI and knowledge dissemination on the internet. Here’s some interesting takeaways from this conversation: - Indexing the web involves complex crawling, content extraction, and ranking using traditional methods like BM25 and newer approaches with LLMs and embeddings. Serving answers with low latency at scale is an engineering challenge. - Perplexity has a web crawler called PerplexityBot that decides which URLs and domains to crawl and how frequently. It has to handle JavaScript rendering and respect publisher policies in robots.txt. Building the right index is key. - Perplexity uses a RAG architecture in which, given a query, it retrieves relevant documents and paragraphs and uses those to generate an answer. The key principle is only to say things that can be cited from the retrieved documents. - There are multiple ways hallucinations can occur in the answers—if the model is not skilled enough to understand the query and paragraphs semantically, if the retrieved snippets are poor quality or outdated, or if too much irrelevant information is provided to the model. Improving retrieval quality, snippet freshness, and model reasoning abilities can reduce hallucinations. - Increasing the context window length (e.g., to 100K+ tokens) allows ingesting more detailed pages while answering. However, if too much information confuses the model, there are tradeoffs with the instruction-following performance. - By incorporating human feedback into the training process via RLHF, Perplexity wants to create an AI knowledge assistant that provides high-quality, relevant answers to user queries. The goal is for the AI to understand the user's intent and give them the information they seek. Here are a few clips from the conversation that you might find insightful: - [How Perplexity works](https://www.youtube.com/watch?v=Q0ncaAwnn-o) - [How web crawlers work](https://www.youtube.com/watch?v=Ci6N3ghLlr0) - [Simple advice for writing academic papers](https://www.youtube.com/watch?v=-NnCWB5EPjk) # 👨🏽‍🔬 Good Research: Is Tokenization the Key to Truly Multimodal Models? ![](https://cdn.sanity.io/images/h6toihm1/production/179b28bc5501122faf547b0b9e03fdf9a913d28b-512x288.gif?auto=format&dpr=2&fit=max&q=75&w=512) A lot of what's being hyped as a "multimodal model" these days is basically just vision-language models, which really aren't multimodal because they're just two modalities. While these models have yielded impressive results and serve as a foundation for more sophisticated architectures, they're basically some type of Frankenstein monster. You glue together a pretrained vision encoder and text encoder, freeze the vision encoder, and let gradients flow through the text encoder during training. Don't get me wrong, VLMs are an important step toward more comprehensive multimodal AI, but this architectural choice has limitations in fully integrating the modalities. I don’t mean to downplay our progress so far, and I fully appreciate how difficult it is to unify diverse modalities into a single model. Modalities be all over the place with their dimensionality, types, and values. Images are typically represented as high-dimensional tensors with spatial relationships. In contrast, text is represented as variable-length sequences of discrete tokens. Structured data like vectors and poses have unique formats and characteristics. Feature maps, intermediate representations learned by neural networks, add another layer of complexity to multimodal learning. Recent approaches to multimodal learning often rely on separate encoders for each modality, such as vision transformers for images and large language models for text. Or, ImageBind, able to encode six modalities, projects four modalities into a frozen CLIP embedding space, essentially aligning them with the vision and language modalities. While these specialized encoders can effectively process their respective modalities, they create a bottleneck when fusing information across modalities, as they lack a common representation space. However, new research from Apple might just change the way we architect multimodal networks. The [_4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities_](https://arxiv.org/abs/2312.06647) introduces an any-to-any model that can handle 21 modalities across the following categories: RGB, geometric, semantic, edges, feature maps, metadata, and text modalities. Check out the demo [here](https://huggingface.co/spaces/EPFL-VILAB/4M). And the big insight in the paper: It all comes down to tokenization. ![](https://cdn.sanity.io/images/h6toihm1/production/8d5fc8d7e52d6611e3d5157bab98116e341ac146-512x341.jpg?auto=format&dpr=2&fit=max&q=75&w=512) To address the challenge of unifying diverse modalities, the 4M-21 model introduces modality-specific tokenization schemes that convert each data type into a sequence of discrete tokens. This tokenization scheme unifies the representation of various modalities into a common space of discrete tokens, allowing a single model to handle all of them with the same architecture and training objective. And, what I think is the coolest part about tokenizing in this way is now all tasks that the model can handle are formulated as a per-token classification problem, which can be trained with a cross-entropy loss using an encoder-decoder based transformer. I want to focus on the tokenization for this post, but [I encourage you to check out the project page](https://4m.epfl.ch/) to learn more. Here are my Cliff's Notes on the tokenizer: - For image-like modalities such as RGB, surface normals, depth, and feature maps from models like CLIP, DINOv2, and ImageBind, they used Transformer-based VQ-VAE tokenizers. These tokenizers compress the dense, high-dimensional image data into a smaller grid of discrete tokens (e.g., 14x14 or 16x16) while preserving the spatial structure. They also used a diffusion decoder for edges to generate more visually plausible reconstructions. The autoencoders learn to encode spatial patches of an image into discrete tokens, effectively capturing local patterns and spatial structure. The VQ-VAEs can map similar patches to the same token using a discrete latent space, providing a compact and semantically meaningful representation. - For non-image-like modalities such as DINOv2 and ImageBind global embeddings and 3D human poses, they employed MLP-based discrete VAEs with Memcodes quantization. This allows compressing the vectors or pose parameters into a small set of discrete tokens (e.g., 16) without imposing any spatial structure. . These autoencoders learn to map continuous inputs to discrete latent variables, which the shared transformer architecture can then process. By discretizing the continuous data, the model can more effectively capture and reason about their underlying structure and relationships. - For text-based modalities like captions, object bounding boxes, image metadata, and color palettes, they utilized a shared WordPiece tokenizer with special token prefixes to encode the type and value of each data field. This tokenizer breaks down words into subword units, allowing the model to handle out-of-vocabulary words and maintain a fixed vocabulary size. Using a shared vocabulary across all modalities, the WordPiece tokenizer enables the model to learn cross-modal associations and alignments. The modality-specific tokenization schemes in 4M-21 seems promising. It provides a common representation space for all modalities, enabling the model to process them using a shared architecture. It also preserves modality-specific information, such as spatial structure in images and semantic meaning in text, which is crucial for effective multimodal learning and reasoning. Finally, converting all data types into sequences of discrete tokens enables cross-modal interactions and attention mechanisms, allowing different modalities to communicate and exchange information. # 🗓️. Upcoming Events Check out these upcoming AI, machine learning and computer vision events! [View the full calendar and register for an event.](https://voxel51.com/computer-vision-events/) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [datasets](https://voxel51.com/blog/tag/datasets) [Filtered Views](https://voxel51.com/blog/tag/filtered-views) [machine learning](https://voxel51.com/blog/tag/machine-learning) [newsletter](https://voxel51.com/blog/tag/newsletter) ![](https://cdn.sanity.io/images/h6toihm1/production/a41a0477c7a98264f600772e9568607d070eea59-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Sahota Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Loading related posts... [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-27-lllmstxt|> ## Dataset Insights and Optimization [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # FiftyOneCuration Get insights into your dataset's distribution, diversity, coverage, and more to optimize AI performance. Analyze billions of samples, hosted securely on your infrastructure, whether in the cloud or on-premise. [Book a demo](https://voxel51.com/sales) [View docs](https://docs.voxel51.com/recipes/index.html) ![](https://cdn.sanity.io/images/h6toihm1/production/20d6894f28b427565b36460217ec59ef40c423fb-1200x1200.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1200,1200&w=350) ![](https://cdn.sanity.io/images/h6toihm1/production/2b7a7b9a4f28c2b7df210382ebe8065abdd2ee2a-1200x1200.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1200,1200&w=350) Data + Models ## Better data leads to better models 80% of AI projects fail due to data issues. FiftyOne allows you to systematically inspect, curate, and analyze your datasets to improve model outcomes. Diversity CoverageDistributionBalanceDeduplicationLabel Accuracy ### Diversity Diverse data helps models generalize effectively across real-world conditions and avoid overfitting. Use similarity search and clustering tools to ensure your dataset captures varied, distinct examples. [Leverage vector search](https://voxel51.com/blog/the-computer-vision-interface-for-vector-search) Multimodal Data ## Get support for multimodal data visualization Visualize images, video, 3D point clouds, geospatial, medical scans, and audio data in an interactive UI. [Learn more](https://docs.voxel51.com/user_guide/groups.html) Compliance & Governance ## Enterprise-grade security, scale, and compliance ### Robust dataset versioning Snapshots systematically maintain different dataset iterations, ensuring traceability and reproducibility. Schema, tags, and metadata are tracked at the dataset and sample level. [Learn more](https://docs.voxel51.com/enterprise/dataset_versioning.html) ![](https://cdn.sanity.io/images/h6toihm1/production/5f846adc7feaa6c2b079f330b7eea37cc84df000-2552x1508.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,2552,1508&w=640) ### Role-based access controls Maintain compliance with configurable roles and fine-grained permissions, enabling secure collaboration across the organization. [Learn more](https://docs.voxel51.com/enterprise/roles_and_permissions.html) ![](https://cdn.sanity.io/images/h6toihm1/production/177c8df384a1b80fb3486a758be30134127e43dc-2552x1508.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,2552,1508&w=640) ML Workflows ## Out-of-the-box workflows for machine learning pipelines ### Streamline visual data discovery Stop waiting days for data teams to deliver samples. Query your data lake and retrieve relevant samples in seconds using Data Lens. [Explore Data Lens](https://voxel51.com/blog/streamline-visual-data-discovery-with-fiftyone-data-lens) ![](https://cdn.sanity.io/images/h6toihm1/production/1c3615991f851f28bbec9f58454a54dd117e8b7e-2226x2226.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,2226,2226&w=600) ### Improve data quality Pinpoint critical dataset issues that quietly sabotage model performance. By enabling rapid diagnosis and targeted data improvements, FiftyOne ensures your models generalize reliably across real-world conditions. [Why data quality matters](https://voxel51.com/blog/data-quality-the-hidden-driver-of-ai-success) ![](https://cdn.sanity.io/images/h6toihm1/production/a6f1a04da69404020c28f9ba6caf3934148863b4-2226x2226.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,2226,2226&w=600) ### Avoid model drift Model drift can degrade performance in production environments. Get active learning workflows to systematically monitor, identify, and correct dataset shifts, ensuring consistent model performance over time. [Explore model evaluation](https://voxel51.com/evaluation) ### Continuously improve model accuracy Model training is never truly complete. With FiftyOne's continuous analysis and embedding-based diagnostics, you can quickly discover and address weaknesses in your data, refining your model iteratively for maximum accuracy. [Explore embeddings](https://docs.voxel51.com/tutorials/image_embeddings.html?highlight=embeddings) ## Questions? We have answers. ### What is data curation, and why is it crucial for computer vision projects? ### How do I choose the best data curation solution for my computer vision project? ### How can data curation improve my machine learning model performance? ### What tools does Voxel51 offer to streamline the data curation process? ## Enough data wrangling.
 Request a demo. [Book a demo](https://voxel51.com/sales) [Explore Verified Auto Labeling](https://voxel51.com/annotation) ![](https://cdn.sanity.io/images/h6toihm1/production/ea42e9b26f49f1cb54bb8aca31dc10e7f74fe11f-3024x960.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-28-lllmstxt|> ## AI Solutions for Manufacturing [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Visual AI in Manufacturing Automation requires reliable execution, even within challenging environments. Developers of cutting-edge machine vision systems turn to FiftyOne to build high-quality datasets and high-performing models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/75818e38d46aba80bffb63ab384017322743ba2f-628x395.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/c2f18551ed102d5dd5e007d090d3a4d9ab619333-628x817.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/317da9fa1e14bde40b62a6b0b91ec0c7ccff153c-628x813.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/25e160cb488c8c60a6d2754a30980832a990b87a-628x393.png?auto=format&dpr=2&fit=max&q=75&w=314) ## Power next generation of intelligent manufacuring with confidence AI isn’t just automating factory tasks—it’s reshaping how decisions are made, errors are caught, and production lines stay agile. ## 78% of manufacturers plan to increase AI investment over the next two years ## 46% of manufacturers are already using generative AI tools ## Visual AI in manufacturing events From industry meetups to 90-minute technical workshops, join our virtual events led by our ML experts. [Register now](https://voxel51.com/events?industry=manufacturing) ![](https://cdn.sanity.io/images/h6toihm1/production/8a1ed58c9faebaa73860b791f079095ab671c65a-840x473.jpg?auto=format&dpr=2&fit=max&q=75&w=420) Use cases ## Manufacturing AI use cases powered by FiftyOne Computer vision plays a critical role in a wide variety of manufacturing use cases. That’s why leaders and innovators building AI/ML solutions for manufacturing rely on FiftyOne. ![](https://cdn.sanity.io/images/h6toihm1/production/b1ca1bf21087414657f981d4ffb45d2481df5e8d-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Defect detection Ensure quality control in industrial processes using computer vision to rapidly detect and identify defects. ![](https://cdn.sanity.io/images/h6toihm1/production/bb964b17136f1a799740b9866f1969b8a7d3707b-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Predictive maintenance Save on maintenance costs by using computer vision and predictive analytics to find and fix potential issues before they result in failures. ![](https://cdn.sanity.io/images/h6toihm1/production/eb297bebbf1cafdcbe64d6cfebb70266b1c1d392-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Bin picking Make automatic bin picking solutions robust and reliable by using computer vision to map the environment and guide robotic arms. ![](https://cdn.sanity.io/images/h6toihm1/production/b9a2cf3b29d778d0982e3b7687b461405c982d39-768x513.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Machine tending Bring high precision to the loading and preparation of raw materials processed by machines. ![](https://cdn.sanity.io/images/h6toihm1/production/c339f927d68526e4555bfae43f0c38468fee6910-768x432.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Palletizing and depalletizing Enable and enhance palletizing and depalletizing by building and training object detection models to provide very high accuracy. ![](https://cdn.sanity.io/images/h6toihm1/production/197bce30bf96be8a6083a2dc7e0bc7f73e8a3d37-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Safety monitoring Revolutionize real-time hazard detection and safety compliance monitoring across manufacturing sites, enhancing worker protection and reducing accidents. Features ## How visual AI can help you Unify multimodal data Scene reconstructionManage workflowsData versioning ### Manage millions of multimodal samples through a unified interface Machine learning engineers spend over half of their time wrangling data, but it doesn’t have to be that way. Use FiftyOne’s powerful dataset import and manipulation capabilities to manage millions of diagnostic images and videos. ### Catch defects before they cost you Surface edge cases, failure modes, and anomalies faster. Prevent defects from slipping through—and from training into your models. ### Boost production Minimize unplanned downtime with smarter data workflows. FiftyOne enables rapid diagnosis and debugging so models stay performant, and production lines keep running without interruption. ### Save time on what matters Automate labeling, surface issues instantly, and cut repetitive visual data tasks in half. FiftyOne frees up your ML team to focus on improving models—not just cleaning up after them. ## Get started today Leading enterprises build using FiftyOne 30% increase in model accuracy 5+ months of development time saved 30% boost in team productivity [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/5691fe278acf8d79c2bb71f6da17f600de33675d-6048x1921.png?auto=format&dpr=2&fit=max&q=75&w=3024) resources ## Learn more about visual AI in manufacturing Get the latest industry news from our ML experts. ### Visual AI for Defect Detection: 4 Common Failures and How to Avoid Them [Download the Paper](https://voxel51.com/whitepapers/visual-ai-for-defect-detection-in-manufacturing) ![](https://cdn.sanity.io/images/h6toihm1/production/2b9cf0dc8ac97d8cfb20b412b7eec7435249ea1f-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=640) ### Visual AI in healthcare blogs [View all](https://voxel51.com/blog/tag/manufacturing) ![](https://cdn.sanity.io/images/h6toihm1/production/1bd4596e43fc708848051c57f484f07ce5f56ef1-1276x754.png?auto=format&dpr=2&fit=crop&fp-x=0.222&fp-y=0.5&h=71&q=75&rect=552,0,144,71&w=71) ![](https://cdn.sanity.io/images/h6toihm1/production/5691fe278acf8d79c2bb71f6da17f600de33675d-6048x1921.png?auto=format&dpr=2&fit=max&q=75&rect=2387,0,3661,1921&w=640) ### Trusted by ML experts [Read customer stories](https://voxel51.com/customers?category=av-physical-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/5691fe278acf8d79c2bb71f6da17f600de33675d-6048x1921.png?auto=format&dpr=2&fit=max&q=75&rect=2387,0,3661,1921&w=640) ![](https://cdn.sanity.io/images/h6toihm1/production/1bd4596e43fc708848051c57f484f07ce5f56ef1-1276x754.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=38&q=75&rect=0,0,64,38&w=38) customer stories ## With FiftyOne, RIOS organized 20TB + multimodal data “Everything we do from the MLOps side interacts with FiftyOne Teams. It’s becoming the hub where all the spokes are connected. _It’s like GitHub for code, but for our datasets._” **-** **Matt Shaffer,** VP of AI & Co-founder @ RIOS Intelligent Machines [Read the story](https://voxel51.com/customers/rios) ![](https://cdn.sanity.io/images/h6toihm1/production/4af6548e5b08c7ee74e983d2995d80ea72fa5ca8-608x609.png?auto=format&dpr=2&fit=max&q=75&rect=0,124,608,360&w=304) resources ## Developer resources ### Datasets - Try on FiftyOne: [MVTecAD](https://try.fiftyone.ai/datasets/mvtec-ad/samples) - Try on FiftyOne: [MVTecAD2](https://demo.fiftyone.ai/datasets/mvtecad2/samples) ### Models - [Anomalib](https://github.com/open-edge-platform/anomalib) ### Tutorials - [Anomaly detection with Anomalib](https://docs.voxel51.com/tutorials/anomaly_detection.html?highlight=anomalib) - [Getting started with FiftyOne](https://github.com/paularamo/awesome-fiftyone/tree/main/getting-started-90min-workshop) ## Data eats models for lunch Talk to our computer vision experts to start building better datasets and models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-29-lllmstxt|> ## RIOS Robotics Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/b225076935c6ada1d3969623e54240936c13b0dc-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) RIOS RIOS’s AI-Powered Robotics Solutions Run on FiftyOne Enterprise May 2, 2025 RIOS Intelligent Machines automated their robotics data engineering pipeline by implementing FiftyOne as the central hub for their MLOps processes. Efficiently organised data of 20 TB+ Customer objects handled Millions Repetitive data transformations Zero Article content In this article [Success story at a glance](https://voxel51.com/customers/rios#5c6f39150bc6) [Introduction](https://voxel51.com/customers/rios#d85349a0094d) [Challenge](https://voxel51.com/customers/rios#b402f966ea76) [Solution](https://voxel51.com/customers/rios#5b2aaaacabb5) [The hub that connects the MLOps spokes](https://voxel51.com/customers/rios#6d7b45a00ab0) [Automation and efficiency across the board](https://voxel51.com/customers/rios#d59d9739e2a1) [Easy scene reconstructions with grouped datasets](https://voxel51.com/customers/rios#a7b05cf10d5b) [Adding real value to synthetic data](https://voxel51.com/customers/rios#caf68b05de07) [Continuously train and deploy models that perform](https://voxel51.com/customers/rios#a3fef29ec370) [Collaborating with outside annotators](https://voxel51.com/customers/rios#684abac83623) [Unparalleled support & documentation](https://voxel51.com/customers/rios#aaf4b12d5ed9) [Results](https://voxel51.com/customers/rios#7d08e5ebaad1) In this article [Success story at a glance](https://voxel51.com/customers/rios#5c6f39150bc6) [Introduction](https://voxel51.com/customers/rios#d85349a0094d) [Challenge](https://voxel51.com/customers/rios#b402f966ea76) [Solution](https://voxel51.com/customers/rios#5b2aaaacabb5) [The hub that connects the MLOps spokes](https://voxel51.com/customers/rios#6d7b45a00ab0) [Automation and efficiency across the board](https://voxel51.com/customers/rios#d59d9739e2a1) [Easy scene reconstructions with grouped datasets](https://voxel51.com/customers/rios#a7b05cf10d5b) [Adding real value to synthetic data](https://voxel51.com/customers/rios#caf68b05de07) [Continuously train and deploy models that perform](https://voxel51.com/customers/rios#a3fef29ec370) [Collaborating with outside annotators](https://voxel51.com/customers/rios#684abac83623) [Unparalleled support & documentation](https://voxel51.com/customers/rios#aaf4b12d5ed9) [Results](https://voxel51.com/customers/rios#7d08e5ebaad1) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) # Success story at a glance **Company:** - Name: RIOS Intelligent Machines - Industry: Robotics - Data capture: Multiple cameras in robot environments, including 360-degree camera views of the scene in which the robot operates - Data sources: Robots from customer sites, headquarters, and synthetic data - Data types: Images, videos, and 3D data **Challenge:** - RIOS supports a variety of robotic solutions deployed across a number of different customer environments - RIOS needed a dataset management solution to efficiently organize and visualize its datasets, with the flexibility to plug into other components in the ML pipeline - As an end-to-end automation company, RIOS wanted to eliminate unnecessary manual work and automate as much of the data engineering pipeline as possible **Solution:** - FiftyOne Enterprise helps RIOS efficiently organize and analyze 20TB+ of data, like GitHub for code, but for datasets - FiftyOne Enterprise seamlessly plugs into RIOS’s ML experiment tracker, cloud data lake, and annotation tools, enabling a slew of automated ML workflows - Data automation eliminates repetitive data transformations, freeing up valuable engineering time - Continuous model improvements support RIOS’s 99.6%+ uptime guarantee ## Introduction [RIOS Intelligent Machines](https://rios.ai/) is a game changer in the automation industry, helping enterprises automate their factories, warehouses, and supply chain operations while increasing production and eliminating defects. RIOS accomplishes this through robust, reliable, and flexible AI-powered robotic solutions that seamlessly adapt to production requirement changes and AI-powered vision for monitoring operations, quality assurance, and handling complex operations previously not possible with mechanical robotic automation. ## Challenge Training different machine learning models on data from robots across customer environments, such as packaged food products, beverage distribution, and wood products, requires subtle changes in the data schema between each use case. If data transformations are done manually, a significant amount of data engineering effort is spent on turning raw data into usable data and then exporting it into a format compatible with model training. Having a central dashboard to organize and visualize datasets and search and sort them would free up valuable engineering time to continue building the robotics and AI solutions that RIOS customers have come to know and love. ## Solution Complex manipulation tasks require 3D scene understanding, and RIOS’s ML team required a dataset management solution with a flexible metadata structure to add custom fields such as camera pose, object pose, and other 3D metadata. Equally important was the ability to export to common data types like YOLO and COCO for downstream model training tasks. Additionally, because RIOS already had an ML experiment tracking tool, annotation tools, and cloud data storage, it was critical to have a data management tool to plug into and supplement them. Enter [FiftyOne Enterprise](https://voxel51.com/). _Not yet familiar with FiftyOne Enterprise? It’s where real AI work happens. FiftyOne Enterprise helps you visualize, augment, manage, and QA data; it also helps you streamline the workflows that make enterprise machine learning possible. FiftyOne Enterprise extends open-source FiftyOne with added features enterprises need, including dataset permissions, versioning, sharing, support, and more._ ## The hub that connects the MLOps spokes RIOS selected FiftyOne Enterprise as its central hub for datasets and integrated its other MLOps activities – annotation, ML experiment tracking, and synthetic data generation – into the FiftyOne platform. RIOS’s data comes in from various sources, including data collected from robots at headquarters, data from robots in the field, and synthetic data, and is put into a cloud data lake. Once datasets are collected, they are processed and loaded into FiftyOne. From there, ML engineers can kick off processes for training, model analysis, annotations, simulations, customer reviews, or anything they need. “Everything we do from the MLOps side interacts with FiftyOne Enterprise. It’s becoming the hub where all the spokes are connected. It’s like GitHub for code, but for our datasets,” said Matt Shaffer, VP of Artificial Intelligence and Co-founder at RIOS. ## Automation and efficiency across the board RIOS is always looking to automate as much as possible, including its own internal ML pipelines. RIOS’s ML team has built processes that understand that data as it’s flowing into the data lake and automatically put it into the FiftyOne format so it’s immediately ready to work with. “The FiftyOne Enterprise API is clean and clear, making it easy to add fields and create samples in grouped datasets. One example is when I created a custom script to load in and render an annotated synthetic video dataset, which was very quick because of FiftyOne,” said Joshua Zorn, Senior Synthetic Data Developer at RIOS. In addition, the ML team makes heavy use of FiftyOne’s [tagging feature](https://docs.voxel51.com/user_guide/app.html#tags-and-tagging) for efficiency. Being able to tag data makes it easy to find the exact data you’re looking for. RIOS tags data by customer, by object, and annotation status so that they have total transparency and visibility into which customer the dataset belongs to, the objects in the dataset, and what stage it’s in in terms of processing. ## Easy scene reconstructions with grouped datasets Working with robots often involves working with data from a 360-degree imaging system, with images from all four sides of an item and sometimes the top and bottom. With FiftyOne, RIOS can group all of those samples together to visualize a reconstruction of the scene. Using the [grouped dataset](https://docs.voxel51.com/user_guide/groups.html) feature of FiftyOne is now a common way for the RIOS team to interact with and view their data. ## Adding real value to synthetic data RIOS uses synthetic data, digital twins, and simulations to rapidly innovate and test new solutions for customer environments while sidestepping some of the constraints around prototyping in the real world, such as shipping times. FiftyOne Enterprise sits at the heart of RIOS’s [synthetic data generation and virtual work cell pipeline](https://www.nvidia.com/en-us/on-demand/session/gtcspring23-s51494/). “We use FiftyOne Enterprise for our dataset management system, which allows us to visualize our synthetic data, organize it, preview it, and ensure that we can work efficiently with it,” explained Joshua. ## Continuously train and deploy models that perform RIOS relies on FiftyOne Enterprise to visualize, analyze, and [evaluate how their models perform](https://docs.voxel51.com/user_guide/evaluation.html) on training data and how they perform on test data from production environments to gain critical insights from training through production. “On one side, we use FiftyOne Enterprise during model training to prep and select the data, then inspect the model’s performance on that data. On the other side, we take the inference results from our production models and test data to visualize and evaluate them in FiftyOne Enterprise. Doing this means we have an end-to-end process that closes the entire loop for us on model analysis,” explained Shubham Kanitkar, Sr. Robotics ML Engineer at RIOS. “Loading the inference results from production models back into FiftyOne Enterprise enables anyone on the ML team to review them to see whether or not they thought the model was performing well. Because of the rich visualizations provided by FiftyOne Enterprise, it’s possible to review the samples quickly and efficiently for insights,” said Matt. ## Collaborating with outside annotators FiftyOne Enterprises’ [roles and permissions](https://docs.voxel51.com/teams/roles_and_permissions.html) set the stage for easy, streamlined collaboration with annotators from the customer’s side. “In the past, we’ve collaborated with customers through self-hosted annotation platforms such as CVAT. This made it difficult to collaborate with outside organizations due to VPN policies and data privacy issues. Using FiftyOne Enterprise as a customer-facing app makes it easier for us to share data with customers to show them the quality of that annotation run,” described Shubham. ## Unparalleled support & documentation RIOS appreciates the quality of support that comes with FiftyOne Enterprise. “I think of Voxel51’s customer success team as an extension of our team. We can work together on solving real-world ML problems and partner in cases where we’re requesting new features in the FiftyOne platform. It’s been a good working relationship for us,” explained Matt. The RIOS team finds the [documentation](https://docs.voxel51.com/) useful, too. “Usually documentation goes unnoticed, but with FiftyOne, the documentation is killer. Having really good documentation is very important. Knowing that FiftyOne is well documented is much appreciated,” explained Shubham. ## Results - A growing data lake of 20TB+ of data supporting millions of customer objects handled - Elimination of repetitive data transformations, freeing up valuable engineering time - Continuously improve model performance to support a 99.6%+ uptime guarantee for customers - Greater confidence in the quality of their datasets ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-30-lllmstxt|> ## METU Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/cd76dddd6d059fdc96846aa6a6cab58ff156b0c1-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) METU METU uses FiftyOne to conduct medical research on inflammatory bowel disease Apr 5, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) The [Deep Learning and Computer Vision group](http://dlcv.ii.metu.edu.tr/) at [METU](https://www.metu.edu.tr/) conducts research on the cutting-edge topics of deep learning and computer vision. Their research is focused on object detection, adversarial attacks, medical imaging, 3D model generation, generative models, and GPU programming. # METU’s success story with FiftyOne METU collaborates with the Department of Gastroenterology at Marmara University, Turkey to conduct medical research on inflammatory bowel disease (IBD). One of the main challenges in IBD diagnosis is the high interobserver and intraobserver variability among the assessments of the doctors. With the aid of deep learning, real-time and objective feedback can be provided to doctors during colonoscopy operations. METU uses FiftyOne at various stages of research. First, they make use of FiftyOne’s CVAT integration to request and load annotations from doctors. Next, they compare doctors’ annotations and their models’ predictions. Then, using FiftyOne Brain’s embedding space visualization tools (t-SNE specifically), they get a better understanding of their data distribution. In a branch of this research, they produce synthetic images with GANs and again, refer to FiftyOne Brain tools for help identifying the most informative/unique images by eliminating “near duplicates”. In the coming months, doctors will start sending videos and METU will start to make use of FiftyOne’s video analysis capabilities. ![](https://cdn.sanity.io/images/h6toihm1/production/85584c2e3962ea8f677613d094ee0a7084c922c2-2017x1340.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) METU uses FiftyOne for cutting-edge deep learning and computer vision research - Voxel51 <|firecrawl-page-31-lllmstxt|> ## Updata Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/2889f465e725d668ea8151e261ccc307a60dcb62-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Updata FiftyOne gives Updata a powerful and flexible framework for ML Apr 21, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Updata](https://updata.ca/) builds AI that works. Their solutions help companies make more money while being kinder to the planet. From strategy to deployment, Updata handles the entire AI lifecycle—whether you need smarter operations, sharper decisions, or AI-powered products. They have transformed businesses in recycling, energy, construction, and food processing. And they don’t just deliver results – they teach your team to maintain and evolve these systems long after they’re gone. > "Updata leverages FiftyOne as an essential component to efficiently sample and analyze the behavior and performance of our computer vision models. Its intuitive UI enables us to effectively demonstrate to clients how machine learning contributes to enhanced manufacturing processes and sustainability initiatives. Additionally, FiftyOne's Python SDK and their innovative concept of a visual dataframe have significantly strengthened our tooling, providing a more powerful and flexible framework for ML experimentation." – David Cardozo, Chief Lead Analyst at Updata ![](https://cdn.sanity.io/images/h6toihm1/production/194ac4e0ee982c62c51acb869d473686a917330b-1344x896.png?auto=format&dpr=2&fit=max&q=75&w=1344) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-32-lllmstxt|> ## Visual AI Healthcare Event [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/4dbe6c37a9ce253eaf07d67643e7feaae0522d45-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/4dbe6c37a9ce253eaf07d67643e7feaae0522d45-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Meetups Healthcare Visual AI in Healthcare – June 27, 2025 This event has ended, but you can still catch up! Watch the on-demand recordings and register for our [future events.](https://voxel51.com/events) Jun 27, 2025 9 AM Pacific Virtually over Zoom! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/e623140494dd0c652c420b0253d977d3ef88b30e-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Aswin Kumar Stanford Bio ![](https://cdn.sanity.io/images/h6toihm1/production/d1873889972e40c637643e56e628ea454af3e3a9-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Maya Varma Stanford Bio ![](https://cdn.sanity.io/images/h6toihm1/production/0e43c562fddce0567e1a72419ae05034e90a976b-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Heather (Dunlop) Couture PixelScientia Bio ![](https://cdn.sanity.io/images/h6toihm1/production/135511c73ad886d388184fcca428eea1bdf99ad0-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Maximilian Rokuss DKFZ German Cancer Research Center Bio ![](https://cdn.sanity.io/images/h6toihm1/production/b364057f3bb6cc61d113b7ea1ab6c47da07f853a-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Gaurav K Gupta Lake County Health Department Bio About this event Hear talks from experts on cutting-edge topics at the intersection of AI, ML, computer vision and healthcare. Schedule MedVAE: Efficient Automated Interpretation of Medical Images with Large-Scale Generalizable Autoencoders ![](https://cdn.sanity.io/images/h6toihm1/production/e623140494dd0c652c420b0253d977d3ef88b30e-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Aswin Kumar Stanford Bio ![](https://cdn.sanity.io/images/h6toihm1/production/d1873889972e40c637643e56e628ea454af3e3a9-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Maya Varma Stanford Bio We present MedVAE, a family of six generalizable 2D and 3D variational autoencoders trained on over one million images from 19 open-source medical imaging datasets using a novel two-stage training strategy. MedVAE downsizes high-dimensional medical images into compact latent representations, reducing storage by up to 512× and accelerating downstream tasks by up to 70× while preserving clinically relevant features. We demonstrate across 20 evaluation tasks that these latent representations can replace high-resolution images in computer-aided diagnosis pipelines without compromising performance. MedVAE is open-source with a streamlined finetuning pipeline and inference engine, enabling scalable model development in resource-constrained medical imaging settings. Leveraging Foundation Models for Pathology: Progress and Pitfalls ![](https://cdn.sanity.io/images/h6toihm1/production/0e43c562fddce0567e1a72419ae05034e90a976b-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Heather (Dunlop) Couture PixelScientia Bio How do you train ML models on pathology slides that are thousands of times larger than standard images? Foundation models offer a breakthrough approach to these gigapixel-scale challenges. This talk explores how self-supervised foundation models trained on broad histopathology datasets are transforming computational pathology. We'll examine their progress in handling weakly-supervised learning, managing tissue preparation variations, and enabling rapid prototyping with minimal labeled examples. However, significant challenges remain: increasing computational demands, the potential for bias, and questions about generalizability across diverse populations. This talk will offer a balanced perspective to help separate foundation model hype from genuine clinical value. LesionLocator: Zero-Shot Universal Tumor Segmentation and Tracking in 3D Whole-Body Imaging ![](https://cdn.sanity.io/images/h6toihm1/production/135511c73ad886d388184fcca428eea1bdf99ad0-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Maximilian Rokuss DKFZ German Cancer Research Center Bio Recent advances in promptable segmentation have transformed medical imaging workflows, yet most existing models are constrained to static 2D or 3D applications. This talk presents LesionLocator, the first end-to-end framework for universal 4D lesion segmentation and tracking using dense spatial prompts. The system enables zero-shot tumor analysis across whole-body 3D scans and multiple timepoints, propagating a single user prompt through longitudinal follow-ups to segment and track lesion progression. Trained on over 23,000 annotated scans and supplemented with a synthetic time-series dataset, LesionLocator achieves human-level performance in segmentation and outperforms state-of-the-art baselines in longitudinal tracking tasks. The presentation also highlights advances in 3D interactive segmentation, including our open-set tool nnInteractive, showing how spatial prompting can scale from user-guided interaction to clinical-grade automation. LLMs for Smarter Diagnosis: Unlocking the Future of AI in Healthcare ![](https://cdn.sanity.io/images/h6toihm1/production/b364057f3bb6cc61d113b7ea1ab6c47da07f853a-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Gaurav K Gupta Lake County Health Department Bio Large Language Models are rapidly transforming the healthcare landscape. In this talk, I will explore how LLMs like GPT-4 and DeepSeek-R1 are being used to support disease diagnosis, predict chronic conditions, and assist medical professionals without relying on sensitive patient data. Drawing from my published research and real-world applications, I’ll discuss the technical challenges, ethical considerations, and the future potential of integrating LLMs in clinical settings. The talk will offer valuable insights for developers, researchers, and healthcare innovators interested in applying AI responsibly and effectively. Join us for several virtual events focused on the latest research, datasets and models at the intersection of visual AI and healthcare. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-33-lllmstxt|> ## Data-Centric Computer Vision Tools [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Whitepaper # The Best Data-Centric Computer Vision Tools for the Enterprise **Choosing the right CV stack is a strategic decision, not a tooling preference.** As computer vision scales from pilots to production, the difference between success and stalled projects often comes down to how well your team manages data—not just models. This whitepaper lays out the critical components of a modern, data-centric CV toolchain: dataset management, versioning, auto-labeling, security, and model feedback loops. **Inside, we break down the top tooling options — from SaaS platforms to open-source and hybrid stacks.** Learn how to reduce labeling spend by 50%, avoid model drift in production, and meet privacy compliance requirements without slowing down development. Whether you're benchmarking tools like CVAT, FiftyOne, Encord, or Roboflow, this guide helps you evaluate tradeoffs and make decisions that scale with your business. ![](https://cdn.sanity.io/images/h6toihm1/production/50a686dd34d9caecdff5e2c62af0814b83710d44-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=1600) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-34-lllmstxt|> ## Customer Success Stories [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/1a672e54a7fb869beb104c2f70bbdee3ec5e5ff7-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=513&q=75&w=912) Featured How SafelyYou detected 350,000 elderly safety events with visual AI using FiftyOne [Read article](https://voxel51.com/customers/safelyyou) ![](https://cdn.sanity.io/images/h6toihm1/production/72a7793467b8ea2576dceb4ce7dce9331cdc561c-2128x2129.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=540&q=75&w=960) Featured 3x Faster robotics dataset investigations: How Berkshire Grey drives productivity with FiftyOne [Read article](https://voxel51.com/customers/berkshire-grey) ![](https://cdn.sanity.io/images/h6toihm1/production/d321cc301ad25d9c4e6890dd2d1eee225ec72bb9-2432x2432.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=540&q=75&w=960) Featured Ancera speeds up pathogen detection model development, regains months of engineering time [Read article](https://voxel51.com/customers/ancera) ![](https://cdn.sanity.io/images/h6toihm1/production/b225076935c6ada1d3969623e54240936c13b0dc-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=513&q=75&w=912) Featured RIOS’s AI-Powered Robotics Solutions Run on FiftyOne Enterprise [Read article](https://voxel51.com/customers/rios) ![](https://cdn.sanity.io/images/h6toihm1/production/112f66cd1604b5cee5d9cb2e4039a25260112f0b-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=513&q=75&w=912) Featured FiftyOne plays a vital role in processing claims at Allstate India [Read article](https://voxel51.com/customers/allstate) All AV & Physical AIRetail & ConsumerDev ToolsSecurity & SafetyAerospace & DefenseHealth & MedicineAgriculture & Sustainability [![](https://cdn.sanity.io/images/h6toihm1/production/2b8949efe1d52b7b699630b0aff09fe3316b5b01-608x608.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How Protex AI built a flexible, scalable ML pipeline for workplace safety with FiftyOne\\ \\ Security & Safety\\ \\ • \\ \\ Jul 15, 2025](https://voxel51.com/customers/protex-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/72a7793467b8ea2576dceb4ce7dce9331cdc561c-2128x2129.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ 3x Faster robotics dataset investigations: How Berkshire Grey drives productivity with FiftyOne\\ \\ AV & Physical AI\\ \\ • \\ \\ Jun 6, 2025](https://voxel51.com/customers/berkshire-grey) [![](https://cdn.sanity.io/images/h6toihm1/production/1a672e54a7fb869beb104c2f70bbdee3ec5e5ff7-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How SafelyYou detected 350,000 elderly safety events with visual AI using FiftyOne\\ \\ Health & Medicine\\ \\ • \\ \\ May 4, 2025](https://voxel51.com/customers/safelyyou) [![](https://cdn.sanity.io/images/h6toihm1/production/d321cc301ad25d9c4e6890dd2d1eee225ec72bb9-2432x2432.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Ancera speeds up pathogen detection model development, regains months of engineering time\\ \\ Dev Tools\\ \\ • \\ \\ May 3, 2025](https://voxel51.com/customers/ancera) [![](https://cdn.sanity.io/images/h6toihm1/production/b225076935c6ada1d3969623e54240936c13b0dc-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ RIOS’s AI-Powered Robotics Solutions Run on FiftyOne Enterprise\\ \\ AV & Physical AI\\ \\ • \\ \\ May 2, 2025](https://voxel51.com/customers/rios) [![](https://cdn.sanity.io/images/h6toihm1/production/af1426b965796134b254c4c1ff02be13740eebc5-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Forsight Finds a Centralized Dataset Management Solution in FiftyOne Teams\\ \\ Security & Safety\\ \\ • \\ \\ May 1, 2025](https://voxel51.com/customers/forsight) [![](https://cdn.sanity.io/images/h6toihm1/production/d286fcf14376e5da670d2fee8f7d5ab8b89fe09e-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Raytheon Technologies Research Center relies on FiftyOne to visualize large computer vision datasets\\ \\ Aerospace & Defense\\ \\ • \\ \\ Apr 27, 2025](https://voxel51.com/customers/raytheon-technologies) [![](https://cdn.sanity.io/images/h6toihm1/production/675fd154dadbf64742785ce336ce4ba3be8722a9-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne helps ADT leverage computer vision for security systems\\ \\ Security & Safety\\ \\ • \\ \\ Apr 26, 2025](https://voxel51.com/customers/adt-commercial) [![](https://cdn.sanity.io/images/h6toihm1/production/112f66cd1604b5cee5d9cb2e4039a25260112f0b-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne plays a vital role in processing claims at Allstate India\\ \\ Retail & Consumer\\ \\ • \\ \\ Apr 26, 2025](https://voxel51.com/customers/allstate) [![](https://cdn.sanity.io/images/h6toihm1/production/adeecdc032b72efa4a348cbcd745aab2973c1e56-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ IBM streamlines the supply chain with FiftyOne\\ \\ AV & Physical AI\\ \\ • \\ \\ Apr 25, 2025](https://voxel51.com/customers/ibm) [![](https://cdn.sanity.io/images/h6toihm1/production/a988d7ae16303abd2bfd52fd474f020a02aa235e-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Vivint relies on FiftyOne Teams for intelligent data insights and smarter home security\\ \\ Security & Safety\\ \\ • \\ \\ Apr 24, 2025](https://voxel51.com/customers/vivint) [![](https://cdn.sanity.io/images/h6toihm1/production/9090ca83d439b677604b11896dbe83589b6a9e98-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Secury360 uses FiftyOne to balance computer vision datasets & improve model performance\\ \\ Security & Safety\\ \\ • \\ \\ Apr 23, 2025](https://voxel51.com/customers/secury360) [![](https://cdn.sanity.io/images/h6toihm1/production/f5bfb518a5b173e357f7c318576cda65d29e7f59-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ RIF Robotics uses FiftyOne to enable robot-assisted surgery\\ \\ AV & Physical AI, Health & Medicine\\ \\ • \\ \\ Apr 22, 2025](https://voxel51.com/customers/rif-robotics) [![](https://cdn.sanity.io/images/h6toihm1/production/2889f465e725d668ea8151e261ccc307a60dcb62-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne gives Updata a powerful and flexible framework for ML\\ \\ Dev Tools\\ \\ • \\ \\ Apr 21, 2025](https://voxel51.com/customers/updata) [![](https://cdn.sanity.io/images/h6toihm1/production/5b420def37b04af0cc5d0167fe91492f7c7362ac-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Smart Eye relies on FiftyOne to build Human Insight AI\\ \\ Aerospace & Defense, AV & Physical AI\\ \\ • \\ \\ Apr 20, 2025](https://voxel51.com/customers/smart-eye) [![](https://cdn.sanity.io/images/h6toihm1/production/899b1fbd73a4f3061789a79484124b63b384d688-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne helps Taranis provide farmers with leaf-level insights for healthier crops and fields\\ \\ Agriculture & Sustainability\\ \\ • \\ \\ Apr 19, 2025](https://voxel51.com/customers/taranis) [![](https://cdn.sanity.io/images/h6toihm1/production/1f113782bbe98bcb928919a47a9f2bb5636402e2-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Fyma relies on FiftyOne to visualize CCTV datasets for object detection tasks\\ \\ Security & Safety, Retail & Consumer\\ \\ • \\ \\ Apr 18, 2025](https://voxel51.com/customers/fyma) [![](https://cdn.sanity.io/images/h6toihm1/production/d2f0a45f9678ebc503ec7b90fbeea293efb719ce-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne enables ArgosAI to increase the accuracy of their ML models for aviation use cases\\ \\ Aerospace & Defense\\ \\ • \\ \\ Apr 18, 2025](https://voxel51.com/customers/argosai) [![](https://cdn.sanity.io/images/h6toihm1/production/4c3b8e9d7736050517b8c81baf652a0566d6c7a2-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Kitro uses FiftyOne to train models to reduce food waste\\ \\ Agriculture & Sustainability\\ \\ • \\ \\ Apr 17, 2025](https://voxel51.com/customers/kitro) [![](https://cdn.sanity.io/images/h6toihm1/production/1ed4c2b8a455ca7707714826c8a8006865e5d2fd-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Aidence gives lung cancer patients a fighting chance using AI\\ \\ Health & Medicine\\ \\ • \\ \\ Apr 15, 2025](https://voxel51.com/customers/aidence) [![](https://cdn.sanity.io/images/h6toihm1/production/1f17885f5ff39ab29f3754810547e6bd2e038f40-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Aquabyte optimizes fish farming operations with AI and FiftyOne\\ \\ Agriculture & Sustainability\\ \\ • \\ \\ Apr 14, 2025](https://voxel51.com/customers/aquabyte) [![](https://cdn.sanity.io/images/h6toihm1/production/6c1873bdb23550a0e00810b5f5246853c39b13db-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Aisprid’s high-precision robots for greenhouse farming rely on FiftyOne\\ \\ Agriculture & Sustainability, AV & Physical AI\\ \\ • \\ \\ Apr 13, 2025](https://voxel51.com/customers/aisprid) [![](https://cdn.sanity.io/images/h6toihm1/production/76bbb00710f5018bbd3eca2d5e71f4675a4ce26d-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne strengthens Seafar’s solutions for autonomous vessels for shipping\\ \\ AV & Physical AI\\ \\ • \\ \\ Apr 12, 2025](https://voxel51.com/customers/seafar) [![](https://cdn.sanity.io/images/h6toihm1/production/d2cffb30499c8bfdf723500dbde020b5b834199d-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Ai.Fish revolutionizes fishery management with FiftyOne\\ \\ Agriculture & Sustainability\\ \\ • \\ \\ Apr 11, 2025](https://voxel51.com/customers/ai-fish) [![](https://cdn.sanity.io/images/h6toihm1/production/b224afdc229760342dd440fdcc31e50b5550d452-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne fuels Finegrain’s ML workflows and data curation\\ \\ Retail & Consumer\\ \\ • \\ \\ Apr 10, 2025](https://voxel51.com/customers/finegrain) [![](https://cdn.sanity.io/images/h6toihm1/production/72ec0f31c4005d0d7c86bbecd6519522fec2d88b-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ G42 gets a superior interface for computer vision data with FiftyOne\\ \\ Dev Tools\\ \\ • \\ \\ Apr 9, 2025](https://voxel51.com/customers/g42) [![](https://cdn.sanity.io/images/h6toihm1/production/1c2c597f18da32cae6956b1bbfc6502efb663b97-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne helps BinIt provide visibility, insights, and automation for efficient material recovery\\ \\ Agriculture & Sustainability\\ \\ • \\ \\ Apr 7, 2025](https://voxel51.com/customers/binit) [![](https://cdn.sanity.io/images/h6toihm1/production/3191b9ce305cb63311a09968bd394af91f5e837a-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Qinecsa’s pharmacovigilance solutions rely on FiftyOne\\ \\ Health & Medicine\\ \\ • \\ \\ Apr 6, 2025](https://voxel51.com/customers/qinecsa) [![](https://cdn.sanity.io/images/h6toihm1/production/cd76dddd6d059fdc96846aa6a6cab58ff156b0c1-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ METU uses FiftyOne to conduct medical research on inflammatory bowel disease\\ \\ Security & Safety\\ \\ • \\ \\ Apr 5, 2025](https://voxel51.com/customers/metu) [![](https://cdn.sanity.io/images/h6toihm1/production/5c5a66794fd02f9b3f3594f57be644e9f7f68ea7-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Cybersecurity research at the Indian Institute of IT & Management runs on FiftyOne\\ \\ Security & Safety\\ \\ • \\ \\ Apr 4, 2025](https://voxel51.com/customers/indian-institute-of-it-and-management) [![](https://cdn.sanity.io/images/h6toihm1/production/60895f85f1668e61f0ed0032a3075ac73d5d0918-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Fast Code AI relies on FiftyOne for managing its data lifecycle\\ \\ Dev Tools\\ \\ • \\ \\ Apr 3, 2025](https://voxel51.com/customers/fast-code-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/31cd5d87baf1e384ed51106820bab1b6de42f927-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ LanceDB finds FiftyOne an invaluable developer tool for building its vector search database\\ \\ Dev Tools\\ \\ • \\ \\ Apr 2, 2025](https://voxel51.com/customers/lancedb) [![](https://cdn.sanity.io/images/h6toihm1/production/ad5c1f638ed4aa77e2de95e816fded79e4b3b0dd-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Wildlife.ai uses FiftyOne to accelerate wildlife conservation\\ \\ Agriculture & Sustainability\\ \\ • \\ \\ Apr 1, 2025](https://voxel51.com/customers/wildlife-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/cd7fe79ef465aa75489f4c049ce2af346f0985b4-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/720f6702611411baf6a274b1010856c78c07235d-240x96.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/57564abbae97325ddeb3bc3b8b7b5690d61cddc5-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/be26ca6233b020b9201ddd482f7d34f53555b51d-276x90.png?auto=format&dpr=2&fit=max&q=75&w=100) Maintained 99% fall detection rates for model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/636eaa49ef62ea3d07ca264f7baa4fe0e494e125-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/3911d157da27a0e996468bf45393d4e01e522934-384x39.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/a54f5857074a8e45684334b38c3e4f1f1528274e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/e845660699edf4deed110e60425659ba8ae73e6e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Foundation for Florence-2 VLM development](https://voxel51.com/plugins/?search=florence) ![](https://cdn.sanity.io/images/h6toihm1/production/c8be8bb13c2be64cd70dda28334501a812eb0bc1-520x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/8c596efcf17de5a9bc02b05b44f55474802f0abe-90x90.png?auto=format&dpr=2&fit=max&q=75&w=90) ![](https://cdn.sanity.io/images/h6toihm1/production/1f4c7beab6c76544152d4f3d990059bc62eb90cb-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/aee0fde76666dd2d134875eb8d5247ee75dab896-153x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cb778b338fac262a1c8a9de0334b256747cfbc5b-500x261.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Eliminated repetitive manual transformations on 20 TB+ of visual data](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/2eed779bf08ffa8ea43cb822cb013674cc05e546-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/6f935624676c89371a1b094385b7b60439e09824-351x60.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/09890f00ed8e1cb9be4a30a0a0a52a80b2d74fee-800x1220.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cbbc7f79176fa84b6e2d1f15eb76b27b84f216b9-393x128.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg)\\ \\ Official partner for visualizing Open Images Dataset V7](https://voxel51.com/blog/exploring-google-open-images-v7/) ![](https://cdn.sanity.io/images/h6toihm1/production/cd7fe79ef465aa75489f4c049ce2af346f0985b4-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/720f6702611411baf6a274b1010856c78c07235d-240x96.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/57564abbae97325ddeb3bc3b8b7b5690d61cddc5-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/be26ca6233b020b9201ddd482f7d34f53555b51d-276x90.png?auto=format&dpr=2&fit=max&q=75&w=100) Maintained 99% fall detection rates for model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/636eaa49ef62ea3d07ca264f7baa4fe0e494e125-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/3911d157da27a0e996468bf45393d4e01e522934-384x39.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/a54f5857074a8e45684334b38c3e4f1f1528274e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/e845660699edf4deed110e60425659ba8ae73e6e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Foundation for Florence-2 VLM development](https://voxel51.com/plugins/?search=florence) ![](https://cdn.sanity.io/images/h6toihm1/production/c8be8bb13c2be64cd70dda28334501a812eb0bc1-520x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/8c596efcf17de5a9bc02b05b44f55474802f0abe-90x90.png?auto=format&dpr=2&fit=max&q=75&w=90) ![](https://cdn.sanity.io/images/h6toihm1/production/1f4c7beab6c76544152d4f3d990059bc62eb90cb-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/aee0fde76666dd2d134875eb8d5247ee75dab896-153x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cb778b338fac262a1c8a9de0334b256747cfbc5b-500x261.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Eliminated repetitive manual transformations on 20 TB+ of visual data](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/2eed779bf08ffa8ea43cb822cb013674cc05e546-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/6f935624676c89371a1b094385b7b60439e09824-351x60.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/09890f00ed8e1cb9be4a30a0a0a52a80b2d74fee-800x1220.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cbbc7f79176fa84b6e2d1f15eb76b27b84f216b9-393x128.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg)\\ \\ Official partner for visualizing Open Images Dataset V7](https://voxel51.com/blog/exploring-google-open-images-v7/) ![](https://cdn.sanity.io/images/h6toihm1/production/cd7fe79ef465aa75489f4c049ce2af346f0985b4-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/720f6702611411baf6a274b1010856c78c07235d-240x96.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/57564abbae97325ddeb3bc3b8b7b5690d61cddc5-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/be26ca6233b020b9201ddd482f7d34f53555b51d-276x90.png?auto=format&dpr=2&fit=max&q=75&w=100) Maintained 99% fall detection rates for model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/636eaa49ef62ea3d07ca264f7baa4fe0e494e125-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/3911d157da27a0e996468bf45393d4e01e522934-384x39.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/a54f5857074a8e45684334b38c3e4f1f1528274e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/e845660699edf4deed110e60425659ba8ae73e6e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Foundation for Florence-2 VLM development](https://voxel51.com/plugins/?search=florence) ![](https://cdn.sanity.io/images/h6toihm1/production/c8be8bb13c2be64cd70dda28334501a812eb0bc1-520x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/8c596efcf17de5a9bc02b05b44f55474802f0abe-90x90.png?auto=format&dpr=2&fit=max&q=75&w=90) ![](https://cdn.sanity.io/images/h6toihm1/production/1f4c7beab6c76544152d4f3d990059bc62eb90cb-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/aee0fde76666dd2d134875eb8d5247ee75dab896-153x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cb778b338fac262a1c8a9de0334b256747cfbc5b-500x261.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Eliminated repetitive manual transformations on 20 TB+ of visual data](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/2eed779bf08ffa8ea43cb822cb013674cc05e546-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/6f935624676c89371a1b094385b7b60439e09824-351x60.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/09890f00ed8e1cb9be4a30a0a0a52a80b2d74fee-800x1220.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cbbc7f79176fa84b6e2d1f15eb76b27b84f216b9-393x128.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg)\\ \\ Official partner for visualizing Open Images Dataset V7](https://voxel51.com/blog/exploring-google-open-images-v7/) ## Talk to a computer vision expert Get in touch to learn about how the best machine learning teams are building visual AI. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-35-lllmstxt|> ## Visual AI in Manufacturing [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/e3b40f53b685e10d34816d355272c984eb12a0d1-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/e3b40f53b685e10d34816d355272c984eb12a0d1-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event Virtual Americas Meetups Manufacturing Visual AI in Manufacturing and Robotics - September 10, 2025 Sep 10, 2025 9 AM Pacific Online. Fill in the form to register! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/0e99a0f5f3e9b669b9985bd84a553cc065a96073-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Matt Puchalsk Bucket Robotics Bio ![](https://cdn.sanity.io/images/h6toihm1/production/d8983f9f8fb1fa8819ab9b8966fb8be988a103bf-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Paula Ramos voxel51 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/0c968c6a456d9cb99369dc64a16ea8754ec540b3-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Frederick Gertz Collide Technology Bio About this event Join us for the first in a series of virtual events to hear talks from experts on the latest developments at the intersection of Visual AI, Manufacturing and Robotics. Schedule Scaling Synthetic Data for Industrial AI: From CAD to Model in Hours ![](https://cdn.sanity.io/images/h6toihm1/production/0e99a0f5f3e9b669b9985bd84a553cc065a96073-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Matt Puchalsk Bucket Robotics Bio This talk explores how we generate high-performance computer vision datasets from CAD—without real-world images or manual labeling. We’ll walk through our synthetic data pipeline, including CPU-optimized defect simulation, material variation, and lighting workflows that scale to thousands of renders per part. While Blender plays a role, our focus is on how industrial data (like STEP files) and procedural generation unlock fast, flexible training sets for manufacturing QA, even on modest hardware. If you're working at the edge of 3D, automation, and vision AI—this is for you Detecting the Unexpected: Practical Approaches to Anomaly Detection in Visual Data ![](https://cdn.sanity.io/images/h6toihm1/production/d8983f9f8fb1fa8819ab9b8966fb8be988a103bf-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Paula Ramos voxel51 Bio Anomaly detection is one of computer vision's most exciting and essential challenges today. From spotting subtle defects in manufacturing to identifying edge cases in model behavior, it is one of computer vision's most exciting and crucial challenges. In this session, we’ll do a hands-on walkthrough using the MVTec AD dataset, showcasing real-world workflows for data curation, exploration, and model evaluation. We’ll also explore the power of embedding visualizations and similarity searches to uncover hidden patterns and surface anomalies that often go unnoticed. This session is packed with actionable strategies to help you make sense of your data and build more robust, reliable models. Join us as we connect the dots between data, models, and real-world deployment—alongside other experts driving innovation in anomaly detection. Swarm Intelligence: Solving Complex Industrial Optimization in Seconds ![](https://cdn.sanity.io/images/h6toihm1/production/0c968c6a456d9cb99369dc64a16ea8754ec540b3-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Frederick Gertz Collide Technology Bio Manufacturing and logistics companies face increasingly complex operational challenges that traditional AI and human planning struggle to solve effectively. Collide Technology harnesses Swarm Intelligence algorithms to transform intractable problems—like scheduling hundreds or thousands of maintenance employees while simultaneously optimizing production capacity, inventory levels, and cross-sector resource allocation—into solutions delivered in seconds rather than weeks. Unlike rigid Operations Research approaches that require specialized expertise and expensive implementations, our platform democratizes industrial optimization by making sophisticated decision-making accessible to any factory or logistics operation. We deliver holistic, data-driven solutions that optimize across multiple business entities and sectors simultaneously, adapting to real-world constraints and evolving operational needs. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-36-lllmstxt|> ## Importance of Dataset Annotation [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) Why Quality Dataset Annotation Is Key to Machine Learning Feb 17, 2025 • 12 min read Article content In this article [Dataset annotation: What is it and why is it important?](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#f25ab8798784) [Building high-quality annotations: best practices](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#7d8e1c3d1ec6) [FiftyOne's role in streamlining annotation](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#bb95b1408c89) [Real-world examples](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#18b75296611e) [FiftyOne: Your partner in streamlining annotation](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#608f81d78672) [Beyond annotation: The future of data labeling](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#4f00289467f5) [FiftyOne's integration with emerging trends](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#24f26e9603f0) In this article [Dataset annotation: What is it and why is it important?](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#f25ab8798784) [Building high-quality annotations: best practices](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#7d8e1c3d1ec6) [FiftyOne's role in streamlining annotation](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#bb95b1408c89) [Real-world examples](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#18b75296611e) [FiftyOne: Your partner in streamlining annotation](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#608f81d78672) [Beyond annotation: The future of data labeling](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#4f00289467f5) [FiftyOne's integration with emerging trends](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning#24f26e9603f0) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Saying that machine learning (ML) is transforming industries worldwide doesn’t quite capture its true impact. ML has brought a 360-degree shift to how businesses operate and solve problems, and experts predict the global ML market will reach USD [500 billion](https://www.statista.com/outlook/tmo/artificial-intelligence/machine-learning/worldwide#:~:text=The%20market%20size%20in%20the,US%24503.40bn%20by%202030.) by 2030. Machine learning models assist in early disease detection and personalized healthcare treatments. Financial institutions also use ML for fraud detection and algorithmic trading. Autonomous vehicles depend on these models for safe navigation. Even customer service is witnessing a revolution with ML powering chatbots and recommendation systems. However, the success of these applications relies heavily on data quality. ML models can only perform as well as the data used for training. High-quality, well-prepared datasets result in more reliable models with high generalizability. But what determines data quality? Annotation (also known as labeling) is one of the most critical factors. Annotations describe the content of each data sample, helping ML models understand what the sample contains. Here, we will discuss what annotation is and why accurate dataset annotation is important for ML development. We will also discuss the risks linked with poor annotation practices, key strategies for high-quality labeling, and how [FiftyOne](https://voxel51.com/) can improve how you annotate data, particularly for computer vision (CV) systems. ## Dataset annotation: What is it and why is it important? **Dataset annotation** adds key information to data samples that enable ML models to learn patterns and make predictions. For example, training a CV model may include adding accurate and relevant labels to images. The model can then learn patterns that help identify the content within each image. ### Types of annotation tasks Multiple annotation strategies exist for different data types. The list below highlights the methods used for labeling visual, text, and audio data for training and evaluating ML models. **Visual annotation** labels visual media samples to train and evaluate models. Types of visual data include still images, video, and 3D scenes, as well as standardized formats like [DICOM](https://www.dicomstandard.org/) for medical imaging. Common computer vision tasks requiring annotation include classification, object detection, and segmentation. ![](https://cdn.sanity.io/images/h6toihm1/production/822a84920d426faca0c7cf93d5bc00231420ddf6-2560x1073.png?auto=format&dpr=2&fit=max&q=75&w=1600) The goal in **classification** is to apply one or more labels to an entire image sample, whereas, in object detection, annotators mark objects with bounding boxes, which are useful to train the model to find individual objects and their locations within the image. **Keypoint annotation** identifies key landmarks, within an object, such as individual joints on a body, for predicting movements. And **segmentation** is even more complex, requiring pixel-level annotation for precise object boundaries. **Semantic segmentation**, for example, assigns a specific label to each pixel based on the object class it belongs to (e.g., “car”, “dog”, or “stopsign”). All these image annotation types are essential for applications like [autonomous driving](https://voxel51.com/computer-vision-use-cases/driving/) and [facial recognition](https://www.knowledgenile.com/blogs/computer-vision-and-facial-recognition-real-world-applications). **Sentiment analysis** identifies emotions or opinions in a text. For instance, it can label customer feedback as “positive,” “negative,” or “neutral”. PoS tagging identifies and labels each word in a text with its corresponding grammatical role or part of speech. In contrast, NER extracts and labels entities like names, dates, or locations. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Text annotations are essential for building [natural language processing (NLP) applications](https://datasciencedojo.com/blog/natural-language-processing-applications/) like chatbots and virtual assistants. Finally, **audio annotation** labels audio files to train models that process and analyze sounds and voices. Key tasks include [speech-to-text transcription](https://huggingface.co/docs/transformers/en/model_doc/speech_to_text), [text-to-speech generation](https://huggingface.co/tasks/text-to-speech), and [speaker identification](https://paperswithcode.com/task/speaker-identification). Labeling audio requires converting audio files into waveforms and spectrograms. Developers can use these formats to label specific regions to indicate frequency, amplitude, and length. ![](https://cdn.sanity.io/images/h6toihm1/production/db3a63098f917d1e84b05e3cd83a033aa763e030-2334x1774.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### Human expertise and the impact of poor annotation Human expertise helps ensure annotations are accurate, relevant, and context-aware. Skilled annotators understand nuances, cultural differences, and domain-specific knowledge that automated tools might miss. Their input helps maintain data quality and align annotations with real-world contexts. However, while human expertise is invaluable, even the most skilled annotators make mistakes. Misunderstandings or inconsistent interpretations of the data can affect annotations. Numerous [examples](https://pmc.ncbi.nlm.nih.gov/articles/PMC9944930/?utm_source=chatgpt.com) in ML datasets illustrate how poor annotation can derail model training. - **Bias and Inaccuracy:** Poor-quality data annotations can introduce biases into datasets. Such biases skew predictions and potentially lead to unintended or [harmful consequences](https://arxiv.org/pdf/2304.07683). - **Decreased Model Performance:** Inaccurate or inconsistent annotations create unreliable training data, impacting a model’s accuracy and generalizability. Models trained on poorly annotated data struggle to perform well on real-world tasks. - **Interpretability Challenges:** Poor annotations can reduce an ML model’s [interpretability](https://www.turing.com/kb/interpretability-methods-in-machine-learning), meaning developers cannot clearly understand a model’s decision-making process when predicting specific outputs. For example, imprecise bounding boxes in radiological data can lead to a model confusing tumors for healthy tissue or otherwise finding “phantom” tumors in healthy regions. The model could then output results that defy medical expectations, such as concluding malignancy from irrelevant organ regions, leading to radiologists viewing the model as untrustworthy and unfit for clinical use. In January 2021, [Koksal et al.](https://arxiv.org/pdf/2004.01059) published a paper on how annotation inconsistencies impact drone detection using the well-known detection model YOLOv3. They used the [CVPR-2020 Anti-UAV Challenge](https://ieeexplore.ieee.org/abstract/document/9615243) dataset to train YOLOv3 and assess its detection and tracking performance. The dataset contains 300 video frames and 580,000 manually annotated bounding boxes. The researchers introduced multiple annotation errors, including additional bounding boxes, missing boxes, and shifted boxes. They assessed the impact of these errors on the model’s tracking accuracy. The results show that with original annotations, tracking accuracy was 73.6%. However, the combined effect of the annotation errors reduced accuracy to 54.2%. The performance degradation clearly highlights the importance of annotation quality, even in robust models like YOLOv3. ## Building high-quality annotations: best practices Although specific annotation procedures may differ by use case, the following sections outline best practices to help you ensure high-quality, reliable annotations for any application. ### Select an appropriate annotation schema Choosing the correct annotation schema is vital. Start by clearly defining the primary task to balance speed and accuracy. For instance, classification tasks, which assign labels to an entire image, are quicker to complete but offer less detail. In contrast, bounding boxes or pixel-level annotations of objects within images require more time and precision but provide richer data for complex models. Documenting attributes and labels is equally critical and ensures consistent labeling across the dataset. Also, consider adopting an existing data schema, like the format used by the [COCO](https://cocodataset.org/) dataset, which offers predefined categories and annotation structures. Using established schemas saves time, ensures compatibility with industry standards, and supports benchmarking against existing ML models. The schema should also address ambiguities and edge cases in the data. For example, annotators need instructions on handling occluded or unclear objects, such as applying fallback labels like “unknown” or “partial.” ### Minimize bias with diverse annotators A diverse pool of annotators can help reduce bias in the annotation process. Different cultural, linguistic, and demographic backgrounds can offer varied perspectives. This reduces the risk of annotations reflecting a narrow or skewed worldview. For example, in image annotation, multiple annotators can reduce biases related to ethnicity, gender, or social context, resulting in more inclusive and fair datasets. ### Implementing quality control **Double-annotation** and **consensus-based** approaches can also boost annotation quality. Double annotation requires multiple annotators to label the same data. The method can quickly reveal discrepancies and inaccuracies that a single annotator may miss. A consensus-based approach takes this further by requiring annotators to agree on the final label. The process may include discussions or automated algorithms that resolve disagreements. ## FiftyOne's role in streamlining annotation Given the importance of data quality, we need good software to manage the datasets being annotated. [FiftyOne](https://voxel51.com/fiftyone/) is a tool that enables visual AI builders to refine data and models in one place. It includes a variety of features to both visualize datasets and improve data quality. For example, FiftyOne makes it easy to find and remediate [mistakes](https://docs.voxel51.com/tutorials/detection_mistakes.html) in your data, such as annotation errors like incorrect or missing object labels. ## Real-world examples The previous sections emphasized the importance of best practices for improving annotation pipelines. Now, let’s go over some real-world scenarios where accurate and consistent annotations directly impact model performance. - **Medical Imaging:** Accurate annotation of [medical images](https://voxel51.com/computer-vision-use-cases/healthcare/) can help build reliable AI models to detect and diagnose conditions like tumors, fractures, or abnormalities. Labeling key features like lesions or organs, annotators allow the model to spot patterns and make precise predictions. The approach boosts diagnostic accuracy and supports faster, more reliable healthcare decisions. - **Self-Driving Cars:** Properly annotated traffic data is critical for [training self-driving car models](https://voxel51.com/computer-vision-use-cases/driving/) to navigate complex environments. Accurate annotations of road signs, pedestrians, vehicles, and obstacles help the model understand its surroundings and make informed decisions. These annotations improve the vehicle’s ability to drive safely, reduce accidents, and adapt to edge-case situations. - **Sentiment Analysis:** Well-annotated text data for [sentiment analysis](https://aclanthology.org/W16-0429.pdf) can help businesses gauge customer emotions accurately. Annotators can gain many insights by labeling customer reviews as positive, negative, or neutral. This can help management get a quick snapshot of how customers view the company’s products or services and reveal improvement areas. ## FiftyOne: Your partner in streamlining annotation We mentioned earlier that FiftyOne is a toolkit for managing and refining the datasets used in your visual AI projects. FiftyOne helps teams curate the samples needing annotation and perform quality control to ensure annotations are correct and relevant to your models, while also integrating with your chosen annotation software. ### Explore and curate data FiftyOne makes it easy to navigate through large datasets both programmatically and in the App in order to curate the data needing annotation. At a minimum, this could mean examining and filtering sample metadata for vital statistics like image size and aspect ratio. It could also mean verifying quality metrics like object clarity and potential duplicate samples before passing the data to annotators. ![](https://cdn.sanity.io/images/h6toihm1/production/9deead09c81f51ccb0c02eafa92ad0c9f67fd7b9-2560x1424.png?auto=format&dpr=2&fit=max&q=75&w=1600) FiftyOne includes a rich plugin ecosystem to assist annotation teams. For example, a team can use the [zero-shot model prediction](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone/) plugin to perform initial predictions on the dataset that are later validated by humans. Then, once annotation teams are ready to begin their work, FiftyOne can hook into any defined annotation backend. For instance, teams can generate [dataset views](https://docs.voxel51.com/user_guide/using_views.html#using-views) containing samples likely needing new or edited annotations, upload those samples to a platform like [CVAT](https://www.cvat.ai/), perform the annotations, and then reload the samples back into FiftyOne. ### Visualize and improve datasets FiftyOne works with visual media like images, videos, 3D scenes, and even geolocation data, and likewise includes many ways to identify possible labeling mistakes. One method is to compute [embeddings](https://docs.voxel51.com/tutorials/clustering.html) from the data. Embeddings are vector representations of image properties. They can be used to identify distributions of unique, similar, or ambiguous samples within a dataset, and also visually identify outliers that could indicate mistaken labels. ![](https://cdn.sanity.io/images/h6toihm1/production/94d0ad7d37d467ebda214d1956c5a0772640b417-2560x1428.png?auto=format&dpr=2&fit=max&q=75&w=1600) You could also use FiftyOne’s interactive plots to create dynamic charts and dashboards. These are viewable in the aAp and can be useful for catching annotation mistakes. In the figure below, the pie chart shows the ratio of model-predicted false positives to total positives in the sample set. Selecting the false positive region of the pie chart creates an equivalent dataset view of false positive samples. From there you can further drill into the samples to determine if some of these false positives are the result of annotation errors. ![](https://cdn.sanity.io/images/h6toihm1/production/7a5e9082763144a73308c35b2f02315d71b4d2ff-2560x1356.png?auto=format&dpr=2&fit=max&q=75&w=1600)![](https://cdn.sanity.io/images/h6toihm1/production/bc378df5f42670fb0a8b4070e3dd6e59ac352450-2560x1461.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### Team collaboration The [enterprise version](https://voxel51.com/enterprise/) of FiftyOne is built for collaboration, with the goal of making it as easy as possible for multiple users to work together to build high-quality datasets and computer vision models. A data scientist, for instance, might use FiftyOne to tag samples as likely label mistakes and then share that dataset with a data quality assurance team to verify, before sending the tagged images for reannotation with one of FiftyOne’s native annotation integrations. The data science and QA teams can then iterate between correcting sample annotations and re-training the model with the improved data. ![](https://cdn.sanity.io/images/h6toihm1/production/ffd1125a3f196eb6ba62d15debf6af647b8f7462-2560x1124.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Beyond annotation: The future of data labeling As use cases become more complex, the availability of labeled data becomes more limited. This phenomenon is making model development a major challenge. However, [ongoing research](https://www.researchgate.net/publication/258650578_A_Study_on_Image_Annotation_Techniques) into alternative approaches is leading to innovative annotation methods that require minimal labeled data for training models. The following section provides an overview of these emerging trends and how FiftyOne integrates with them to offer a comprehensive end-to-end labeling solution. ### Active learning [Active learning](https://arxiv.org/pdf/2211.14819) allows models to suggest which data points to annotate next. Instead of labeling random samples, the model predicts labels for a subset of unlabeled data. It then identifies predictions for which it has low confidence. ![](https://cdn.sanity.io/images/h6toihm1/production/2eaeb9124cf728b88d4f655bd793b05fcc9a2d60-2280x1340.png?auto=format&dpr=2&fit=max&q=75&w=1600) Lastly, it asks the user to provide the labels for such data points. The technique reduces the amount of labeled data needed and makes the process more resource-efficient and cost-effective. ### Semi-supervised learning [Semi-supervised learning (SSL)](https://link.springer.com/content/pdf/10.1007/s11704-019-8452-2.pdf) trains models on labeled and unlabeled data. In this approach, annotators provide a small amount of labeled data to help the model learn and improve. ![](https://cdn.sanity.io/images/h6toihm1/production/54804806541b07117d0f13ba5574d08750ca0d6d-1780x1222.png?auto=format&dpr=2&fit=max&q=75&w=1600) Additionally, large amounts of unlabeled data help refine the model's understanding. SSL is beneficial when labeled data is scarce or costly to obtain. It is a more efficient and scalable alternative to traditional, fully supervised learning. ## FiftyOne's integration with emerging trends FiftyOne integrates with active learning and SSL frameworks. This includes a [plugin](https://github.com/jacobmarks/active-learning-plugin?tab=readme-ov-file) for the [modAL](https://modal-python.readthedocs.io/en/latest/) library for building active learning pipelines. You can easily integrate it into your data annotation projects to enhance labeling accuracy with minimal additional coding effort. Additionally, the modAL plugin is compatible with industry-standard annotation tools like CVAT, Labelbox, and Label Studio, ensuring a smooth and efficient annotation process. FiftyOne also supports state-of-the-art models like [CoTracker3](https://voxel51.com/blog/cotracker3-a-point-tracker-using-real-videos/?utm_content=317088110&utm_medium=social&utm_source=twitter&hss_channel=tw-901403524868825088), an innovative point tracker that predicts a point’s trajectory in a video using its existing location. Unlike other tracking frameworks that use synthetic data, CoTrakcer3 uses SSL to predict locations based on real-world video feeds. With [FiftyOne’s model evaluation features](https://docs.voxel51.com/user_guide/evaluation.html#), developers can use CoTracker3’s SSL ability to improve a model’s generalization performance. ### Summary and next steps With rising concerns around AI’s reliability in mission-critical applications like healthcare and autonomous cars, high-quality annotations are key ingredients in determining a model’s success. Accurate and consistent annotations ensure models learn effectively and perform well in unknown situations. Organizations can invest in tools like FiftyOne to help developers build scalable labeling pipelines for multiple use cases. The platform empowers AI experts by streamlining and integrating the annotation process with advanced techniques like active and semi-supervised learning. If you want to take your AI workflows to the next level, you can explore the following resources to help you get started: - Join FiftyOne’s [Discord](https://discord.com/invite/fiftyone-community) community to learn directly from AI practitioners in the industry. - [Book a demo](https://voxel51.com/book-a-demo/) with a computer vision expert to learn how FiftyOne provides a scalable solution to your enterprise AI workflows - [Install the open source edition](https://docs.voxel51.com/getting_started/install.html) and explore the [documentation](https://docs.voxel51.com/) and [guided tutorials](https://docs.voxel51.com/tutorials/index.html) [dataset management](https://voxel51.com/blog/tag/dataset-management) [machine learning](https://voxel51.com/blog/tag/machine-learning) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/53c3b307e202d573a6fccf19494ba35c43a3dc7e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Best Practices for Evaluating AI Models Accurately\\ \\ Learn\\ \\ • \\ \\ Dec 17, 2024](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately) [![](https://cdn.sanity.io/images/h6toihm1/production/d2e24d0a14de508f8ccd36eaffd3b909c9f193b9-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Image Embeddings Transform Computer Vision Capabilities\\ \\ Learn\\ \\ • \\ \\ Nov 25, 2024](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities) [![](https://cdn.sanity.io/images/h6toihm1/production/ebd118c3d8181dcba0532200d024759e23c7ec1e-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Computer vision in healthcare: 12 breakthrough case studies\\ \\ Industry Solutions\\ \\ • \\ \\ Jul 15, 2025](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-37-lllmstxt|> ## BinIt Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/1c2c597f18da32cae6956b1bbfc6502efb663b97-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) BinIt FiftyOne helps BinIt provide visibility, insights, and automation for efficient material recovery Apr 7, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [BinIt](https://binit.ai/) is the first fully integrated analytics and data platform for waste management, helping Materials Recovery Facilities analyze materials and boost revenues via data. > "FiftyOne helps us stay on top of our incredibly diverse vision dataset, allowing us to build better models faster!" – Raghav Mecheri, CEO at BinIt ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-38-lllmstxt|> ## Point Cloud Management [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # FiftyOnePoint Cloud Support Get a data engine for 3D visualization in autonomous driving, industrial, defense, and physical AI use cases. [Book a demo](https://voxel51.com/sales) [Learn more](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e) ![](https://cdn.sanity.io/images/h6toihm1/production/ec8449933ece119d748172cc66ea7daf8d0cb825-715x521.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=350&q=75&w=350) ![](https://cdn.sanity.io/images/h6toihm1/production/e8b3011557a742a185138295a1dbc5ac48f0e5e6-564x340.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=340) Point Cloud Pain Points ## Working with point cloud data is challenging Modern ML teams working with 3D and sensor data—such as point clouds from LiDAR—struggle with fragmented workflows, clunky tooling, and poor dataset understanding. ![](https://cdn.sanity.io/images/h6toihm1/production/e1fd1039e37398073314736814885ffaeff91151-1208x689.gif?auto=format&dpr=2&fit=max&q=75&w=600) Point Cloud Support ## FiftyOne provides a unified interface for analyzing point clouds ### Visualize point cloud data at scale Seamlessly inspect point clouds, LiDAR annotations, and multimodal sensor data in a powerful, interactive interface. ### Debug and explore model outputs Identify failure modes, compare predictions vs. ground truth, and drill into edge cases across complex 3D scenes. ### Compare datasets and experiments Spot differences across versions, identify covariate shift, and track performance on curated subsets—all in one place. Point Cloud Features ## How FiftyOne helps you manage point cloud data ### Effortlessly manage millions of point cloud samples FiftyOne makes it easy to organize large-scale point cloud data alongside images, video, and metadata—essential for autonomous vehicles and sensor-driven systems. Point cloud-first: Manage LiDAR, radar, and depth-based point cloud data at scale Multimodal support: Combine point clouds with images, videos, frames, geolocation, and more Custom metadata: Track time of day, sensor IDs, location, weather, and other context critical to your AI models Compatible with any model: Object detection, lane detection, semantic segmentation, and more ![](https://cdn.sanity.io/images/h6toihm1/production/efc2dd080cda4cfab1e59377f015f2b984d9c703-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Quickly find the subsets of data you want Sifting through massive amounts of data is like searching for a needle in a haystack. Pinpoint samples of interest in seconds using FiftyOne. Create meaningful, balanced datasets: query samples by metadata to correct for imbalances Accelerate training data selection: quickly find unique scenarios and anomalies in your data streams Cover the edge cases: identify hard samples to strengthen your datasets and model performance ![](https://cdn.sanity.io/images/h6toihm1/production/22885a78a7c957147b1e93fc91cd4bb1736a1b2e-2560x1806.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Improve model performance with confidence Models can struggle on new data. FiftyOne helps you visualize, compare, and debug performance—so you can deploy with confidence. - **Find failure modes**: Explore point cloud predictions at the sample level - **Continuously evaluate**: Track model and dataset quality across training cycles - **Version datasets**: Manage and roll back dataset changes with ease Find failure modes: Explore model predictions and debug at the sample level Embrace active learning: integrate FiftyOne into your training pipeline to improve model performance and datasets with every update Version datasets: Manage and roll back dataset changes with ease ![](https://cdn.sanity.io/images/h6toihm1/production/4fb68f28115173c89811e57ddde635273c8c8e99-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-39-lllmstxt|> ## Building GUI Agents Workshop [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/03ebd9c87f7cc01872aaf1d562b1d3cdb5d10754-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/03ebd9c87f7cc01872aaf1d562b1d3cdb5d10754-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event Virtual Americas Webinars & Workshops From Research to Reality: Building GUI Agents That Actually Work - August 22, 2025 Aug 22, 2025 9 AM Pacific Online. Register for the Zoom! About this event Welcome to the Visual Agents Workshop Series, your virtual pass to learn about visual agents - how they work, how to develop them and how to fine-tune them. Host ![](https://cdn.sanity.io/images/h6toihm1/production/c96544cfe7c8fc1e23601a34ca6a5fc11ccd6aa5-320x320.png?auto=format&dpr=2&fit=max&q=75&w=96) Harpreet Sahota Voxel51 Bio ### Part 2: From Pixels to Predictions - Building Your GUI Dataset Hands-On Dataset Creation and Curation with FiftyOne The best GUI models are only as good as their training data, and the best datasets are built by understanding what makes GUI interactions fundamentally different from natural images. In this practical session, you'll build a complete GUI dataset from scratch, learning to capture the precise annotations that GUI agents need. Using FiftyOne as your data management backbone, you'll import diverse GUI screenshots, explore annotation strategies that go beyond bounding boxes, and implement efficient labeling workflows. We'll tackle the real challenges: handling platform differences, managing annotation quality, and creating datasets that transfer to new domains. You'll also learn advanced techniques like synthetic data generation and automated prelabeling to scale your annotation efforts. Walk away with a production-ready dataset and the skills to build more—because in GUI agents, data quality determines everything. By the end, you'll have both a dataset and the methodology to build the next generation of GUI training data. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-40-lllmstxt|> ## FiftyOne Visual AI [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Powering visual AI with FiftyOne Built on a powerful open source core, FiftyOne surfaces the data insights needed to build high-quality computer vision datasets and robust models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/e5f53bf5142f3960cbcaa8fd84fc2a03b6089327-2560x961.png?auto=format&dpr=2&fit=max&q=75&rect=682,0,1495,961&w=748) ![](https://cdn.sanity.io/images/h6toihm1/production/d286d778ffac5e30c2af62755808bf566dc5d3b6-2048x1148.webp?auto=format&dpr=2&fit=max&q=75&rect=412,301,1636,847&w=818) ![](https://cdn.sanity.io/images/h6toihm1/production/4b11ac1529f6fd395f12053d7f1da33c43aeb416-451x524.png?auto=format&dpr=2&fit=max&q=75&rect=44,0,323,376&w=162) ![](https://cdn.sanity.io/images/h6toihm1/production/3ca1b330c63e537c1217fd35d2e7cdf338c9637c-1600x1600.png?auto=format&dpr=2&fit=max&q=75&rect=71,822,1422,778&w=711) Why FiftyOne ## Poor quality data is the biggest obstacle to AI success FiftyOne is the refinery for building visual AI that harmonizes data and models to create quality datasets and accurate models. Reveal data gaps, spot failure modes, and streamline workflows to maximize model performance. customers ## Loved by developers 0+ product installs 0k GitHub stars 0+ enterprise customers ![](https://cdn.sanity.io/images/h6toihm1/production/cd7fe79ef465aa75489f4c049ce2af346f0985b4-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/720f6702611411baf6a274b1010856c78c07235d-240x96.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/57564abbae97325ddeb3bc3b8b7b5690d61cddc5-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/be26ca6233b020b9201ddd482f7d34f53555b51d-276x90.png?auto=format&dpr=2&fit=max&q=75&w=100) Maintained 99% fall detection rates for model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/636eaa49ef62ea3d07ca264f7baa4fe0e494e125-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/3911d157da27a0e996468bf45393d4e01e522934-384x39.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/a54f5857074a8e45684334b38c3e4f1f1528274e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/e845660699edf4deed110e60425659ba8ae73e6e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Foundation for Florence-2 VLM development](https://voxel51.com/plugins/?search=florence) ![](https://cdn.sanity.io/images/h6toihm1/production/c8be8bb13c2be64cd70dda28334501a812eb0bc1-520x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/8c596efcf17de5a9bc02b05b44f55474802f0abe-90x90.png?auto=format&dpr=2&fit=max&q=75&w=90) ![](https://cdn.sanity.io/images/h6toihm1/production/1f4c7beab6c76544152d4f3d990059bc62eb90cb-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/aee0fde76666dd2d134875eb8d5247ee75dab896-153x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cb778b338fac262a1c8a9de0334b256747cfbc5b-500x261.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Eliminated repetitive manual transformations on 20 TB+ of visual data](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/2eed779bf08ffa8ea43cb822cb013674cc05e546-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/6f935624676c89371a1b094385b7b60439e09824-351x60.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/09890f00ed8e1cb9be4a30a0a0a52a80b2d74fee-800x1220.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cbbc7f79176fa84b6e2d1f15eb76b27b84f216b9-393x128.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg)\\ \\ Official partner for visualizing Open Images Dataset V7](https://voxel51.com/blog/exploring-google-open-images-v7/) ![](https://cdn.sanity.io/images/h6toihm1/production/cd7fe79ef465aa75489f4c049ce2af346f0985b4-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/720f6702611411baf6a274b1010856c78c07235d-240x96.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/57564abbae97325ddeb3bc3b8b7b5690d61cddc5-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/be26ca6233b020b9201ddd482f7d34f53555b51d-276x90.png?auto=format&dpr=2&fit=max&q=75&w=100) Maintained 99% fall detection rates for model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/636eaa49ef62ea3d07ca264f7baa4fe0e494e125-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/3911d157da27a0e996468bf45393d4e01e522934-384x39.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/a54f5857074a8e45684334b38c3e4f1f1528274e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/e845660699edf4deed110e60425659ba8ae73e6e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Foundation for Florence-2 VLM development](https://voxel51.com/plugins/?search=florence) ![](https://cdn.sanity.io/images/h6toihm1/production/c8be8bb13c2be64cd70dda28334501a812eb0bc1-520x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/8c596efcf17de5a9bc02b05b44f55474802f0abe-90x90.png?auto=format&dpr=2&fit=max&q=75&w=90) ![](https://cdn.sanity.io/images/h6toihm1/production/1f4c7beab6c76544152d4f3d990059bc62eb90cb-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/aee0fde76666dd2d134875eb8d5247ee75dab896-153x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cb778b338fac262a1c8a9de0334b256747cfbc5b-500x261.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Eliminated repetitive manual transformations on 20 TB+ of visual data](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/2eed779bf08ffa8ea43cb822cb013674cc05e546-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/6f935624676c89371a1b094385b7b60439e09824-351x60.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/09890f00ed8e1cb9be4a30a0a0a52a80b2d74fee-800x1220.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cbbc7f79176fa84b6e2d1f15eb76b27b84f216b9-393x128.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg)\\ \\ Official partner for visualizing Open Images Dataset V7](https://voxel51.com/blog/exploring-google-open-images-v7/) ![](https://cdn.sanity.io/images/h6toihm1/production/cd7fe79ef465aa75489f4c049ce2af346f0985b4-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/720f6702611411baf6a274b1010856c78c07235d-240x96.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/57564abbae97325ddeb3bc3b8b7b5690d61cddc5-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/be26ca6233b020b9201ddd482f7d34f53555b51d-276x90.png?auto=format&dpr=2&fit=max&q=75&w=100) Maintained 99% fall detection rates for model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/636eaa49ef62ea3d07ca264f7baa4fe0e494e125-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/3911d157da27a0e996468bf45393d4e01e522934-384x39.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/a54f5857074a8e45684334b38c3e4f1f1528274e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/e845660699edf4deed110e60425659ba8ae73e6e-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Foundation for Florence-2 VLM development](https://voxel51.com/plugins/?search=florence) ![](https://cdn.sanity.io/images/h6toihm1/production/c8be8bb13c2be64cd70dda28334501a812eb0bc1-520x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/8c596efcf17de5a9bc02b05b44f55474802f0abe-90x90.png?auto=format&dpr=2&fit=max&q=75&w=90) ![](https://cdn.sanity.io/images/h6toihm1/production/1f4c7beab6c76544152d4f3d990059bc62eb90cb-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/aee0fde76666dd2d134875eb8d5247ee75dab896-153x96.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/9d107edd3dcfaf321e55af64ced1ff6f0647d484-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cb778b338fac262a1c8a9de0334b256747cfbc5b-500x261.png?auto=format&dpr=2&fit=max&q=75&w=100)\\ \\ Eliminated repetitive manual transformations on 20 TB+ of visual data](https://voxel51.com/blog/rios-ai-powered-robotics-run-on-fiftyone-teams/) ![](https://cdn.sanity.io/images/h6toihm1/production/2eed779bf08ffa8ea43cb822cb013674cc05e546-520x680.png?auto=format&dpr=2&fit=max&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/6f935624676c89371a1b094385b7b60439e09824-351x60.png?auto=format&dpr=2&fit=max&q=75&w=100) ![](https://cdn.sanity.io/images/h6toihm1/production/09890f00ed8e1cb9be4a30a0a0a52a80b2d74fee-800x1220.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/cbbc7f79176fa84b6e2d1f15eb76b27b84f216b9-393x128.png?auto=format&dpr=2&fit=max&q=75&w=100) [![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=340&q=75&rect=1229,0,1795,961&w=260)![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg)\\ \\ Official partner for visualizing Open Images Dataset V7](https://voxel51.com/blog/exploring-google-open-images-v7/) fiftyone platform [Smarter Annotation](https://voxel51.com/annotation) [Data Curation & Management](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) Deploy anywhere ![](https://cdn.sanity.io/images/h6toihm1/production/79fa40b25cb38d7c9c85b3f565dcda26356a4f15-1276x634.png?auto=format&dpr=2&fit=max&q=75&w=640) Fully customizable and extensible ![](https://cdn.sanity.io/images/h6toihm1/production/c5d011a5d71e23096ce07385b12ddebe5e66cb42-1268x634.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1268,634&w=640) Support for billions of samples ![FiftyOne supports billions of samples and metadata](https://cdn.sanity.io/images/h6toihm1/production/9c891a7d08584ac04cb8f033bae7bc7c3941778b-951x1152.png?auto=format&dpr=2&fit=max&q=75&w=317) Dataset versioning ![FiftyOne includes robust data versioning capabilities](https://cdn.sanity.io/images/h6toihm1/production/e9f5d5d86653a3ccf36dc923a034c091cdf99ff1-951x1152.png?auto=format&dpr=2&fit=max&q=75&w=317) Role-based access controls ![](https://cdn.sanity.io/images/h6toihm1/production/ad606a31d0b9b15441e8528aa3d8836e307083ee-951x1152.png?auto=format&dpr=2&fit=max&q=75&w=317) ISO 27001 certification ![FiftyOne is ISO 27001 certified](https://cdn.sanity.io/images/h6toihm1/production/c498202dd43ae6504021355074de1d97950b5daa-951x1152.png?auto=format&dpr=2&fit=max&q=75&w=317) Annotation Overview ## Verified Auto Labeling Get auto-labeling combined with intelligent QA to deliver human-quality labels at a fraction of the price. [Explore data labeling](https://voxel51.com/annotation) Curation Overview ## Data curation and management Curate high-quality visual datasets with tools that surface insights, uncover issues, and optimize data for model performance. [Explore data curation](https://voxel51.com/curation) ### **Data Quality** Detect label errors, uncover biases, and fill coverage gaps. FiftyOne gives you the visibility to improve dataset quality and ensure your models are learning from the right data. ![Voxel51 Data Quality - Exact Duplicates](https://cdn.sanity.io/images/h6toihm1/production/c816227ddef10ae3a71b750a46301a49361be8f2-2226x2226.png?auto=format&dpr=2&fit=max&q=75&w=640) ### **Embeddings & Data Visualizations** Explore your dataset like never before with powerful embeddings and interactive visualizations. Quickly surface patterns, clusters, and edge cases to guide model development and refinement. ![Computer vision embeddings](https://cdn.sanity.io/images/h6toihm1/production/eda6b1f41ad09ea94312fe8902f8837fa89cb74a-1200x1200.png?auto=format&dpr=2&fit=max&q=75&w=640) MODEL EVALUATION ## Understand model strengths and weaknesses to improve performance Assess overall model performance and inspect individual samples — all in one interactive workflow. Identify failure modes, bias, and blind spots with ease. [Explore model evaluation](https://voxel51.com/evaluation) ![](https://cdn.sanity.io/images/h6toihm1/production/3e795ee996356246388c3bad919ae3363db07ec4-2954x1206.png?auto=format&dpr=2&fit=max&q=75&w=1477) #### Scenario analysis Compare model metrics on multiple models to understand which model performs best and why #### Sample-level analysis Uncover model failure modes by conducting sample-level analysis to view the best/worst-performing samples and edge cases **Performance metrics** Evaluate and analyze model detections and aggregate metrics such as mAP, precision, recall, f1 scores, and more INTEGRATIONS ## Integrate with your existing ML stack [Explore integrations](https://voxel51.com/integrations) ![](https://cdn.sanity.io/images/h6toihm1/production/4262375676fd146768064041982c53f191063e6d-2560x1200.jpg?auto=format&dpr=2&fit=max&q=75&w=1280) > “FiftyOne is our primary resource for machine learning research. Thanks to FiftyOne's convenient field visualizations and filtering capabilities, we can easily distinguish incorrect labels and predictions, and therefore iterate on models faster than ever. As a result, we've achieved a 77% reduction in images sent for manual verification.” > > **Ryan Szeto** > > Senior Computer Vision Engineer at SafelyYou ![](https://cdn.sanity.io/images/h6toihm1/production/2b86407c4eee830da8897a6d94e08ad7f776c698-2550x834.png?auto=format&dpr=2&fit=max&q=75&rect=42,117,2439,582&w=100) > “As we developed our Florence-2 model, FiftyOne proved invaluable for data management and visualization. Its powerful capabilities helped streamline our workflow, ensuring we built a robust foundation for our models. Now, as we dive into the development of Florence-5B, we're relying on FiftyOne more than ever. The tool's intuitive interface and rich feature set are essential for effectively managing our large datasets and gaining critical insights.” > > **Bin Xiao** > > AI Researcher, Meta (formerly Principal Research Manager, Microsoft GenAI) ![](https://cdn.sanity.io/images/h6toihm1/production/d9b4bdb75662f7f850d7abbf017918438f9f4885-3333x713.png?auto=format&dpr=2&fit=max&q=75&w=100) > “FiftyOne has helped us speed up investigations by 3x. For example, if we see a wrong suction cup grasping an item, we can quickly visualize the issue across all data sources and identify what went wrong.” > > **Dimitry Pechyoni** > > Senior Principal Machine Learning Engineer at Berkshire Grey ![](https://cdn.sanity.io/images/h6toihm1/production/30b0acc56eeba0fd43bf86f4a390dd985d1916cd-640x152.png?auto=format&dpr=2&fit=max&q=75&w=100) > "FiftyOne enables researchers to analyze and improve the quality of their datasets rapidly, replacing the weeks of manual labor that would otherwise be required without this technology. High-quality data is critical to the success of machine learning systems. Without the right tools to analyze and curate datasets, machine learning development can be inefficient and ineffective.” > > **Jordi Pont-Tuset** > > Research Scientist at Google ![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg) > “At Allstate, my team works on auto vehicle damage inspection. Verifying the damage to a vehicle can take an insurance claim agent hours to verify, but using computer vision and FiftyOne, we can segment the parts of vehicles first, then detect the damages, and finally match the damage to repair costs and generate reports for the adjusters.” > > **Pavan Nanjundappa** > > Data Science Manager, Allstate India ![](https://cdn.sanity.io/images/h6toihm1/production/876897e6347fb3fae892d3c91ff18031dddb2aef-600x132.svg) > “We use FiftyOne to organize large research datasets. My favorite feature is the ability to view distributions over image attributes in the dataset, and filter the dataset by those attributes.” > > **Brett Israelsen** > > Principal Research Scientist, AI, Raytheon ![](https://cdn.sanity.io/images/h6toihm1/production/293dc34ac3ccdd7973ac04b5a27e66a0c0862260-158x61.svg) > “What really stands out about FiftyOne is the flexibility. The plugin framework lets us customize our workflows based on our unique needs, and the mature SDK lets us consolidate more of our pipeline into one tool, avoiding the cost of stitching together multiple systems. FiftyOne integrates directly into our production pipeline to drive 80% reductions in workplace incidents .” > > **Patrick Rowsome** > > Head of Computer Vision Operations, Protex AI ![](https://cdn.sanity.io/images/h6toihm1/production/edafe4cdb4e8e73984ea47990c1625dd354071d3-1602x365.png?auto=format&dpr=2&fit=max&q=75&w=100) ## Join our AI, ML, and data science community Join more than 20,000 developers and scientists around the world who are leveling up their AI skills in
GenAI, Computer Vision, LLMs, MLOps, vector search, and more. [Join upcoming events](https://voxel51.com/events) ![](https://cdn.sanity.io/images/h6toihm1/production/dea067d4cded2a986cc0f4173ec818c64e0d57bc-925x444.svg) ## Enough data wrangling.
 Request a demo. [Book a demo](https://voxel51.com/sales) [Check out pricing](https://voxel51.com/pricing) ![](https://cdn.sanity.io/images/h6toihm1/production/ea42e9b26f49f1cb54bb8aca31dc10e7f74fe11f-3024x960.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-41-lllmstxt|> ## Visual AI in Sports [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Visual AI in Sports Build solutions that delight fans and give teams and athletes the winning edge. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/37d4804ae9cd901cf2b6ec022455381b253c5c6c-628x395.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/2712b7d006baefb08ca2a8e612cfde9a96923468-628x817.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/dba6d318ad01ffcec31bfe0a058e039ef5260fd1-628x813.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/17c367d2ed23c23aa294dcd8974a2b3b7a2bf000-628x393.png?auto=format&dpr=2&fit=max&q=75&w=314) Benefits ## Take the game to the next level with solutions powered by computer vision Revolutionize how teams, coaches and athletes elevate performance and how fans engage with vision-based innovations built using FiftyOne. ### Win more matches Visualize and analyze every jump, sprint, throw, and maneuver to gain an advantage over the competition ### Improve player safety Use data-centric AI to better detect and prevent potential injuries and improve protective gear and equipment ### Revolutionize the game Build AI-driven tools to transform experiences, both on and off the field, and bring people even closer to the sports they love Use cases ## Sports use cases powered by Visual AI Computer vision plays a critical role in a growing array of use cases in sports. That’s why leaders and innovators building retail solutions rely on FiftyOne. ![](https://cdn.sanity.io/images/h6toihm1/production/bd159f0df96491b16491c4c5496f242095ffae2d-768x768.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Sports analytics and strategy Optimize strategies, identify performance trends, and make data-driven decisions by analyzing footage and extracting insights using computer vision. ![](https://cdn.sanity.io/images/h6toihm1/production/0517668d2f12d22928908fe54b50d65be5c17e61-768x432.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Injury prevention and rehabilitation Identify injury risks and create personalized rehabilitation plans that improve safety and recovery outcomes with computer vision. ![](https://cdn.sanity.io/images/h6toihm1/production/228e83221e0ef48c6d530dcf6441f951130744ed-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Referee assistance Use real-time insights on player positioning, ball trajectory, and line calls to reduce human error and improve accuracy. ![](https://cdn.sanity.io/images/h6toihm1/production/ccea5b413c99a6ea9d8baaab1d9cb81139dca93c-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Fan experience enhancement Provide immersive viewing experiences, real-time statistics and personalized content to make sports more engaging and interactive. ![](https://cdn.sanity.io/images/h6toihm1/production/eb2c8a162f40b9bfa08b8d8b32b447068ee7a851-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Broadcast enhancement Provide real-time analysis, immersive graphics, and enhanced viewing experiences powered by computer vision. ![](https://cdn.sanity.io/images/h6toihm1/production/11616532edb66aa6666b152c4f4086d37f8ffb31-768x432.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Personalized fitness and training Analyze real-time movement, provide tailored feedback, and create customized workout plans to optimize performance and prevent injuries. Features ## How visual Al can help you ### Visualize, sort, and select the samples you need Visualize in-game data from every angle and easily find a specific player, match, or any action you specify with FiftyOne. ![](https://cdn.sanity.io/images/h6toihm1/production/e675c9f2c517e93954f3b8819bb12b03f81b7876-2560x1806.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Analyze and get actionable insights into your sports data Use state of the art computer vision techniques to perform video analysis, determine player speed, acceleration, separation distance, and more. Track players during the game to find total distance traveled, figure out shot angles, and determine how to prevent injuries. ![](https://cdn.sanity.io/images/h6toihm1/production/039530715f22da0bf342f7a1048108d5fa94bf8b-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Continuously build models that perform Use FiftyOne to evaluate several different models over several different datasets to find the highest performing ones for your needs and get an edge up on your competition. ![](https://cdn.sanity.io/images/h6toihm1/production/09ea06899eb7f144117c8d3a0995fdd251cf88a3-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Unlimited flexibility to capture every moment Add as much visual data as you want, millions of samples, with as many labels and fields as you need. Use FiftyOne’s label features such as pose estimation, tracking, temporal events, and custom fields to capture every aspect of the game. ![](https://cdn.sanity.io/images/h6toihm1/production/4c618700944ef8bee9d0e71767ee105199d218f0-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) Resources ## Learn more about visual AI in sports Learn how leading organizations are using visual AI to revolutionize the game for teams, coaches, athletes, and fans. ### How computer vision is changing sports Applying computer vision and AI to sports opens up exciting new technologies for teams, athletes, coaches, sports analytics teams, and fans. [Read more](https://voxel51.com/blog/how-computer-vision-is-changing-sports) ![](https://cdn.sanity.io/images/h6toihm1/production/d94b00bbfadfd5cf59c4debb3d48da6eea3a3a4f-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=960) ## Data eats models for lunch Talk to our computer vision experts to start building better datasets and models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-42-lllmstxt|> ## Fast Code AI Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/60895f85f1668e61f0ed0032a3075ac73d5d0918-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Fast Code AI Fast Code AI relies on FiftyOne for managing its data lifecycle Apr 3, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Fast Code AI](https://www.fastcode.ai/) builds computer vision solutions for driverless cars, sports analytics, inspection, and surveillance – empowering the world’s largest companies to interface with the visual world. > "Fast Code AI relies on FiftyOne for managing our data lifecycle. Upon receiving initial data from our data labeling partner, IndiVillage, we load it into FiftyOne and scrutinize the labeling for any initial inconsistencies. We then provide feedback and obtain the next batch of data. This is where we train our baseline models and identify any outliers. We sort the data by their ground truth class confidence values in increasing order and display it in FiftyOne. This helps us easily spot and correct any incorrect labels. With approximately 30 attributes per image and over 40 possible values for each attribute, visualizing the data elsewhere is challenging. We greatly appreciate FiftyOne’s contribution to our workflow, as it has significantly increased our efficiency. We cannot imagine managing our data without this tool, as it offers an unmatched level of customization and functionality." – Arjun Jain, Founder and Chief Scientist at Fast Code AI ![](https://cdn.sanity.io/images/h6toihm1/production/810731f9e7e69d3454c8a7064af94931fd4f272f-512x288.jpg?auto=format&dpr=2&fit=max&q=75&w=512) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-43-lllmstxt|> ## Boston AI/ML Workshop [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/0cd905823479352908c60eb3a07cb4230442a758-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/0cd905823479352908c60eb3a07cb4230442a758-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event In-person Americas Meetups Boston AI, ML and Computer Vision Meetup - September 25, 2025 Sep 25, 2025 5:30 - 8:00 PM Microsoft Research Lab – New England (NERD) at MIT Deborah Sampson Conference Room One Memorial Drive, Cambridge, MA, 02142 About this event This hands-on workshop explores computer vision techniques for automotive damage detection using the [CarDD dataset](https://cardd-ustc.github.io/) \- the largest public dataset specifically designed for vehicle damage analysis. Participants will learn how to leverage [FiftyOne](https://docs.voxel51.com/), a powerful computer vision experimentation platform, to explore, visualize, and build models for detecting six common types of vehicle damage: dents, scratches, cracks, glass shatter, tire flats, and broken lamps. The workshop consists of five comprehensive modules that guide participants from initial setup to advanced model deployment: 1. **Environment Setup:** Configure a Python environment with essential libraries including FiftyOne, PyTorch, and other dependencies needed for computer vision tasks. 2. **Dataset Exploration:** Load and analyze the CarDD dataset, which contains 4,000 high-resolution images with over 9,000 expertly annotated instances of vehicle damage. Explore dataset statistics, annotation formats (COCO), and understand the dataset's split into training, validation and test sets. 3. **Visual Embeddings Analysis:** Implement state-of-the-art vision models (CLIP, SigLIP) to gain deeper understanding of the dataset. Visualize relationships between damage types using dimensionality reduction techniques and explore similarities through natural language search capabilities. 4. **Model Evaluation:** Learn techniques for evaluating model performance on instance segmentation and detection tasks. Export data for training and evaluate models using FiftyOne's built-in evaluation tools. 5. **Extended Functionality:** Enhance your workflow using the FiftyOne plugin ecosystem for specialized visualizations, search capabilities, and integration with other tools. Deploy pre-trained "zoo" models for real-world car damage detection applications. This workshop is designed for computer vision practitioners, automotive industry professionals, and data scientists interested in applying modern deep learning techniques to vehicle damage detection. Participants will gain practical experience working with a production-grade dataset while learning best practices for model development, evaluation, and deployment. Host ![](https://cdn.sanity.io/images/h6toihm1/production/2b9d150f45e9d15e9fd4f7867032cb29b0aa566f-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Harpreet Sahota Voxel51 Bio [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-44-lllmstxt|> ## Import Kaggle Datasets [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tutorials](https://voxel51.com/blog/category/tutorials) Import Kaggle Datasets into FiftyOne & Publish to Hugging Face Hub — Step‑by‑Step Tutorial with ASL‑MNIST Jul 17, 2025 • 7 min read Article content In this article [Quick Start: Load ASL‑MNIST into FiftyOne](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#c63ba0fd5e68) [Step 1: Set Up Your Python Environment](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#7f21a6920de4) [Step 2: Configure the Kaggle API Credentials](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#4aae249e9f49) [Step 3: Download & Process the ASL‑MNIST Dataset](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#6ce0607c111e) [Quirks of ASL-MNIST](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#60fa1f99fbb1) [Step 4: Build a FiftyOne Dataset from ASL‑MNIST Images](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#88dbcedfe91b) [Step 5: Explore & Visualize the Dataset in FiftyOne](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#7b81b57d1a08) [Step 6: Save the FiftyOne Dataset Locally](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#c4db7c69b057) [Step 7: Publish the Dataset to Hugging Face Hub](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#e8ebc8322cf5) [Step 8: Verify & Reload the Dataset from Hugging Face Hub](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#455a573e5171) [Pro Tip: Move the Dataset to a Hugging Face Organization](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#26dd37e79780) [Full Google Colab Notebook and Source Code](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#e739f06c27df) [Key Takeaways](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#7a475b2c7c65) [Next steps](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#b602357c380c) In this article [Quick Start: Load ASL‑MNIST into FiftyOne](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#c63ba0fd5e68) [Step 1: Set Up Your Python Environment](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#7f21a6920de4) [Step 2: Configure the Kaggle API Credentials](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#4aae249e9f49) [Step 3: Download & Process the ASL‑MNIST Dataset](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#6ce0607c111e) [Quirks of ASL-MNIST](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#60fa1f99fbb1) [Step 4: Build a FiftyOne Dataset from ASL‑MNIST Images](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#88dbcedfe91b) [Step 5: Explore & Visualize the Dataset in FiftyOne](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#7b81b57d1a08) [Step 6: Save the FiftyOne Dataset Locally](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#c4db7c69b057) [Step 7: Publish the Dataset to Hugging Face Hub](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#e8ebc8322cf5) [Step 8: Verify & Reload the Dataset from Hugging Face Hub](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#455a573e5171) [Pro Tip: Move the Dataset to a Hugging Face Organization](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#26dd37e79780) [Full Google Colab Notebook and Source Code](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#e739f06c27df) [Key Takeaways](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#7a475b2c7c65) [Next steps](https://voxel51.com/blog/import-kaggle-datasets-into-fiftyone-and-publish-to-hugging-face-hub#b602357c380c) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) In this tutorial, you will learn how to prepare and visualize a dataset using FiftyOne after obtaining data from Kaggle. We will use the [American Sign Language MNIST (ASL-MNIST) dataset](https://www.kaggle.com/datasets/datamunge/sign-language-mnist) as our example. FiftyOne helps you curate high-quality datasets by combining code-based analysis with visual exploration. With it we can filter samples, remove duplicates, fix annotations, and add metadata programmatically. Then we can launch an interactive app to explore our results visually and slice or aggregate the data as we’d like. ### What You will Need - A virtual environment with Python 3.9-3.11 installed (or willingness to use a Google Colab notebook) - [Kaggle](https://www.kaggle.com/) account with API access - [HuggingFace](https://huggingface.co/) account - Basic familiarity with Pandas and PIL ### Time Required 90 minutes start-to-finish ### Installing FiftyOne Local FiftyOne installation instructions are [available here](https://docs.voxel51.com/getting_started/install.html). You can also follow the tutorial using [this Google Colab notebook](https://githubtocolab.com/andandandand/practical-computer-vision/blob/main/notebooks/ASL_from_Kaggle_to_FiftyOne.ipynb). ## Quick Start: Load ASL‑MNIST into FiftyOne If you just want to [get the dataset](https://huggingface.co/datasets/Voxel51/American-Sign-Language-MNIST) and explore it in FiftyOne (and already have a HuggingFace account set up), you can simply execute: This is the artifact that we will produce at step 8 of this tutorial: a FiftyOne dataset published on HuggingFace Hub with all the images from the test and training set. To visualize it, we can start the FiftyOne app Try filtering the samples by label, here we get a view of images that share the label “V” However, if what you want is to learn how to get your data from Kaggle to FiftyOne and HuggingFace Hub, please keep reading. There are only eight steps ahead :) Even though this dataset is already available on HuggingFace, this tutorial is valuable if you want to upload your own datasets or understand the underlying steps. ## **Step 1: Set Up Your Python Environment** Begin by installing the necessary libraries: We pin versions to ensure the code in this tutorial runs exactly as shown. For your own projects, you may be able to use more recent versions. These versions are known to be compatible with each other as of July 2025. ## **Step 2: Configure the Kaggle API Credentials** To access datasets from Kaggle: 1. Log into Kaggle and navigate to your account settings. 2. Generate a new API token (kaggle.json). 3. Move the downloaded file to ~/.kaggle/ and set permissions: ## **Step 3: Download & Process the ASL‑MNIST Dataset** Download and extract the dataset: ## Quirks of ASL-MNIST The images are _not_ stored as JPEG or PNG files but as rows inside two CSV files. This format is unusual and I found it unintuitive, but is done on variants of the MNIST dataset. Each row of the csv file encoding the training set has 784 pixel intensity values in uint8 and a column indicating the ground truth label. The ground truth label is the letter that the image of the gesture is representing. Gestures in the dataset do not represent numbers. These images are **_really tiny_**. They only have 28x28 pixels and that’s where the “MNIST” part of the name comes from. The images have the same dimensions as the [original MNIST dataset](https://docs.voxel51.com/dataset_zoo/datasets.html#mnist) of grayscale handwritten digits. FiftyOne supports importing datasets using more standard formats, such as COCO or having directory structures where the name of the folder maps to the label. In this case, we need to do some manual processing to get around the quirks of ASL-MNIST. Test images will have the label “unknown”. Knowing all this, we process the CSV files into images with pandas and PIL: ## **Step 4: Build a FiftyOne Dataset from ASL‑MNIST Images** We import the processed jpg images into a [FiftyOne Dataset](https://docs.voxel51.com/user_guide/using_datasets.html) and map numerical labels to their corresponding letters. **Note:** The letters **'J' (9)** and **'Z' (25)** are not included in this dataset as they involve motion, and the images that we have are static (single frames). [Each image is a Sample](https://docs.voxel51.com/user_guide/using_datasets.html#samples) within the FiftyOne dataset. We can associate metadata, labels, and tags to each of them. Samples are initialized with a filepath to the corresponding data on disk. We had to save our images to the local hard drive to create the FiftyOne Dataset (a [collection](https://docs.voxel51.com/api/fiftyone.core.collections.html#module-fiftyone.core.collections) of samples). Note that asl\_dataset.persistent = True only affects in-memory persistence across Python sessions, not across system reboots. This is why we also export the dataset to disk in Step 6. ## **Step 5: Explore & Visualize the Dataset in FiftyOne** We want to launch the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html) to explore the dataset: We can now visualize, query, and analyze our dataset interactively using FiftyOne. The FiftyOne app is a powerful graphical user interface that allows us to browse, tag, aggregate, and interact directly with the dataset. ### Producing Histograms After getting the dataset into FiftyOne, try producing an aggregation of its ground\_truth.label field by going to the Histograms panel and selecting the field. You can then click on the “Split Horizontally” button to see the Samples next to the Histogram panel. Here is a short video demoing this. I encourage you to try creating histograms on other fields, such as metadata.size\_bytes. ## Step 6: Save the FiftyOne Dataset Locally To save our FiftyOne dataset for future use or for sharing, we can export it locally. The dataset.export() method creates a portable and self-contained archive and allows us to save the dataset in various formats. In this case, we will save it in the FiftyOneDataset format, which preserves the full FiftyOne dataset structure. Note that this local export is also needed to make the dataset available for ourselves after we turn off our computer. The persistence that we defined in step 4 (with asl\_dataset.persistent=True) is only valid across different Python sessions. With export\_media=True, our export ensures portability by copying the images into a self-contained folder. This will create a new directory named asl-mnist-fiftyone-dataset containing the JSON definition of your dataset and copies of the images. The export directory will contain both the label files AND copies of all the original jpg images. This creates a self-contained dataset that you can share or move without worrying about broken file paths. This exported dataset can be easily reloaded into FiftyOne later using fo.Dataset.from\_dir(). For more information on loading and using datasets, see the [documentation](https://docs.voxel51.com/user_guide/using_datasets.html). Following these steps, we have moved from a raw Kaggle dataset stored in CSV files to a portable, visual, and queryable dataset ready for interpretable computer vision workflows. ## Step 7: Publish the Dataset to Hugging Face Hub We can share the dataset with others through HuggingFace Hub, which is a great resource both for open models and data. When we push our FiftyOne dataset to the HF Hub, a fiftyone.yml is generated and the dataset remains in FiftyOne format, with all its fields: annotations, tags, and metadata. If you are new to HuggingFace, you can [follow this guide](https://github.com/andandandand/practical-computer-vision/blob/main/docs/huggingface-account-and-token.md) to set up your authentication token. You will need to have it set up with write permissions to publish your dataset. Be sure to check the data's licensing rights. The American Sign Language dataset has a [CC0 License](https://creativecommons.org/public-domain/cc0/), meaning that is in the public domain. The original content is CC0, and we apply an MIT license only to the packaging and any additional code or annotations. We can modify a public domain dataset and redistribute it under a compatible license (such as the [MIT license](https://opensource.org/license/mit)). Public domain works have no copyright restrictions, so anyone can use, modify, and redistribute them without permission. When we modify public domain content, we “create” a new work that we own the copyright to. We can license our modifications under MIT, but the original public domain portions remain **public domain** and cannot be relicensed. Remember to follow best practices to avoid complications: - Document your changes: Clearly state which portions are your modifications versus the original public domain content. - Apply the license correctly: Ensure your MIT license applies only to your contributions. - Verify the source: Double-check that the original dataset is truly in the public domain. After uploading the dataset, be sure to fill-in its dataset card with all the details on the data collection process. Dataset cards serve as documentation for how your dataset was collected, cleaned, and used. You can use [the one that we have created for this example](https://huggingface.co/datasets/Voxel51/American-Sign-Language-MNIST/blob/main/README.md) as a template. ## Step 8: Verify & Reload the Dataset from Hugging Face Hub Finally, we can check that the dataset is available on our HuggingFace user account. ## Pro Tip: Move the Dataset to a Hugging Face Organization A detail that might be important to you is that push\_to\_hub will only allow you to push to personal accounts, not organizations. To transfer the dataset to an organization, you will first need to push it to a personal account and then transfer ownership through the dataset’s page. ## Full Google Colab Notebook and Source Code You can find the full Google Colab notebook to run this example [through this link](https://github.com/andandandand/practical-computer-vision/blob/main/notebooks/ASL_from_Kaggle_to_FiftyOne.ipynb), the fully processed dataset is [on our HuggingFace organization page](https://huggingface.co/datasets/Voxel51/American-Sign-Language-MNIST). ## Key Takeaways Congratulations! You have successfully taken a dataset from Kaggle, processed it into a usable format, curated it within FiftyOne, and shared it with the community on Hugging Face Hub. This workflow is a powerful pattern for any computer vision project, enabling better data understanding, collaboration, and reproducibility. Try applying these steps to your own dataset and share it on HuggingFace Hub! ## Next steps In the following blog posts, we will go into training neural networks using the integration of FiftyOne and PyTorch. Be sure to try those techniques on this data! Meanwhile, explore the [FiftyOne documentation](https://docs.voxel51.com/) and our [Tutorials](https://docs.voxel51.com/tutorials/index.html). [Kaggle](https://voxel51.com/blog/tag/kaggle) [Hugging Face](https://voxel51.com/blog/tag/hugging-face) ![](https://cdn.sanity.io/images/h6toihm1/production/33b22e3f1c5c695dc77026bcf45e2aea73909e83-800x800.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Antonio Rueda-Toicen Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/8cb892c18f65a83d75021c76f891400ac219dcef-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The ML Menu for Model Selection: Hugging Face, Weights & Biases, and FiftyOne\\ \\ Tutorials\\ \\ • \\ \\ May 1, 2023](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/896cf6d1a00466cd9403a4af77cef2bddc75fe1e-1400x813.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ CVAT <> FiftyOne: Data-Centric Machine Learning with Two Open Source Tools\\ \\ Tutorials\\ \\ • \\ \\ Nov 30, 2022](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools) [![](https://cdn.sanity.io/images/h6toihm1/production/cfa6067062cae206570d98a5e688951723545822-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why FiftyOne is the pandas of Computer Vision\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ Nov 23, 2022](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-45-lllmstxt|> ## CVPR 2025 Highlights Day 1 [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision) The Best of CVPR 2025 Series – Day 1 May 29, 2025 • 10 min read Article content In this article [Building Smarter, Safer, and More Grounded Vision AI](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#ab5e2fbdc043) [Teaching AI to See the Unseeable — OpticalNet \[1\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#f4b132c4196f) [Compositional Reasoning You Can See — Factored Generative AI \[2\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#ce0c7c8a4794) [Teaching AI to See the Farm with Just a Handful of Images — Few-Shot Grounding DINO \[3\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#ef13fc903e6b) [Can Your AI Really Drive? — Drive4C Breaks It Down \[4\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#d36aaeae6999) [Why These Papers Matter — A New Era in Computer Vision](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#78be174eb83c) [What is next?](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#0e2fb01a4f8a) [References](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#dc05a483bd14) In this article [Building Smarter, Safer, and More Grounded Vision AI](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#ab5e2fbdc043) [Teaching AI to See the Unseeable — OpticalNet \[1\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#f4b132c4196f) [Compositional Reasoning You Can See — Factored Generative AI \[2\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#ce0c7c8a4794) [Teaching AI to See the Farm with Just a Handful of Images — Few-Shot Grounding DINO \[3\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#ef13fc903e6b) [Can Your AI Really Drive? — Drive4C Breaks It Down \[4\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#d36aaeae6999) [Why These Papers Matter — A New Era in Computer Vision](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#78be174eb83c) [What is next?](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#0e2fb01a4f8a) [References](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1#dc05a483bd14) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Building Smarter, Safer, and More Grounded Vision AI CVPR brings together some of the most exciting research in computer vision, but sometimes it’s hard to keep up, especially when great ideas are shared quickly or secret in complex papers. That’s why we created “The Best of CVPR,” a three-part [virtual meetup](https://voxel51.com/events/best-of-cvpr-july-9-2025)and blog series. We want to give more visibility to the people behind the work and help everyone see these projects’ potential beyond the conference. The goal is simple: to help more people understand how this research connects to real-life problems — things like farming, healthcare, driving, and imaging — and to highlight its potential to impact our communities positively. We want to show who’s doing this work, what they’re building, and the promising future that might come next. This blog is written in a clear and relaxed tone. We’re keeping things easy to follow, with highlights from four excellent papers, short summaries, and ample space to explore how we can learn from each other and work together. We invite you to be part of this collaborative journey. This is the first in a three-part series. We hope it inspires new ideas and opens the door for more conversations, collaboration, and space to lift this fantastic community. Please consider how your unique perspective and expertise could contribute to these exciting developments in computer vision. Check out [Day 2](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2) and [Day 3](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3) blog posts as well. ![](https://cdn.sanity.io/images/h6toihm1/production/ad18aeed61586213b5ea579aa60c8f17d9fc61ea-1306x628.png?auto=format&dpr=2&fit=max&q=75&w=1306) ## Teaching AI to See the Unseeable — OpticalNet \[1\] **_Paper Title:_** _[OpticalNet: An Optical Imaging Dataset and Benchmark Beyond the Diffraction Limit](https://deep-see.github.io/OpticalNet/assets/paper.pdf)._ **_Paper Authors:_** _Benquan Wang, Ruyi An, Jin-Kyu So, Sergei Kurdiumov, Eng Aik Chan, Giorgio Adamo, Yuhan Peng, Yewen Li, Bo An._ **_Institutions:_** _Nanyang Technological University, Skywork AI, Singapore, University of Southampton, The University of Texas at Austin._ What if we could image nanoscale structures invisible to traditional optics — without dyes, electron beams, or damage? OpticalNet delivers a transformative AI benchmark for breaking the diffraction limit, using modular deep learning and a first-of-its-kind dataset. ![](https://cdn.sanity.io/images/h6toihm1/production/af05572b3988c9a4f41d2783b5cddd4c5adf090a-1400x735.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ### What It’s About OpticalNet is the first AI benchmark designed to reconstruct ultra-tiny objects — smaller than light can typically resolve — from blurry diffraction images. By combining experimental and simulated data, it trains deep learning models to translate invisible light patterns into clear, interpretable images. ### Why It Matters Due to light’s wave nature, optical resolution is traditionally capped at ~200nm. This limits real-time, noninvasive imaging of nanoscale biological structures (like viruses) and nanomaterials. OpticalNet enables conventional microscopes to break that limit using AI without expensive, invasive add-ons, potentially revolutionizing biomedicine, materials science, and manufacturing imaging. ### How It Works - Data Collection: Real subwavelength objects are fabricated with Focused Ion Beam (FIB) on gold films, then scanned with a high-precision optical microscope to generate diffraction images. - Simulation Framework: A Python-based tool mimics optical wave propagation to generate synthetic data for model training and proof-of-concept validation. - Learning Task: Formulated as an image-to-image translation problem, where models learn to convert diffraction images into binary object images. - Evaluation: Predictions are stitched together to reconstruct full objects and tested on synthetic “Light” (curved shape) and Siemens Star (rotation benchmark) datasets. ### Key Result Transformer-based vision models outperform CNNs, successfully reconstructing complex subwavelength structures from diffraction images, even under experimental noise. The models trained on simple square blocks generalized well to unseen complex shapes, validating the “building block” approach. ### Broader Impact OpticalNet lays the groundwork for AI-powered subwavelength imaging using existing hardware, enabling affordable, non-invasive diagnostics and quality control in fields ranging from virology to semiconductor inspection. It bridges optical physics and computer vision, inviting interdisciplinary collaboration to push the boundaries of what light-based imaging can achieve. ## Compositional Reasoning You Can See — Factored Generative AI \[2\] **_Paper title:_**_[Nonisotropic Gaussian diffusion for realistic 3D human motion prediction](https://arxiv.org/abs/2501.06035)._**_Paper Authors:_** _Cecilia Curreli, Dominik Muhle, Abhishek Saroha, Zhenzhang Ye, Riccardo Marin, Daniel Cremers._ **_Institutions:_** _Technical University of Munich, Munich Center for Machine Learning_ What if AI could think like humans — separating what’s in a scene from where it is — to reason and create new images more logically and safely? This paper introduces a generative model that does just that. ![](https://cdn.sanity.io/images/h6toihm1/production/d283d191a2bc6708202aa6d2a487e4f021ff105a-1400x733.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ### What It’s About The paper presents SkeletonDiffusion, a novel latent diffusion model for probabilistic human motion prediction. Unlike prior approaches, which often generate implausible poses (like stretched or jittery limbs), SkeletonDiffusion introduces a nonisotropic Gaussian diffusion that better reflects the structure and relationships between human body parts. ### Why It Matters Predicting human motion accurately and realistically has critical implications for autonomous driving, robotics, virtual reality, healthcare, and human-computer interaction. This method improves realism in forecasted human poses and addresses a significant shortcoming of previous models: inconsistent or anatomically incorrect body movements. ### How It Works - SkeletonDiffusion learns to generate human motion in a latent space, using a nonisotropic diffusion process tailored to the skeleton’s kinematic structure. - The model emphasizes bone-aware motion synthesis, enforcing realism and diversity through architectural bias and improved training strategies. - The authors also critique commonly used diversity metrics, highlighting how some models gain higher diversity scores at the cost of anatomical accuracy. ### Key Result SkeletonDiffusion outperforms isotropic diffusion baselines across multiple benchmarks, including three real-world datasets, producing more plausible and diverse predictions. It sets a new state-of-the-art in balancing realism and diversity without compromising physical consistency. ### Broader Impact By incorporating structural awareness into generative modeling, SkeletonDiffusion enables safer, more accurate motion forecasting for critical applications such as assistive robotics, virtual avatars, and surveillance systems. It also opens the door for redefining evaluation standards in generative human motion modeling. ## Teaching AI to See the Farm with Just a Handful of Images — Few-Shot Grounding DINO \[3\] **_Paper title:_** _[Few-Shot Adaptation of Grounding DINO for Agricultural Domain](https://arxiv.org/abs/2504.07252)._ **_Paper Authors:_** _Rajhans Singh, Rafael Bidese Puhl, Kshitiz Dhakal, Sudhir Sornapudi._ **_Institutions:_** _Corteva Agriscience, Indianapolis, USA_ Why rely on expensive, labor-intensive annotations when AI can learn crop detection from just a few photos? This paper turns Grounding-DINO into a fast, prompt-free few-shot learner for agriculture. ![](https://cdn.sanity.io/images/h6toihm1/production/9b6ea937dbfc3467d5fccba96007ec18b7447860-1400x508.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ### What It’s About The paper introduces a lightweight, few-shot adaptation of the Grounding-DINO open-set object detection model, tailored explicitly for agricultural applications. The method eliminates the text encoder (BERT) and uses randomly initialized trainable embeddings instead of hand-crafted text prompts, enabling accurate detection from minimal annotated data. ### Why It Matters High-performing agricultural AI often demands large, diverse annotated datasets, which are expensive and time-consuming. This method rapidly adapts a powerful foundation model to diverse agricultural tasks using only a few images, reducing costs and accelerating model deployment in farming and phenotyping scenarios. ### How It Works - Grounding-DINO typically uses a vision-language architecture (image and text encoders). - This adaptation removes the BERT-based text encoder and replaces it with randomly initialized fine-tuned embeddings with minimal training images. - Only these new embeddings are trained, while the rest of the model remains frozen. - This simplified design avoids the complexities of manual text prompt engineering and dramatically reduces training overhead. ### Key Result Across eight agricultural datasets, including PhenoBench, Crop-Weed, BUP20, and DIOR, the few-shot method consistently outperforms: - Zero-shot Grounding-DINO, particularly in cluttered or occluded scenarios. - YOLOv11, by up to 24% higher mAP with just four training images. - Prior state-of-the-art few-shot detectors in remote sensing benchmarks ### Broader Impact This work presents a scalable and cost-effective way to deploy deep learning in agriculture, even with limited data. It demonstrates how foundation models can be tailored to real-world domains like plant counting, insect detection, fruit recognition, and remote sensing, making AI more accessible and valuable for sustainable and efficient farming practices. ## Can Your AI Really Drive? — Drive4C Breaks It Down \[4\] **_Paper title:_** _Drive4C: A Closed-Loop Benchmark on What Foundation Models Really Need to Be Capable of for Language-Guided Autonomous Driving._ **_Paper Authors:_** _Tin Stribor Sohn, Maximilian Dillitzer, Johannes Bach, Jason J. Corso, Tim Brühl, Robin Schwager, Tim Dieter Eberhardt, Eric Sax._ **_Institutions:_** _Dr. Ing. h.c. F. Porsche AG, University of Applied Science Esslingen, University of Michigan, Voxel51 Inc., Karlsruhe Institute of Technology_ As language-guided autonomous driving becomes more common, one big question remains: What exactly should foundation models understand to drive safely? Drive4C answers that. ### What It’s About The paper introduces Drive4C, a closed-loop benchmark that systematically evaluates multimodal large language models (MLLMs) for language-guided autonomous driving (E2E-AD). It isolates and tests four essential capabilities: semantic, spatial, temporal, and physical understanding, along with scenario anticipation and language-guided motion (LGM) ### Why It Matters Language-guided driving is an emerging AI paradigm, but existing evaluations miss critical skills needed for safe autonomy. Drive4C allows for fine-grained performance breakdowns, helping researchers understand and fix weaknesses in modern MLLMs — an essential step for making autonomous vehicles robust and trustworthy. ### How It Works - Drive4C is built on the CARLA simulator and: - Separates evaluation into two stages: (1) QA-based scenario understanding and (2) instruction-based driving execution. - Covers 380 scenarios with 165K QA pairs and 87 question templates. - Evaluates models using multiple-choice and free-form questions, scored with correctness and GPT-based similarity. - Adds driving performance scores based on compliance with natural language instructions (LGM). - Supports multimodal input (e.g., video, LiDAR, radar, GPS) and is compatible with real-world sensor setups ### Key Result All evaluated models (including GPT-4o, SmolVLM, Llama-3.2, DriveMM, and Dolphins) perform well in semantic understanding and scenario anticipation, but struggle significantly with spatial, temporal, and physical understanding, and fail at complex driving maneuvers in LGM. GPT-4o had the best score (0.3012), but all models were not ready for real-world use. ### Broader Impact Drive4C exposes the core weaknesses in current AI driving agents and provides a capability-driven framework to guide future model improvements. It advocates for structured inductive biases (like physical laws and spatial layout models) to move toward safe and generalizable autonomous systems. The benchmark is open-source and aims to become a standard for evaluating foundation models in autonomous driving. ## Why These Papers Matter — A New Era in Computer Vision These four CVPR 2025 papers, which cover fields as diverse as microscopy, farming, autonomous driving, and visual reasoning, share a common thread: interpretability over black boxes. Low-data solutions over brute force scale, domain-specific realism over generic accuracy. Together, they signal a shift in the computer vision landscape: - From scaling up to smart modularity - From end-to-end pipelines to compositional transparency - From curated datasets to real-world applications As AI moves deeper into high-stakes domains, we’ll need systems that build them. The ‘real-world applications’ here refer to the use of AI in fields such as healthcare, agriculture, and autonomous driving, where the ability to reason, adapt, and explain is crucial. CVPR 2025 is not just about seeing better. It’s about thinking better with vision. Read [Day 2](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2)and [Day 3](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3) of the Best of CVPR Series and register for the[virtual meetup](https://voxel51.com/events/best-of-cvpr-july-9-2025). ## What is next? If you’re interested in following along as I dive deeper into the world of AI and continue to grow professionally, feel free to connect or follow me on [LinkedIn](https://www.linkedin.com/in/paula-ramos-phd/). Let’s inspire each other to embrace change and reach new heights! You can find me at some [Voxel51 events](https://voxel51.com/events), or if you want to join this fantastic team, it’s worth taking a look at this page: [https://voxel51.app/careers](https://voxel51.com/careers) ![](https://cdn.sanity.io/images/h6toihm1/production/571af7476f44954ee67389529f23791f4d50ff97-990x990.png?auto=format&dpr=2&fit=max&q=75&w=990) ## References \[1\] B. Wang, R. An, J.-K. So, S. Kurdiumov, E. A. Chan, G. Adamo, Y. Peng, Y. Li, and B. An, “OpticalNet: An Optical Imaging Dataset and Benchmark Beyond the Diffraction Limit,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025. Temporal Link: [https://deep-see.github.io/OpticalNet/assets/paper.pdf](https://deep-see.github.io/OpticalNet/assets/paper.pdf) \[2\] C. Curreli, Z. Ye, D. Muhle, R. Marin, A. Saroha, and D. Cremers, “Nonisotropic Gaussian diffusion for realistic 3D human motion prediction,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025. Temporal link: [https://arxiv.org/abs/2501.06035](https://arxiv.org/abs/2501.06035) \[3\] R. Singh, R. B. Puhl, K. Dhakal, and S. Sornapudi, “Few-Shot Adaptation of Grounding DINO for Agricultural Domain,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025. Temporal link: [https://arxiv.org/abs/2504.07252v1](https://arxiv.org/abs/2504.07252v1) \[4\] T. S. Sohn, M. Dillitzer, J. Bach, J. J. Corso, T. Brühl, R. Schwager, T. D. Eberhardt, and E. Sax, “Drive4C: A Closed-Loop Benchmark on What Foundation Models Really Need to Be Capable of for Language-Guided Autonomous Driving,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025. [CVPR](https://voxel51.com/blog/tag/cvpr) [Visual AI](https://voxel51.com/blog/tag/visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/e926c07c7d1426c0fde8fdefa637c528d47b16f4-512x512.webp?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/c63b546423b6cec32ccccc7df6bc4e0fced0b1a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 2\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2) [![](https://cdn.sanity.io/images/h6toihm1/production/3a661345dfbc596f7118a7b8ec8375ea1f73d138-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 3\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3) [![](https://cdn.sanity.io/images/h6toihm1/production/b9e8bcb7b44ddb46a05144cd1e370b837e6aa401-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Visual Agents at CVPR 2025\\ \\ Computer Vision\\ \\ • \\ \\ May 28, 2025](https://voxel51.com/blog/visual-agents-at-cvpr-2025) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-46-lllmstxt|> ## Tokyo AI Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/080da3344bd6912fe4584874428d247baa7968b5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/080da3344bd6912fe4584874428d247baa7968b5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) In-person Meetups Tokyo AI, Machine Learning and Computer Vision Meetup - July 31, 2025 Jul 31, 2025 17:00 - 20:00 Grand Hyatt Tokyo Coriander Room 6 Chome-10-3 Roppongi, Minato City, Tokyo Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/f520ebf5dbb593e5333fadca31bdbbd2eb1cd043-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Hideo Yoshimi Microsoft Bio ![](https://cdn.sanity.io/images/h6toihm1/production/d4eb1c0c476964bf0b1966398b35a6e7ffc13a56-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Mate Szarvas NVIDIA Bio ![](https://cdn.sanity.io/images/h6toihm1/production/21ff1a46ec3d10712fc9a805d2cedb39baa586ad-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Keisuke Kamata Weights & Biases Bio ![](https://cdn.sanity.io/images/h6toihm1/production/789c4d2e94a74f14ce09ab95663668161039216f-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Dan Gural Voxel51 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/cf538618661ada1278bfac62a8ad1592bf0faa21-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Masaaki Kameyama Wavye Bio About this event ご登録の上、お席を確保してください! Schedule 自動運転開発の再発明: 生成AIによる創造と変革 ![](https://cdn.sanity.io/images/h6toihm1/production/f520ebf5dbb593e5333fadca31bdbbd2eb1cd043-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Hideo Yoshimi Microsoft Bio 本セッションでは、生成AIによって自動運転開発に革命をもたらすマイクロソフトの最先端アプローチを紹介します。 マイクロソフトの自動運転開発プラットフォーム「AvOps」は、シナリオ生成、データ処理、機械学習評価、次アクション抽出まで、開発ライフサイクルの全フェーズにAIを統合し、これまで数週間かかっていたワークフローを数時間で完了できるように変革します。 さらに生成AIによって要件定義、スプリント計画、実装、テストを加速し、かつてない開発スピードを実現するHypervelocity Engineeringの導入によってこれまでにないスピードの開発を可能にします。 加えて、複数の生成AIエージェントが協調し、開発ワークフローを強化・効率化するAgent AIフレームワークも紹介します。 実際のケーススタディを通じて、生成AIが自動運転開発にもたらす具体的なインパクトを探り、モビリティの未来をどう変えていくかについて議論します。 ニューラル再構築とワールド・ファウンデーション・モデルによる自律走行車開発の推進 ![](https://cdn.sanity.io/images/h6toihm1/production/d4eb1c0c476964bf0b1966398b35a6e7ffc13a56-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Mate Szarvas NVIDIA Bio 最先端のニューラル再構築とワールド・ファウンデーション・モデルが、シミュレーションワークフローを合理化し、AVモデルの開発、テスト、検証における重要な課題に対処することで、自律走行車開発にどのような変革をもたらすかをご覧ください。本セッションでは、NVIDIA NuRecおよびCosmosの画期的な技術と、次世代AVシステムのためのスケーラブルなシミュレーションパイプラインを可能にするVoxel51の統合を紹介します。 GenAI評価の世界 ![](https://cdn.sanity.io/images/h6toihm1/production/21ff1a46ec3d10712fc9a805d2cedb39baa586ad-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Keisuke Kamata Weights & Biases Bio 生成AIモデルは急速に進化しており、現在では幅広いタスクをサポートしています。その能力を理解するために、大規模な評価活動が行われています。Weights & Biases Japanでは、日本最大級の公開LLMリーダーボードとビジョン言語モデルリーダーボードを運営しています。また最近では、生成AIシステム全体の評価に関するホワイトペーパーも発表しています。本講演では、これらの取り組みから得られた知見を共有し、モデルやシステムの評価における我々の幅広いコミュニティの取り組みについて紹介します。 現代のドライビングデータセット ![](https://cdn.sanity.io/images/h6toihm1/production/789c4d2e94a74f14ce09ab95663668161039216f-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Dan Gural Voxel51 Bio 自律走行車ほど、物理AIを急速に前進させているものはないと考えています。本講演では、合成データ、NeRF、よりスマートなキュレーション、ベクトル検索を使用して、最先端のAVデータセットを構築する方法を説明します。Nvidia Omniverse、最新モデル、そして最高のツールを駆使した本講演にご期待ください。 拡張性が高いエンボディドAI:単一のエンドツーエンドの運転モデルでロンドンから東京を走破 ![](https://cdn.sanity.io/images/h6toihm1/production/cf538618661ada1278bfac62a8ad1592bf0faa21-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Masaaki Kameyama Wavye Bio 次なるAIの波となっている「エンボディドAI」が最初に提供する体験こそ、自動運転です。 Wayveの製品アーキテクチャ&セーフティ責任者から、同社のエンボディドAI技術とその商用化に向けたグローバル展開、地図を必要としないAV2.0のアプローチ、単一のエンドツーエンドのニューラルネットワークを用いてデータからの直接認識と判断によって運転を学習していくAIモデルが、どのように米国、カナダ、ドイツ、ヨーロッパ、さらに日本でほとんど事前学習がない状態、あるいは全く必要とせずに運転できるのかを、解説していきます。 この基盤モデルは、安全性を設計段階から考慮した安全対策と共に多様な市場からのデータによって、高解像度地図や高額なセンサーなしで迅速かつ汎化を促進させているのかも紹介します。最後に、消費者レベルでの自動運転の実装へと進む中、この動向が日本の自動車メーカーとモビリティエコシステムにとって何を意味するのかについても解説します。 [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-47-lllmstxt|> ## Visual AI for Retail [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Visual AI in Retail and commerce Struggling with visual data? FiftyOne helps you find the insights that matter, so you can spend more time building what your customers love. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/690e197d2b808dbac66488655fb530cc16ad1088-628x395.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/11258f76b4d5094596108464c5959f43a7b8d7b9-628x813.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/2d1da4a20c862887d570ed14999d247542311450-628x817.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/815c2d0cee4047adfba1f7e3eb09d66ded278835-628x393.png?auto=format&dpr=2&fit=max&q=75&w=314) Benefits ## Build retail and ecommerce solutions that engage and scale AI is helping retailers increase engagement, effectiveness, and efficiency. Retailers that need superior data quality and model performance rely on FiftyOne to innovate and differentiate. ### Optimize the value chain Get better insights into what shoppers want, and meet those needs more sustainably ### Delight your customers Personalize and optimize touchpoints with engaging and delightful experiences ### Stand out from the crowd Build and deploy powerful visual AI solutions to stay ahead of the competition Use cases ## Power any visual AI use case in retail Computer vision is changing how retail works. But, building effective AI solutions takes more than a one-size-fits-all approach. That’s why AI builders turn to FiftyOne. It gives them the power to build exactly what they need. ![](https://cdn.sanity.io/images/h6toihm1/production/7262d174d393895aef89fe7c027e0bc978cdc047-2000x1333.jpg?auto=format&dpr=2&fit=max&q=75&w=1000) Inventory management Monitor and analyze inventory levels in real time using visual data from cameras and other sensors. ![](https://cdn.sanity.io/images/h6toihm1/production/2eee4517df40403cf8da4261eb0290a6e4e5b3c0-1600x1067.jpg?auto=format&dpr=2&fit=max&q=75&w=800) Autonomous checkout Instantly recognize and tally products through the use of cameras, scanners, sensors, and object recognition concepts. ![](https://cdn.sanity.io/images/h6toihm1/production/ef43cdda3e565b77acf40efda76d9a5d82f0dec4-2000x1104.jpg?auto=format&dpr=2&fit=max&q=75&w=1000) Virtual try-ons Create virtual try-on experiences that shoppers can use anywhere, including the comfort of their own homes. ![](https://cdn.sanity.io/images/h6toihm1/production/cc4a7575575e0746af69655c492085f0d1975022-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Customer behavior analysis Get insights into shopper activity, including shopping patterns and product preferences. ![](https://cdn.sanity.io/images/h6toihm1/production/5316f34b24b8bdba194e866599ba8b872bd9dee2-768x436.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Product recommendations Tailor the shopping experience to an individual consumer’s style and preferences based on visual similarity and history. ![](https://cdn.sanity.io/images/h6toihm1/production/67019f6321fb0a901575b192bd7ed3229d7ca914-768x513.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Shoplifting prevention Detect anomalies, improve checkout processes, and ensure smooth operations to deliver safer, more efficient environments. Features ## How visual AI can help you ### Re-identification & tracking made easy Whether it is customers or products you are aiming to identify or recommend, FiftyOne allows you to manage all your tracking needs and provides builtin tools to evaluate model performance on re-identification. ![](https://cdn.sanity.io/images/h6toihm1/production/3c6bf9f30619fd6163f0084cb2034fbcc397e320-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Recommendation systems without limits Have images but don’t know how to store all 50 associated labels with each image? With FiftyOne, there’s no limit on how many attributes or descriptions you can add to each sample, allowing you to filter, sort, and slice through even the largest, most complex datasets with ease. ![](https://cdn.sanity.io/images/h6toihm1/production/f13680b665447aa258d725c8544b0983049a8039-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Effortless zero-shot detection Leverage the best zero-shot detection models available today with FiftyOne Plugins! Detect broken products, pre-label your checkout data, and more by applying zero shot detection to your data. ![](https://cdn.sanity.io/images/h6toihm1/production/aef419eb3858edc04de34a183b50aa52ebc00299-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Embeddings visualization in just a few clicks Compute embeddings on your dataset to unlock similarity and natural language search on your entire dataset. You can even visualize the embeddings to find the hidden structures of your dataset and understand how your model sees your data. ![](https://cdn.sanity.io/images/h6toihm1/production/5167813d75c1afc6a2bc7d2a9d822f9b5decfa22-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) Resources ## Learn more about visual AI in retail Learn how leading companies are using visual AI to enhance customer experiences and transform retail operations. ### How computer vision is changing retail Explore popular use cases for Computer Vision and AI in retail and get a glimpse into companies at the cutting edge of retail tech. [Read more](https://voxel51.com/blog/how-computer-vision-is-changing-retail) ![](https://cdn.sanity.io/images/h6toihm1/production/311ab9b40bad2e432ca993f4ac1c1f62b111adba-1024x576.png?auto=format&dpr=2&fit=max&q=75&w=512) ## Data eats models for lunch Talk to our computer vision experts to start building better datasets and models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-48-lllmstxt|> ## Brussels AI Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/4350c9376583d48cbe893d6181100c919685bce2-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/4350c9376583d48cbe893d6181100c919685bce2-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) In-person EMEA Meetups Brussels AI, ML and Computer Vision Meetup - July 15 Jul 15, 2025 5:30-8:30 PM Commons Hub - Rue de la Madeleine 51, Brussels Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/82986dbc4104b54ebf56cf67c3d188da25d311d7-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Kais Bedioui SEAFAR AI Bio ![](https://cdn.sanity.io/images/h6toihm1/production/57a3d610fcf365789c30937b03739b2d4f67e2f1-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Niels Rogge Hugging Face Bio ![](https://cdn.sanity.io/images/h6toihm1/production/e1993c0f69fc690ab5c2c68c9caa3abbfe2fca2a-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Joy Timmermans Secury360 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/5b9770ebf879163fd519d57717c074eab4e49cca-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Sahota Voxel51 Bio About this event Hear talks from experts on cutting-edge topics in AI, ML, and computer vision on July 15 Schedule Enhancing Data Quality with Zero-Shot Detections ![](https://cdn.sanity.io/images/h6toihm1/production/82986dbc4104b54ebf56cf67c3d188da25d311d7-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Kais Bedioui SEAFAR AI Bio Presenting the winning solution of the [Wake Vision Challenge 2025](https://www.edgeaifoundation.org/posts/challenge-edge-wake-vision-data-challenge-now-open) where we showcase how smart and efficient data curation and exploration, can generate good model performance from small subsets. Computer Vision Ecosystem at Hugging Face ![](https://cdn.sanity.io/images/h6toihm1/production/57a3d610fcf365789c30937b03739b2d4f67e2f1-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Niels Rogge Hugging Face Bio Niels will describe all things computer vision at Hugging Face, from models you can use out-of-the-box, models you can use to fine-tune on your specific data, as well as ongoing and upcoming trends Adding Temporal Data to Object Detection Models ![](https://cdn.sanity.io/images/h6toihm1/production/e1993c0f69fc690ab5c2c68c9caa3abbfe2fca2a-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Joy Timmermans Secury360 Bio Object detection models are used for a lot of variety of tasks. However they are limited in seeing only the current image making it hard to detect objects in non ideal condition like: poor lighting, motion blur, etc. In this talk we will look at adding temporal data into existing object detection models without adding a lot of overhead. And how Fiftyone makes it easy to compare the model trained on temporal data vs static image data Visual Agents: What it takes to build an agent that can navigate GUIs like humans ![](https://cdn.sanity.io/images/h6toihm1/production/5b9770ebf879163fd519d57717c074eab4e49cca-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Harpreet Sahota Voxel51 Bio We’ll examine conceptual frameworks, potential applications, and future directions of technologies that can “see” and “act” with increasing independence. The discussion will touch on both current limitations and promising horizons in this evolving field. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-49-lllmstxt|> ## AI, ML, Computer Vision Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/619dc7486f70359ac3237c38c591b68808c05386-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/619dc7486f70359ac3237c38c591b68808c05386-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Meetups Agriculture AI, ML and Computer Vision Meetup – June 19, 2025 This event has ended, but you can still catch up! Watch the on-demand recordings and register for our [future events.](https://voxel51.com/events) Jun 19, 2025 10:00 - 11:30 AM Pacific Virtually over Zoom! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/ded4911b758aae8bb0294887b9c660e377c97c7c-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Wolfgang Schulz Continental Bio ![](https://cdn.sanity.io/images/h6toihm1/production/e4b14b45a0cec57f9e99dedb28df266efa9ab4d8-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Ashshak Sharifdeen MBZUAI Bio ![](https://cdn.sanity.io/images/h6toihm1/production/4815e4a8e0735d5a4f73cf8b349fc79a093e79ec-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Xiongkun Linghu Beijing Institute for General Artificial Intelligence Bio ![](https://cdn.sanity.io/images/h6toihm1/production/afec08d42bcc55cccbe749af5b77284bbb3c33f5-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Dan Gural Voxel51 Bio About this event Join the Meetup to hear talks from experts on cutting-edge topics across AI, ML, and computer vision. Schedule Multi-Modal Rare Events Detection for SAE L2+ to L4 ![](https://cdn.sanity.io/images/h6toihm1/production/ded4911b758aae8bb0294887b9c660e377c97c7c-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Wolfgang Schulz Continental Bio A burst tire on the highway or a fallen motorbiker occur rarely and thus pose extra efforts to Autonomous vehicles. Methods to tackle such Edge cases in Road scenarios are explained. O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models ![](https://cdn.sanity.io/images/h6toihm1/production/e4b14b45a0cec57f9e99dedb28df266efa9ab4d8-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Ashshak Sharifdeen MBZUAI Bio We propose O-TPT, a method to improve the calibration of vision-language models (VLMs) during test-time prompt tuning. While prompt tuning improves accuracy, it often leads to overconfident predictions. O-TPT introduces orthogonality constraints on textual features, enhancing feature separation and significantly reducing calibration error across multiple datasets and model backbones. Advancing MLLMs for 3D Scene Understanding ![](https://cdn.sanity.io/images/h6toihm1/production/4815e4a8e0735d5a4f73cf8b349fc79a093e79ec-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Xiongkun Linghu Beijing Institute for General Artificial Intelligence Bio Recent advances in Multimodal Large Language Models (MLLMs) have shown impressive reasoning capabilities in 2D image and video understanding. However, these models still face significant challenges in achieving holistic comprehension of complex 3D scenes. In this talk, we present our recent progress toward enabling global 3D scene understanding for MLLMs. We will cover newly developed benchmarks, evaluation protocols, and methods designed to bridge the gap between language and 3D perception. Voxel51 + NVIDIA Omniverse: Exploring the Future of Synthetic Data ![](https://cdn.sanity.io/images/h6toihm1/production/afec08d42bcc55cccbe749af5b77284bbb3c33f5-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Dan Gural Voxel51 Bio Join us for a lightning talk on one of the most exciting frontiers in Visual AI: synthetic data. We’ll showcase a sneak peek of the new integration between FiftyOne and NVIDIA Omniverse, featuring fully synthetic downtown scenes of Santa Jose. NVIDIA Omniverse is enabling the generation of ultra-precise synthetic sensor data, including LiDAR, RADAR, and camera feeds, while FiftyOne is making it easy to extract value from these rich datasets. Come see the future of sensor simulation and dataset curation in action, with pixel-perfect labels to match. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-50-lllmstxt|> ## Embodied AI Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Embodied Computer Vision at CVPR 2025: The Next AI Frontier Jun 30, 2025 • 5 min read Article content In this article [1\. RoBoSpatial: Training AIto Reason About Space](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#3948f67b7f25) [2\. GROVE: Teaching Robots with Generalized Rewards](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#ade977ae5dab) [3\. Navigation World Models: Planning in Simulation](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#b0eb4f95a305) [4\. Gemini Robotics: Bridging Foundation Models and Physical Action](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#38c7e1e5c358) [Why Embodied AI is the Next Big Boom](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#7b0b43cdb036) [A Call to the Research Community: Prepare to Validate the Future](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#e24e18b72b16) [Final Thoughts](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#58e35dde604c) [What is next?](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#06cac91a384d) In this article [1\. RoBoSpatial: Training AIto Reason About Space](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#3948f67b7f25) [2\. GROVE: Teaching Robots with Generalized Rewards](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#ade977ae5dab) [3\. Navigation World Models: Planning in Simulation](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#b0eb4f95a305) [4\. Gemini Robotics: Bridging Foundation Models and Physical Action](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#38c7e1e5c358) [Why Embodied AI is the Next Big Boom](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#7b0b43cdb036) [A Call to the Research Community: Prepare to Validate the Future](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#e24e18b72b16) [Final Thoughts](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#58e35dde604c) [What is next?](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier#06cac91a384d) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _**CVPR 2025 Insights \#4:** Embodied Intelligence_ The E **mbodied Computer Vision** session at CVPR 2025 illuminated a transformative shift in AI, from passive perception to **intelligent, context-aware action**. As someone involved in applying robotics to agriculture, manufacturing, and elderly action recognition, I found this session to resonate profoundly with my journey. It reaffirmed what many of us working in the real world have long known: the future of AI is about moving, reasoning, and adapting alongside us. **Dr. Carolina Parada’s keynote** from Google DeepMind anchored this vision, highlighting how **embodied AI** represents the next great leap in artificial intelligence. These systems don’t just interpret the world; they interact with it, learn from it, and evolve within it. Some talks I followed at the beginning of the conference brought this idea to life with compelling, real-world progress: - **RoBoSpatial** introduced a benchmark for spatial reasoning, an essential yet often overlooked capability in robotics. It demonstrated how current vision-language models fall short when answering spatial questions across multiple frames, emphasizing the need for more grounded reasoning in dynamic environments. - **GROVE** addressed the challenge of reward design in reinforcement learning by using vision-language prompts and high-level goals. This approach allows robots to learn diverse behaviors without handcrafted engineering, making them more adaptable and scalable. - **Navigation World Models** presented a diffusion-based model that enables agents to simulate outcomes, predict consequences, and plan trajectories before acting, bringing human-like foresight to autonomous systems. Together, these contributions paint a clear picture: **embodied intelligence** is an active transformation across agriculture, manufacturing, healthcare, and beyond. This blog explores those meaningful insights and concludes with a reflection on why **Embodied Computer Vision** is not just the next big thing in AI; it’s the bridge between perception and purposeful action. As a community, it’s time to prepare for what’s next. ## 1\. RoBoSpatial: Training AIto Reason About Space The [RoBoSpatial](https://openaccess.thecvf.com/content/CVPR2025/papers/Song_RoboSpatial_Teaching_Spatial_Understanding_to_2D_and_3D_Vision-Language_Models_CVPR_2025_paper.pdf) benchmark offers a novel dataset and evaluation framework for spatial reasoning in robotics. It enables models to answer grounded questions like: - “Can the chair fit in front of the cabinet?” - “What’s behind the object in frame X?” - “Is this space compatible with object Y?” These spatial questions are generated heuristically with 3D scene grounding across multiple reference frames, leading to accurate, data-efficient annotations. When tested, even state-of-the-art vision-language models struggled. However, models trained on RoBoSpatial showed significantly improved spatial understanding. Spatial reasoning is essential for embodied AI. RoBoSpatial provides a foundational tool to benchmark and improve it. ![](https://cdn.sanity.io/images/h6toihm1/production/2231e6151e7652d75377f6f8416c48d17445c255-1400x1115.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ## 2\. GROVE: Teaching Robots with Generalized Rewards [GROVE](https://jiemingcui.github.io/grove/) (Generalized Reward Learning) bypasses the tedious need for hand-crafted reward functions in reinforcement learning. By leveraging vision-language prompts and diffusion planning, GROVE can: - Translate open-ended instructions (e.g.,“box with both hands”) into rewards. - Train robots across varied embodiments (humanoids, quadrupeds). - Adapt across tasks like motion imitation and locomotion. This approach means robots can learn complex behaviors (like agile movement or interaction) with minimal supervision, making them more adaptable, scalable, and intelligent. GROVE proves that embodied intelligence doesn’t need bespoke engineering; it can be taught through language, context, and high-level goals. ![](https://cdn.sanity.io/images/h6toihm1/production/ed5d094d49877eaeecc6bfddf6ea2a15451251e4-1400x784.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ## 3\. Navigation World Models: Planning in Simulation [This talk](https://openaccess.thecvf.com/content/CVPR2025/papers/Bar_Navigation_World_Models_CVPR_2025_paper.pdf) introduced a Conditional Diffusion Transformer (CDT) trained on 700+ hours of multimodal robot data. It learns to: - Predict future visual frames given the current context and action. - Simulate outcomes across diverse environments. - Evaluate and plan trajectories before acting in the real world Such models empower agents with an “imagination”, they can simulate action consequences before execution, making decisions that align with long-term goals and environment dynamics. Prediction, not just perception, is at the heart of embodied AI. ![](https://cdn.sanity.io/images/h6toihm1/production/320f41321956006cf8933b50576d9b0293a33ab4-1400x1224.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ## 4\. Gemini Robotics: Bridging Foundation Models and Physical Action Dr. Carolina Parada’s keynote was a watershed moment. As the Director of Robotics at Google DeepMind, she unveiled how Gemini Robotics brings Google’s flagship multimodal model into the physical world. > “Gemini Robotics draws from Gemini’s world understanding and brings it to the physical world by adding actions as a new modality.” Highlights from the keynote include: - **Visual-Language-Action Models (VLA):** Robots interpret commands like “slam dunk the basketball” and execute nuanced physical actions, even with previously unseen objects. - **Generalization Across Embodiments:** The same VLA model can control different robots, from arms like Aloha and Franca to full humanoids like Apollo. - **Zero-shot dexterity:** With a few demonstrations, robots folded origami, fit timing belts, and performed delicate tasks, proving that fine motor skills can be learned from visual feedback alone. - **Safety-first design:** With the ASIMOV Benchmark and proactive safety monitors, Gemini ensures robots operate ethically, intelligently, and securely. Gemini Robotics marks a leap in embodied intelligence, combining world knowledge, multimodal reasoning, and real-time adaptation. ![](https://cdn.sanity.io/images/h6toihm1/production/ace8d62f1d9e8a5b0e845aad2bb38486e90909e1-1400x784.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ## Why Embodied AI is the Next Big Boom In just a few years, we’ve seen vision models evolve into multimodal agents that understand, plan, and act. With models like Gemini, GROVE, and RoBoSpatial, the boundary between perception and action is vanishing. We’re entering a world where: - Robots interpret ambiguous language and respond in context. - Models reason about geometry, physics, and human intent. - Intelligence isn’t trapped in the cloud — it’s embodied, reactive, and adaptive. **Embodied AI is not hype, it’s here, and it’s the future. This is not another AI milestone; this is a paradigm shift.** ## A Call to the Research Community: Prepare to Validate the Future As Dr. Parada emphasized, “We’re riding the wave of foundation models — but for robotics, we still need breakthroughs.” To get there, we must: - Build stronger VLMs that understand the physical world. - Validate generalization with diverse and multimodal data. - Benchmark safety, social understanding, and fine motor control (dexterity). - Create shared community tools like [**ASIMOV**](https://asimov-benchmark.github.io/) and **RoBoSpatial**. - Shift from 2D VQA to **embodied evaluation**. > “In order to build truly helpful robots, we must ground our research in the physical world and validate our models through embodied interaction.” — Carolina Parada ## Final Thoughts CVPR 2025 clarified that the fusion of visual understanding, language, and physical action is no longer science fiction. Embodied AI is rising, and it demands **new data, new benchmarks, new ethics, and new imagination**. Let’s shape that future together. ## What is next? If you’re interested in following along as I dive deeper into the world of AI and continue to grow professionally, feel free to connect or follow me on [LinkedIn](https://www.linkedin.com/in/paula-ramos-phd/). Let’s inspire each other to embrace change and reach new heights! You can find me at some Voxel51 events ( [https://voxel51.com/computer-vision-events/](https://voxel51.com/events)), or if you want to join this fantastic team, it’s worth taking a look at this page: [https://voxel51.com/jobs/](https://voxel51.com/careers) [CVPR](https://voxel51.com/blog/tag/cvpr) [multimodal](https://voxel51.com/blog/tag/multimodal) [human perception](https://voxel51.com/blog/tag/human-perception) [visual agents](https://voxel51.com/blog/tag/visual-agents) ![](https://cdn.sanity.io/images/h6toihm1/production/e926c07c7d1426c0fde8fdefa637c528d47b16f4-512x512.webp?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a558b86370f2f17212fb2f2c894d590101458a85-5760x3241.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Multimodal Frontier in Computer Vision, Medicine, and Agriculture— CVPR 2025 Reflections\\ \\ Event Recaps, Industry Solutions\\ \\ • \\ \\ Jun 24, 2025](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections) [![](https://cdn.sanity.io/images/h6toihm1/production/8bf3c50a5edd9f83e1013ad5df86ae159d519dae-1200x626.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Best of CVPR 2025: Conversations at the Cutting Edge of AI\\ \\ Event Recaps\\ \\ • \\ \\ Jul 3, 2025](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/7ec61c80f387b16f24b4b2fe33f804264a38b486-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Rethinking How We Evaluate Multimodal AI\\ \\ Event Recaps\\ \\ • \\ \\ Jun 12, 2025](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-51-lllmstxt|> ## Visual AI in Agriculture [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Visual AI in Agriculture From sprawling fields to tiny seedlings, agriculture data piles up fast. FiftyOne helps you find exactly the samples you need for model training, smart annotation, model evaluation, and more. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/3df3dc138314e30e0a0ab1bd8d4406f5487a659b-628x394.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,628,394&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/5d8783b20c0d35f9530dc9cf11e1afe854a919ca-628x816.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/4f5d908e09a611efe9999483dc93b66fad6d4120-628x812.png?auto=format&dpr=2&fit=max&q=75&rect=0,85,628,727&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/79ad86bcc5b7695ba0085b514f336c295cac10f9-629x393.png?auto=format&dpr=2&fit=max&q=75&w=315) Benefits ## FiftyOne helps visual AI projects in agriculture thrive Leaders and innovators in agriculture rely on FiftyOne to accelerate the development of visual AI applications, including crop and livestock health detection and monitoring, automated farm equipment, precision agriculture, grading and sorting, and more. 0% increase in model accuracy 0+ months of development time saved 0% boost in team productivity features ## Designed for AI builders FiftyOne natively supports and enables the computer vision building blocks needed to develop robust agriculture AI solutions. ![](https://cdn.sanity.io/images/h6toihm1/production/0870ccefc24a1487f124ffc9bf87701c5cf6abbe-2560x961.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,2560,961&w=1280) - Classification - Detection - Segmentation - Polygons and polylines - Keypoints - Pointclouds - Heatmaps - Geolocation - Embeddings - Multiview datasets - Images, videos, and 3D data > “We’ve been using FiftyOne for over a year and it has drastically changed the way we work. The ability to easily display and analyze our images and their metadata, including experiment results, has been a refreshing change compared to the way we’ve worked before – mainly writing our own metrics and viewers. I’ve personally used FiftyOne for a segmentation model I’ve trained – trying to analyze the results and visually see what my model outputs has been really easy and fluid thanks to FiftyOne.” > > **Ido Greenfeld** > > AI Team Lead, Taranis ![](https://cdn.sanity.io/images/h6toihm1/production/c3e918765cd30f58112e6f3167f9cee2eaa99498-717x348.png?auto=format&dpr=2&fit=max&q=75&w=100) > “FiftyOne has been super effective in helping me to quickly understand the issues in my datasets, models, and latent representations.” > > **Dan Erez** > > Computer Vision Expert, Taranis ![](https://cdn.sanity.io/images/h6toihm1/production/c3e918765cd30f58112e6f3167f9cee2eaa99498-717x348.png?auto=format&dpr=2&fit=max&q=75&w=100) Resources ## Discover how visual AI is transforming the future of farming. Learn about the latest AI-driven technologies that are enhancing crop monitoring, improving yield predictions, and making farming more efficient and sustainable. ### Why Computer Vision in Agriculture is the Future Computer vision and artificial intelligence are driving innovation in agriculture. Learn more. [Read more](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023) ![](https://cdn.sanity.io/images/h6toihm1/production/225cdfec3976c8f7b4b1b1c661ccb711994651a3-1024x576.png?auto=format&dpr=2&fit=max&q=75&w=512) ## Data eats models for lunch Talk to our computer vision experts to start building better datasets and models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-52-lllmstxt|> ## CVPR 2025 Highlights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps), [Product & News](https://voxel51.com/blog/category/product-news) Voxel51 @CVPR 2025: Smarter, Faster Visual AI Jun 3, 2025 • 6 min read Article content In this article [Demos](https://voxel51.com/blog/cvpr-2025#22f044d7c2b7) [Workshops & Tutorials](https://voxel51.com/blog/cvpr-2025#3f91a568a49d) [Lightning Talks](https://voxel51.com/blog/cvpr-2025#fea01986d416) [Simulation-to-Reality with NVIDIA Omniverse + FiftyOne](https://voxel51.com/blog/cvpr-2025#21f389c94e1f) [Zero-shot auto-labeling with near-human labeling performance at a fraction of the cost](https://voxel51.com/blog/cvpr-2025#ffd225ca3032) [Solving the Video Understanding Challenge with Voxel51, Twelve Labs, and Databricks](https://voxel51.com/blog/cvpr-2025#5188482fcf09) [Voxel51 Workshops and Talks](https://voxel51.com/blog/cvpr-2025#a6410a5abac6) [Interesting CVPR Research Papers](https://voxel51.com/blog/cvpr-2025#9359003d104b) [Community Happy Hour](https://voxel51.com/blog/cvpr-2025#8f024132511a) In this article [Demos](https://voxel51.com/blog/cvpr-2025#22f044d7c2b7) [Workshops & Tutorials](https://voxel51.com/blog/cvpr-2025#3f91a568a49d) [Lightning Talks](https://voxel51.com/blog/cvpr-2025#fea01986d416) [Simulation-to-Reality with NVIDIA Omniverse + FiftyOne](https://voxel51.com/blog/cvpr-2025#21f389c94e1f) [Zero-shot auto-labeling with near-human labeling performance at a fraction of the cost](https://voxel51.com/blog/cvpr-2025#ffd225ca3032) [Solving the Video Understanding Challenge with Voxel51, Twelve Labs, and Databricks](https://voxel51.com/blog/cvpr-2025#5188482fcf09) [Voxel51 Workshops and Talks](https://voxel51.com/blog/cvpr-2025#a6410a5abac6) [Interesting CVPR Research Papers](https://voxel51.com/blog/cvpr-2025#9359003d104b) [Community Happy Hour](https://voxel51.com/blog/cvpr-2025#8f024132511a) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Visual and multimodal AI applications are quickly evolving from research experiments into core drivers of real-world innovation. Powering everything from autonomous vehicles, robots, to industrial automation and analytics, the success of these systems in production relies on high-quality visual data and scalable workflows. At [CVPR 2025](https://cvpr.thecvf.com/), the premier conference in computer vision, we’re excited to showcase practical tools, integrations, and workflows that meet the rising demands for building faster, more accurate AI models and datasets. Join us in Nashville, TN, from June 17-21 to get hands-on with the tools engineered to simplify your data processes and accelerate model development. Here's a glance at what we’ll be showcasing at the conference. ## **Demos** - [Simulation-to-reality](https://voxel51.com/blog/cvpr-2025#21f389c94e1f): Smart simulation, curation, visualization with [NVIDIA Omniverse](https://www.nvidia.com/en-us/omniverse/) and [Voxel51](https://voxel51.com/) - [Verified Auto-Labeling:](https://voxel51.com/blog/cvpr-2025#ffd225ca3032) New zero-shot annotation workflows in FiftyOne to label visual data 5,000X faster with near-human accuracy - [Video Search](https://voxel51.com/blog/cvpr-2025#5188482fcf09):Understand and efficiently search video content with [Databricks](https://www.databricks.com/), [TwelveLabs](https://www.twelvelabs.io/), and [Voxel51](https://voxel51.com/) - Explore and curate datasets, improve and debug AI models using FiftyOne structured and actionable insights. ## **Workshops & Tutorials** - [VAND 3.0](https://voxel51.com/events/vand-3-0-cvpr-2025-workshop): Visual Anomaly and Novelty Detection workshop - Agriculture-Vision: Visual AI in the Field - [Tutorial](https://www.agriculture-vision.com/agriculture-vision-2025/tutorial-2025) \+ [Workshop](https://www.agriculture-vision.com/) ## **Lightning Talks** - Practical tips and discussion focused on solving data curation and model performance challenges in use cases across medical, agtech, and manufacturing. Come for the talks and demos, and stay for the drinks, swag, and insightful conversations! 🎉 ## **Simulation-to-Reality with NVIDIA Omniverse + FiftyOne** Synthetic data is becoming a critical component in building robust visual AI systems, especially in domains like AVs, industrial robotics, and physical automation, where collecting diverse real-world data is often costly, impractical, or even unsafe. It’s slowly becoming a powerful alternative in scenarios where precise control over scene composition, object placement, and conditions is necessary. By leveraging synthetic datasets using [NVIDIA Omniverse](https://www.nvidia.com/en-us/omniverse/), OpenUSD, and [FiftyOne data curation](https://voxel51.com/curation),ML engineers can achieve comprehensive coverage and significantly augment training sets, without the burden of extensive manual data collection or labeling. Come see a live demo of **FiftyOne and NVIDIA Omniverse t** o learn how you can bridge the gap between synthetic and real-world data curation: - Visualize and inspect synthetic scenes at scale - Organize and curate multimodal datasets for training and testing - Analyze 3D reconstructions of real scenes using [NVIDIA’s Neural Reconstruction Engine](https://www.nvidia.com/en-us/on-demand/session/gtcfall22-a4d9010/) - Augment your data with [NVIDIA Cosmos](https://www.nvidia.com/en-us/ai/cosmos/) world foundation models to imagine your samples like never before 📍 **See it in action** at the CVPR **booth #1417** with **Voxel51** and **NVIDIA**, and meet with the experts from both companies. ## **Zero-shot auto-labeling with near-human labeling performance at a fraction of the cost** Manual labeling has traditionally been one of the biggest bottlenecks in getting computer vision models into production due to the time, cost, and accuracy implications. As foundation models advance, zero-shot labeling is emerging as a viable, scalable path to building high-quality datasets—without the manual burden. We’re excited to introduce a new approach to [AI-assisted annotation](https://voxel51.com/annotation) that combines Voxel51’s expertise in data curation with automated labeling and QA workflows. Our [research on Verified Auto-Labeling](https://voxel51.com/reports/auto-labeling-data-for-object-detection) shows that this approach achieves 95% of human-level performance while being **5,000X faster** and **cutting costs** up to **$100,000X.** To put it in perspective, a dataset such as [COCO](https://cocodataset.org/#home), consisting of ~850K objects, takes **human labeling** a total of **1,653** **hours** and **$30,598 in costs, compared to ~27 minutes and just $0.42 with Verified Auto-Labeling.** And with almost the same model performance! That’s mind-blowing and changes the paradigm of computer vision data workflows! Curious to see how it works? 📍Stop by the booth #1417 to get hands-on with the product. ## **Solving the Video Understanding Challenge with Voxel51, Twelve Labs, and Databricks** Video analysis techniques have historically involved a labor-intensive pipeline that includes manual annotation of frames or the use of scripts scrubbing through hours of footage using timestamps or scene changes. These manual approaches are expensive and error-prone, making it extremely challenging to detect precise scenarios, e.g., “a red car stopping at a traffic light” across an unlabeled video collection. State-of-the-art embedding techniques encode video content into searchable representations so you can perform similarity searches without the need for explicit video annotations. Learn about an integrated approach to understanding your video content using rich embeddings, fast similarity searches, and a visual user interface that streamlines the exploration of relevant video segments. In this demo with [TwelveLabs](https://www.twelvelabs.io/), [Databricks](https://www.databricks.com/), and [Voxel51](https://voxel51.com/), you can simplify video understanding and curation by learning how to: - Generate rich, multimodal embeddings (video+audio+text) using Twelve Labs’ foundation models - Index and search those embeddings for fast, cloud-scale retrieval with Databricks Vector Search - Run searches, visualize results, and refine datasets with FiftyOne 📍 **See this workflow in action** at the Voxel51 booth ## **Voxel51 Workshops and Talks** Join Voxel51 researchers to explore high-impact computer vision challenges ranging from anomaly detection to real-world deployment in agriculture. ### **VAND 3.0: Visual Anomaly and Novelty Detection** Researchers from academia and industries such as Bosch will present and discuss recent developments, opportunities and open challenges in the area of anomaly detection. Encourage the development and benchmarking of new algorithms by participating in the anomaly detection challenge! Organized by top researchers from AWS AI Labs, Voxel51, Intel, Durham University, and many others. **When**: June 18, 2025 **Where**: In person and on Zoom [Workshop Info](https://voxel51.com/events/vand-3-0-cvpr-2025-workshop) ### **Agriculture-Vision: Visual AI in the Field** Learn how researchers and practitioners are applying computer vision to real-world agricultural challenges—from crop segmentation to weed detection and disease prediction. **Tutorial**: June 12, 2025 @8:30 am \| [Tutorial Info](https://www.agriculture-vision.com/agriculture-vision-2025/tutorial-2025) **Workshop**: June 12, 2025 @1pm \| [Workshop Info](https://www.agriculture-vision.com/) Our partners at the NVIDIA research team have over [60 papers and 15+ workshops](https://www.nvidia.com/en-us/events/cvpr/). Check out the workshop on “ [Exploring the Next Generation of Data”](https://sites.google.com/view/nexd25/home) where the NVIDIA team addresses the challenges of curating NeXD25, high-quality, scalable, and unbiased data for foundation models in safety-critical applications. ## **Interesting CVPR Research Papers** Here are some interesting CVPR papers we’re watching that we think you should too! - [OpticalNet: An Optical Imaging Dataset and Benchmark Beyond the Diffraction Limit](https://deep-see.github.io/OpticalNet/assets/paper.pdf) B. Wang, R. An, J.-K. So, S. Kurdiumov, E. A. Chan, G. Adamo, Y. Peng, Y. Li, and B. An - [“Few-Shot Adaptation of Grounding DINO for Agricultural Domain,”](https://arxiv.org/abs/2504.07252v1) R. Singh, R. B. Puhl, K. Dhakal, and S. Sornapudi - “ [FLAIR: VLM with Fine-grained Language-informed Image Representations,”](https://arxiv.org/abs/2412.03561) R. Xiao, S. Kim, M.-I. Georgescu, Z. Akata, and S. Alaniz, - [“RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings,”](https://arxiv.org/pdf/2502.19781) A. Dhakal, S. Sastry, S. Khanal, A. Ahmad, E. Xing, and N. Jacobs - [“Interactive Medical Image Analysis with Concept-based Similarity Reasoning”](https://arxiv.org/abs/2503.06873) Ta Duc Huy, Sen Kim Tran, Phan Nguyen, Nguyen Hoang Tran, Tran Bao Sam, Anton van den Hengel, Zhibin Liao, Johan W. Verjans, Minh-Son To, and Vu Minh Hieu Phan ### **Best of CVPR 2025 Series: 12 must-read papers** Read more in **The Best of CVPR 2025 Series** three-part series ( [Part 1](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1), [Part 2](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2), [Part 3](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3)) that highlights papers that rethink how vision models interact with complexity, ambiguity, including papers that address safety, trust, usability across use cases and industries. ### **Visual agents: Key research wave at CVPR** CVPR 2025 marks a turning point for visual agents: this year's papers are signaling a shift from academic curiosity to tangible systems capable of interacting and controlling visual environments. [Read the full blog](https://cbbk204.na1.hubspotlinks.com/Ctc/W0+113/cBBk204/VWpB3t4R-3VbW7_hBBl907D09W5g6HvJ5xjS2XN61zzVd3qn9gW7lCdLW6lZ3n1W6q-KxM5XBzcsW5XdBvX5BrhC5F89jqHWXVz3W3GsR004fKbCbW37JGgy5X7BbxW1hFqms1mdt1-W7YXm2W11pPBVW8l0MCZ3_5c1fW1wMTN17WXQ5pN6MrMBxJ85PJW8CJP6_4T7sN_VwNH938YMj88W2Nn_VL2VnpglW8DcwyS2pv6pJW2b_cL63Vy6WsW5nbp6_5qz4HSW8PGVRN74sbcgN95tZ99p3RBXN7-dXl7VZJlgW1s4lgt1Zv0JjW5T_r4Z8NRXN_W8MxYmM6zVKcnW3pgNSg7rSBZ7W6-xrwv2c77dtf4LDxrT04) for key research papers driving the development. ### **Compositional Image Retrieval: The future of visual search** Compositional Image Retrieval (CIR) brings image search closer to how humans naturally describe visual concepts, opening new doors for e-commerce and creative applications. [Check out the blog](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025) for an in-depth look. ## **Community Happy Hour** Presenting a paper at CVPR and using FiftyOne in the course of your research? Meet the FiftyOne ML experts and have a drink. Spots are limited. [Register here](https://share.hsforms.com/1_CP3dm4SQxykBJQibNendQ2ykyk)! 🎸 We can’t wait to meet you! [CVPR](https://voxel51.com/blog/tag/cvpr) [synthetic data](https://voxel51.com/blog/tag/synthetic-data) [embeddings](https://voxel51.com/blog/tag/embeddings) [robotics](https://voxel51.com/blog/tag/robotics) [auto-labeling](https://voxel51.com/blog/tag/auto-labeling) ![](https://cdn.sanity.io/images/h6toihm1/production/217576d5aa661fd7bb82f46798a5f1b9902d637d-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Kirti Joshi Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/990a31a85d850d41fb482ae29960a8fa10b18ecd-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Behind the math: how we built the annotation savings calculator\\ \\ Product & News\\ \\ • \\ \\ Jun 4, 2025](https://voxel51.com/blog/how-we-built-annotation-savings-estimator) [![](https://cdn.sanity.io/images/h6toihm1/production/8bf3c50a5edd9f83e1013ad5df86ae159d519dae-1200x626.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Best of CVPR 2025: Conversations at the Cutting Edge of AI\\ \\ Event Recaps\\ \\ • \\ \\ Jul 3, 2025](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/a558b86370f2f17212fb2f2c894d590101458a85-5760x3241.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Multimodal Frontier in Computer Vision, Medicine, and Agriculture— CVPR 2025 Reflections\\ \\ Event Recaps, Industry Solutions\\ \\ • \\ \\ Jun 24, 2025](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-53-lllmstxt|> ## Hidden Costs of Outsourcing [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Your data, your advantage: the hidden cost of outsourced data annotation Jul 14, 2025 • 8 min read Article content In this article [Outsourcing your data annotation is handing competitors the keys](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#3e069e08c923) [Proprietary data creates unbreachable competitive moats](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#721a1a50d200) [Foundation models can now label your data nearly as well as humans](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#85414343507f) [Five questions every executive should ask before outsourcing data annotation](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#42058e56d145) [A strategic framework for implementing Data-Centric AI](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#c47766b02de0) [Your data, your way: the future belongs to companies that control their data destiny](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#75ae0a372828) In this article [Outsourcing your data annotation is handing competitors the keys](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#3e069e08c923) [Proprietary data creates unbreachable competitive moats](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#721a1a50d200) [Foundation models can now label your data nearly as well as humans](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#85414343507f) [Five questions every executive should ask before outsourcing data annotation](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#42058e56d145) [A strategic framework for implementing Data-Centric AI](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#c47766b02de0) [Your data, your way: the future belongs to companies that control their data destiny](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation#75ae0a372828) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) The recent shakeup in the AI industry offers a critical lesson for executives: your competitive advantage in artificial intelligence isn't just about algorithms—it's about who has access to your data and how you use it. When [Meta acquired a 49% stake in Scale AI](https://scale.com/blog/scale-ai-announces-next-phase-of-company-evolution), one of the largest data labeling companies, the ripple effects were immediate and telling. Google canceled a $200 million contract overnight. [Google](https://www.reuters.com/business/google-scale-ais-largest-customer-plans-split-after-meta-deal-sources-say-2025-06-13/), [xAI](https://www.reuters.com/business/google-scale-ais-largest-customer-plans-split-after-meta-deal-sources-say-2025-06-13/), and [OpenAI](https://www.bloomberg.com/news/articles/2025-06-18/openai-is-phasing-out-its-work-with-scale-ai-after-meta-deal) retreated from partnerships. Demand for Scale's competitors tripled within weeks. The message was unmistakable: companies suddenly realized they had handed over their most valuable asset—proprietary training data—to a potential competitor. This episode illuminates a broader strategic question that every AI-driven organization must answer: Should you outsource your data annotation, or bring this critical capability in-house? ## **Outsourcing your data annotation is handing competitors the keys** Most executives view data labeling as a necessary but mundane operational task—like facilities management or payroll processing. This perspective misses a fundamental truth: in AI development, your training data is your competitive moat. When you outsource its preparation, you're not just buying convenience; you're potentially surrendering strategic advantage. Consider what happens when you send proprietary datasets to external vendors. Your data contains unique insights about your operations, customer behaviors, and market dynamics. The very act of selecting and labeling this data reveals your business logic and strategic priorities. External annotators gain intimate knowledge of your approach to problems, which could inadvertently benefit your competitors if that vendor serves multiple clients in your industry. The risks extend beyond competitive intelligence. Data breaches at third-party vendors can expose sensitive information. Regulatory compliance becomes more complex when data crosses organizational boundaries, particularly in highly regulated industries like healthcare and financial services. And perhaps most critically, you lose control over a process that directly impacts your AI system's performance and reliability. ![](https://cdn.sanity.io/images/h6toihm1/production/53ba6e6ac82668a30fd1a0273a83de957f4602a2-1600x1489.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## **Proprietary data creates unbreachable competitive moats** The companies achieving sustainable competitive advantages through AI share a common trait: they zealously guard their proprietary data and the insights it generates. ![RETFound is a foundation model trained on over 1.6 million NHS retinal scans that can diagnose eye diseases and predict systemic conditions like stroke and Parkinson’s. Its self-supervised learning approach enables high performance with minimal labeled data and strong generalization across diverse populations.](https://cdn.sanity.io/images/h6toihm1/production/779f81bdb30fc86667e2b6651ee03cbda7f980db-1600x809.png?auto=format&dpr=2&fit=max&q=75&w=1600)RETFound is a foundation model trained on over 1.6 million NHS retinal scans that can diagnose eye diseases and predict systemic conditions like stroke and Parkinson’s. Its self-supervised learning approach enables high performance with minimal labeled data and strong generalization across diverse populations. Moorfields Eye Hospital and the UCL Institute of Ophthalmology received exclusive access to 1.6 million retinal scans from Moorfields Eye Hospital, the world's largest ophthalmic imaging dataset. Their [RETFound model](https://www.nature.com/articles/s41586-023-06555-x) doesn't just outperform competitors in detecting eye diseases; it can identify general health conditions across diverse patient populations—a breakthrough made possible by their unique data advantage. [Amazon's StyleSnap visual search](https://www.amazon.science/latest-news/the-science-behind-amazons-new-stylesnap-for-home-feature#:~:text=overcome%20some%20of%20the%20same,snapshots%20that%20customers%20might%20take) feature demonstrates how proprietary data creates customer value that's difficult to copy. Trained on Amazon's extensive product catalog, it matches customer-uploaded images to similar products with remarkable accuracy. Competitors can't replicate this capability because they don't have access to Amazon's comprehensive product image database. John Deere subsidiary [Blue River Technology](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#d738fb9a63c9) built its precision agriculture business on a massive, proprietary dataset of plant images—hundreds of thousands of photos of crops and weeds. This data enables their farming robots to identify and target weeds with pinpoint accuracy, a capability competitors struggle to replicate because they lack equivalent training data. These examples illustrate a fundamental principle: while AI algorithms and model architectures are increasingly commoditized through open-source availability, proprietary data remains uniquely yours. The question becomes how to maximize this advantage while maintaining operational efficiency. > “I do believe the models are getting commoditized…models by themselves are not sufficient, but having a full system stack and great successful products, those are the two places” where companies need to focus now. > > — Satya Nadella, Microsoft CEO ## **Foundation models can now label your data nearly as well as humans** Until recently, organizations faced a binary choice: either accept the risks of outsourcing data annotation or commit to expensive workforce investments in-house. Advances in foundation models have created a third option that changes the strategic calculus entirely. Modern AI systems can now label data for AI development. Pre-trained vision models can automatically identify and segment objects in images, while language models can classify and tag text data. Human experts then verify and refine these automated labels, focusing their expertise where it adds the most value. This approach delivers remarkable efficiency gains. Where manual labeling might require thousands of hours, AI-assisted workflows can reduce the task to hundreds of hours of human review. The cost differential is equally dramatic. In a paper titled, [“Auto-Labeling for Object Detection,”](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance) machine learning researchers found that automated labeling can reduce annotation costs by orders of magnitude while maintaining accuracy standards. Labelling 3.4 million objects on a single NVIDIA L40S GPU costs $1.18 and took just over an hour. Manually labeling the same dataset via AWS SageMaker, which has among the least expensive annotation costs, would cost roughly $124,092 and take nearly 7,000 hours. ![Models trained with Voxel51’s Verified Auto-Labeling (VAL) approach achieve near-human accuracy on data annotation tasks—mAP50 scores of 0.768 on VOC and 0.538 on COCO—while reducing annotation costs by up to 100,000×, proving that high-quality training data no longer has to come at high cost.](https://cdn.sanity.io/images/h6toihm1/production/294c1e1d7b2bebad54dbc95c18de28a02818c126-901x517.png?auto=format&dpr=2&fit=max&q=75&w=901)Models trained with Voxel51’s Verified Auto-Labeling (VAL) approach achieve near-human accuracy on data annotation tasks—mAP50 scores of 0.768 on VOC and 0.538 on COCO—while reducing annotation costs by up to 100,000×, proving that high-quality training data no longer has to come at high cost. Critically, this entire process can occur within your secure environment. You can deploy open-source models on your infrastructure or use vendor tools designed for on-premises operation. Your data never leaves your control, eliminating the security and competitive risks of traditional outsourcing. ## **Five questions every executive should ask before outsourcing data annotation** For executives considering this shift, the decision framework should address several key questions. First, assess the sensitivity of your training data. Industries like healthcare, finance, and defense have obvious sensitivity concerns, but even consumer-focused companies may have competitive intelligence embedded in their datasets. Next, analyze your competitive landscape—are your competitors likely to use the same data labeling vendors? The more concentrated your industry's outsourcing relationships, the higher the risk of indirect data sharing. Consider the scale and resource requirements carefully. AI-assisted labeling requires upfront investment in technology and process development, but these costs typically scale sublinearly with data volume, creating long-term economic advantages. Evaluate your regulatory environment as well, since industries subject to strict data protection requirements may find in-house labeling simplifies compliance and reduces regulatory risk. Finally, factor in speed to market considerations. While initial setup takes time, auto-labeling ultimately provides greater agility in responding to changing requirements and market conditions. ## **A strategic framework for implementing Data-Centric AI** The transition toward data sovereignty represents part of a broader shift to data-centric AI practices—an approach that recognizes data quality and control as the primary drivers of competitive advantage. This philosophy extends beyond labeling to encompass the entire data lifecycle, from collection and curation to model training and validation. Organizations embracing data-centric AI begin by auditing their current data practices across all AI initiatives. This involves mapping data flows, identifying where proprietary information travels outside organizational boundaries, and assessing the strategic value of different datasets. The goal is understanding not just what data you have, but how effectively you're leveraging it as a competitive asset. The implementation typically starts with automated labeling for data annotation. Organizations should begin with less sensitive datasets to test auto-labeling solutions and validate downstream model performance on their specific use cases. Smart companies leverage data curation techniques to strategically slice their datasets, determining which portions can be efficiently auto-labeled and which specialized classes require human-in-the-loop workflows that require critical domain expertise. Beyond automated labeling, the real competitive advantage comes from how thoughtfully you organize and refine your data over time. Most companies treat data preparation as a one-time task, but the smartest organizations recognize that data curation is an ongoing strategic process that compounds in value. [Amazon's computer vision research](https://www.amazon.science/blog/how-computer-vision-will-help-amazon-customers-shop-online) demonstrates this principle in action. The company has developed sophisticated visual search systems that allow customers to refine product queries by describing variations on images—saying something like "I want it to have a light floral pattern" to modify search results. This capability emerged from years of carefully curating product images with detailed, nuanced labels that capture style attributes, textures, and aesthetic qualities that competitors using generic categorization schemes cannot match. Effective [data curation](https://voxel51.com/curation) requires establishing domain-specific labeling standards that reflect your business priorities, implementing systematic quality control to catch inconsistencies, and creating feedback mechanisms where AI performance informs labeling improvements. The strategic advantage emerges because well-curated data becomes increasingly difficult for competitors to replicate. While they can copy your algorithms, they cannot easily recreate years of thoughtful data organization that reflects your unique operational context and domain expertise. The ideal partner provides tools that seamlessly integrate into your existing model development workflows rather than forcing you to adopt entirely new systems. Look for platforms that offer APIs, SDKs, and plugin architectures that work with your current machine learning infrastructure. Extensibility becomes crucial as your AI capabilities mature, enabling you to scale from simple auto-labeling tasks to complex, multi-stage data pipelines without requiring complete infrastructure overhauls. Most importantly, seek partners who embrace collaborative development approaches where their AI capabilities enhance rather than replace your domain expertise. The goal isn't to hand over your data to a black box system, but to amplify your team's knowledge and insights. ## **Your data, your way: the future belongs to companies that control their data destiny** The Meta-Scale AI episode is more than an industry anecdote—it's a preview of how AI competition will intensify around data control. As AI becomes central to competitive advantage across industries, the organizations that maintain sovereignty over their data assets will be best positioned for long-term success. [Dive deeper with the whitepaper](https://voxel51.com/whitepapers/your-data-your-advantage) This shift requires executives to reconceptualize data labeling from a procurement decision to a strategic capability. Like research and development or product design, data preparation directly impacts your ability to compete and should be managed accordingly. The technology now exists to make this transition to model-led data annotation practical and economically viable. The question isn't whether you can afford to bring data labeling in-house—it's whether you can afford not to. In an AI-driven economy, your data is your most important competitive advantage. [auto-labeling](https://voxel51.com/blog/tag/auto-labeling) [data annotation](https://voxel51.com/blog/tag/data-annotation) ![](https://cdn.sanity.io/images/h6toihm1/production/8d61ff90b31d151405f9e21a33c2802509f34651-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Brian Moore Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/e60eea36edcff16d65c62e2c3fea99a66d36a8d5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 @CVPR 2025: Smarter, Faster Visual AI\\ \\ Event Recaps, Product & News\\ \\ • \\ \\ Jun 3, 2025](https://voxel51.com/blog/cvpr-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/990a31a85d850d41fb482ae29960a8fa10b18ecd-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Behind the math: how we built the annotation savings calculator\\ \\ Product & News\\ \\ • \\ \\ Jun 4, 2025](https://voxel51.com/blog/how-we-built-annotation-savings-estimator) [![](https://cdn.sanity.io/images/h6toihm1/production/7d027bddb314b23d6afedfbcbdd8e8770784661a-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ NVIDIA AI Podcast: ADAS, Visual AI, and the Road Ahead with Porsche and Voxel51\\ \\ Product & News\\ \\ • \\ \\ Jul 30, 2025](https://voxel51.com/blog/nvidia-ai-podcast-adas-with-porsche) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-54-lllmstxt|> ## Vision AI Model Failures [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Whitepaper # Why Vision AI Models Fail Why do vision AI models break down in the real world—even when they look perfect on paper? To find out, we analyzed the most common failure patterns across high-stakes domains like autonomous driving, retail, and healthcare. This guide outlines why most breakdowns are _data failures in disguise_—and what it really takes to build robust, production-ready models. Get practical takeaways and field-tested strategies, including: - **Data failure modes:** How label noise, imbalance, and bias quietly derail model accuracy - **Real-world incidents:** What Walmart, Tesla, and TSMC got wrong—and what it cost them - **Debugging the invisible:** Techniques for spotting silent failures standard metrics miss - **Visual QA workflows:** Tools to find mislabeled, biased, or low-quality samples—before deployment - **Fix-before-fail strategy:** How leading teams use data-centric practices to prevent outages and regressions ![](https://cdn.sanity.io/images/h6toihm1/production/60424cea21d1d2754c2e14e6e80f55bc6bac1447-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=1600) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-55-lllmstxt|> ## Aquabyte Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/1f17885f5ff39ab29f3754810547e6bd2e038f40-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Aquabyte Aquabyte optimizes fish farming operations with AI and FiftyOne Apr 14, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Aquabyte](https://aquabyte.ai/) is one of the few companies applying machine learning and computer vision to directly solve the world’s food sustainability issues. Fish farming is the #1 fastest growing sector of food production—a $300B worldwide industry. By improving fish farm efficiency, Aquabyte is helping close the world’s impending protein deficit. > "FiftyOne is a critical tool for supporting our computer vision work at Aquabyte. We love the speed and flexibility FiftyOne offers as a tool which helps support all sorts of workflows. It's just enough coding to be really powerful, but not so much that it turns me off from using it throughout my work. FiftyOne has supported my efforts in dataset EDA, data curation, model failure analysis, and more." – Samuel Weitzman, Lead Machine Learning Engineer at Aquabyte ![](https://cdn.sanity.io/images/h6toihm1/production/0ca5f2c1f2f730854eaf08fc4688a3c4ffe77e05-600x400.jpg?auto=format&dpr=2&fit=max&q=75&w=600) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-56-lllmstxt|> ## Robotics Visual AI Solutions [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Visual AI in Robotics Robots need to reliably and accurately sense and understand their environments. Achieving human-level perception requires both high-quality data and high-performing models. FiftyOne is here to help! [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/caee075c792d771749a6d73602100695748e0115-628x395.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/f4aafc5ae0f2922bc1f0cd3579991712cd37236f-628x817.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/d104b716562693daad284a775390aa000ceec950-628x813.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/4c4ba0371d60de4437352f7f8fe2682648a67022-628x393.png?auto=format&dpr=2&fit=max&q=75&w=314) Benefits ## Revolutionize robotics development with Voxel51 Delivering new advances in robotics solutions requires high-quality data and high-quality models. Leading AI builders use FiftyOne to enable them to refine and utilize visual data throughout development. ### Increase quality and accuracy Curate high-quality datasets and use them to evaluate and tune models to deliver the accuracy and reliability needed for production solutions. ### Accelerate time to production Streamline and automate the data pipelines that supply your development work, enabling faster iteration and getting to production-ready quality faster and more easily. ### Do more of what matters Save valuable engineering time spent on laborious data management and quality control so that engineers can focus on building more of what your customers love. Use cases ## FiftyOne helps visual AI projects execute with precision Leaders and innovators in robotics rely on
FiftyOne to boost data quality and model performance across computer vision use cases crucial to robotics solutions. ![](https://cdn.sanity.io/images/h6toihm1/production/53370b5acef3e28948292cd6b203dda362f4fb73-768x496.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Object detection and recognition Key tasks in robotics involve sensing and perceiving the environment to inform decision-making. Object detection and recognition techniques help robots perform their actions with precision. ![](https://cdn.sanity.io/images/h6toihm1/production/4e8f0fc822aa0f1dd2293b985290a8aa772bce55-710x430.png?auto=format&dpr=2&fit=max&q=75&w=355) Environmental monitoring and surveillance Enable robotics systems to actively monitor and analyze their environment to detect obstacles, disruptions, and changing conditions to trigger alerts or active responses. ![](https://cdn.sanity.io/images/h6toihm1/production/98cdacdaac46d4aece85a27b8e2c2efb72bfa359-710x430.png?auto=format&dpr=2&fit=max&q=75&w=355) Visual inspection and quality control Robots equipped with sophisticated visual systems can handle critical inspection tasks. These robots thrive even in harsh environments, accomplishing inspections swiftly and safely. ![](https://cdn.sanity.io/images/h6toihm1/production/6f88858bd7fd24d85eda628e9878ea8864862f2b-710x430.png?auto=format&dpr=2&fit=max&q=75&w=355) Autonomous navigation Computer vision is a crucial foundation for autonomous navigation systems, allowing robots to perceive, understand, and maneuver through their surroundings. ![](https://cdn.sanity.io/images/h6toihm1/production/9e734ce875c4c65209327bfbb0cf23efa3792e81-710x430.png?auto=format&dpr=2&fit=max&q=75&w=355) UAVs and UUVs Deliver solutions that navigate autonomously in complex environments and perceive their surroundings, identify obstacles, and maneuver safely through challenging settings. ![](https://cdn.sanity.io/images/h6toihm1/production/d0bdd82558f323ba7833f9c248dd0550eba702d2-710x430.png?auto=format&dpr=2&fit=max&q=75&w=355) Product assembly and manufacturing Bring high precision to automated manufacturing robots, as well as to the loading, preparation, and assembly of raw materials and components for processing. Features ## How visual AI can help you ### Understand and curate all your sensor data Utilize powerful visualizations to understand your data like never before. Load image, video, radar, and lidar samples into FiftyOne to curate your sensor data all in one place. ![](https://cdn.sanity.io/images/h6toihm1/production/930954101f5011b6ba6de1c8508b75098cb05607-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Quickly find what you’re looking for Sort, slice, and search your dataset any way you like — with natural language, an image, or filtering operations — to quickly find the samples you’re looking for. ![](https://cdn.sanity.io/images/h6toihm1/production/84dd890372baa4febf49a673b7359eb75cd08ba5-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Instantly improve data quality issues Easily spot image quality issues across your dataset. Find the pesky images that are hurting your performance, then apply transformations instantly using FiftyOne’s Albumentations integration. ![](https://cdn.sanity.io/images/h6toihm1/production/b3454b107f5751d615ed12e586003be1acdda282-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Easily compute and visualize embeddings Reveal hidden patterns and spot outliers in your samples through embeddings visualization. Choose from multiple ways to generate embeddings and specify the dimensionality reduction method to use. ![](https://cdn.sanity.io/images/h6toihm1/production/f782717051453ff6a582d6cc8c82fa48a6b82fcd-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) features ## Designed for AI builders FiftyOne natively supports and enables the computer vision building blocks needed to develop robust automotive AI solutions. ![](https://cdn.sanity.io/images/h6toihm1/production/b65da79f5913eb1e20235f7bc09363e94b9dca14-2560x960.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,2560,960&w=1280) - Classification - Detection - Segmentation - Polygons and polylines - Keypoints - Pointclouds - Heatmaps - Geolocation - Embeddings - Multiview datasets - Images, videos, and 3D data > “Everything we do from the MLOps side interacts with FiftyOne Teams. It’s becoming the hub where all the spokes are connected. It’s like GitHub for code, but for our datasets.” > > **Matt Shaffer** > > VP of AI & Co-founder, RIOS Intelligent Machines ![](https://cdn.sanity.io/images/h6toihm1/production/2e9b68952ebe661b86a7f609dad10532632bab5a-1817x920.png?auto=format&dpr=2&fit=max&q=75&w=100) > “We use FiftyOne Teams during model training to prep and select the data, then inspect the model’s performance on that data. On the other side, we take the inference results from our production models and test data to visualize and evaluate them in FiftyOne Teams. Doing this means we have an end-to-end process that closes the entire loop for us on model analysis.” > > **Shubham Kanitkar** > > Sr. Robotics ML Engineer, RIOS Intelligent Machines ![](https://cdn.sanity.io/images/h6toihm1/production/2e9b68952ebe661b86a7f609dad10532632bab5a-1817x920.png?auto=format&dpr=2&fit=max&q=75&w=100) Resources ## Learn more about visual AI in robotics Take a deeper dive into how computer vision is transforming robotics, and how FiftyOne can help your teams deliver robust visual AI solutions. ### RIOS’s AI-Powered Robotics Solutions Run on FiftyOne Enterprise RIOS selected FiftyOne Teams as its dataset management solution to efficiently organize and visualize 20TB+ of images, videos, and 3D data [Read more](https://voxel51.com/customers/rios) ![](https://cdn.sanity.io/images/h6toihm1/production/b225076935c6ada1d3969623e54240936c13b0dc-912x913.png?auto=format&dpr=2&fit=max&q=75&rect=0,214,912,499&w=456) ### How Computer Vision Is Transforming Robotics Explore how computer vision & AI are transforming the Robotics Industry, including use cases, technologies & companies at the cutting edge. [Read more](https://voxel51.com/blog/how-computer-vision-is-transforming-robotics) ![](https://cdn.sanity.io/images/h6toihm1/production/2accce7a17c676c06e75cc8647e211cfdac7690a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=960) ## Data eats models for lunch Talk to our computer vision experts to start building better datasets and models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-57-lllmstxt|> ## Annotation Savings Estimator [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Behind the math: how we built the annotation savings calculator Jun 4, 2025 • 5 min read Article content In this article [The inputs: what drives labeling cost?](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#2dfe41c2a257) [Benchmark data sources](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#5ebd5be5e640) [Human labeling estimator model](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#3205bacce4a6) [Verified Auto Labeling estimator model](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#c73a0537dbc7) [Putting it together: an example](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#8b159f609598) [The hidden costs of human annotation](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#a4f277df28d5) [Takeaways](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#9d8a46ccb05c) In this article [The inputs: what drives labeling cost?](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#2dfe41c2a257) [Benchmark data sources](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#5ebd5be5e640) [Human labeling estimator model](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#3205bacce4a6) [Verified Auto Labeling estimator model](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#c73a0537dbc7) [Putting it together: an example](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#8b159f609598) [The hidden costs of human annotation](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#a4f277df28d5) [Takeaways](https://voxel51.com/blog/how-we-built-annotation-savings-estimator#9d8a46ccb05c) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) When we launched **Verified Auto Labeling** we kept hearing the same two questions: 1. _“How much cheaper is auto-labeling than hiring humans for the same job?”_ 2. _“Where do those numbers actually come from?”_ To answer both at once, we published an [open-source spreadsheet](https://docs.google.com/spreadsheets/d/12CYDO-4oBrAf9JBiX4j98ncMnpb7lSWfklNTyVETTao/edit?usp=sharing) and a web-based [Annotation Savings Estimator.](https://voxel51.com/annotation#calculator) This post provides a walkthrough of how the estimator works: every user input, assumption, and formula that powers the calculator, so you can audit or adapt the model for your own projects. ## **The inputs: what drives labeling cost?** The Annotation Savings Estimator asks for three key inputs: - **Number of images to label (** to determine scale for costs and throughput) - **Task type** (Classification, Detection, or Segmentation) - **Task complexity** (Simple, Moderate, Complex, or Custom) based on the number of objects per image The task complexity is defined as the following: - 1–5 objects/image → Simple task Example: binary or small multi-class detection of basic objects, such as in [CIFAR-10](https://www.cs.toronto.edu/~kriz/cifar.html) dataset - 6–20 objects/image → Moderate task Example: multi-label objects and scenes, such as in [Cityscapes](https://paperswithcode.com/paper/cityscapes-3d-dataset-and-benchmark-for-9-dof) or [Pascal VOC](https://www.lvisdataset.org/dataset) datasets - 21-200+ objects/image → Complex task Example: multi-label objects with fine-grained objects in highly cluttered scenes or requiring specific domain knowledge, such as medical imaging, aerial views. Example datasets include [Open Images V4](https://storage.googleapis.com/openimages/web/index.html), [LVIS](https://www.lvisdataset.org/) The estimator also provides a “Custom” input field to enter the average number of objects per images if the user has a good idea of that for their dataset. The tool maps that to a complexity tier based on the object density for a more precise estimation. ## **Benchmark data sources** All references are enumerated [in-shee](https://docs.google.com/spreadsheets/d/12CYDO-4oBrAf9JBiX4j98ncMnpb7lSWfklNTyVETTao/edit?usp=sharing) t so you can trace every calculation. **Human labeling base cost and time per annotation** ![](https://cdn.sanity.io/images/h6toihm1/production/026bcfe1be5dd409624aaf27ef20e147de33969c-1024x185.png?auto=format&dpr=2&fit=max&q=75&w=1024) \*Source: Base time benchmark: [Research paper](https://openaccess.thecvf.com/content_iccv_2013/papers/Jain_Predicting_Sufficient_Annotation_2013_ICCV_paper.pdf) \*Source: Base price benchmark: [AWS Mechanical Turk ground truth pricing](https://aws.amazon.com/sagemaker-ai/groundtruth/pricing/) ### **Auto-labeling compute cost** The experiments conducted for the [Verified Auto Labeling research paper](https://voxel51.com/reports/auto-labeling-data-for-object-detection) use NVIDIA L40s GPU for compute. The cost for renting the GPUs is $0.93/hour. The foundation models used for benchmarking label generation are [YOLOE](https://docs.voxel51.com/integrations/ultralytics.html#open-vocabulary-segmentation), [YOLO-World](https://docs.voxel51.com/integrations/ultralytics.html#open-vocabulary-detection), and [Grounding DINO](https://huggingface.co/docs/transformers/en/model_doc/grounding-dino). ## **Human labeling estimator model** For each task type and complexity level, we use benchmarked human annotation time and standard labeling service costs to estimate: - Time per label (in seconds) - Total number of objects to annotate - Total human labeling hours - Total human labeling cost ### **For classification tasks** Human Labeling time (hrs) = base\_time x number of images to label / 3600 Human Labeling cost (USD) = base\_price x number of images to label ### **For detection tasks** Based on task complexity (simple, moderate, complex) we can use the lower and upper bounds for calculating the number of objects per image and then calculate the labeling time and cost range. If the user provides the average number of objects/image we use that in the equation. See the estimated obj/image for each tier of task complexity in the benchmark data sources section above. Total number of objects = number of objects per image x number of images to label Human labeling time (hrs) = total number of objects x base\_time / 3600 Human labeling cost (USD) = total number of objects x base\_cost ### **For segmentation tasks** The effort for instance and semantic segmentation not only depends on the number of images to label but also needs to factor in the scene complexity. This complexity depends vastly on how dense the scene is, intricate textures, whether the objects in the scene are well-defined, whether there is any background noise, or varying lighting conditions, etc. We derive heuristic estimates of the scene complexity factor based on image segmentation algorithm [research studies](https://arxiv.org/html/2504.04435v1) and overall empirical observations. ![](https://cdn.sanity.io/images/h6toihm1/production/c496aa7691c65f9cae8df003bf78e08348539429-1024x185.png?auto=format&dpr=2&fit=max&q=75&w=1024) If the user provides an estimated number of obj/img, we use that to tier the scene complexity. Human Labeling time (hrs) = base\_time x number of images to label x scene\_complexity / 3600 Human Labeling cost (USD) = base\_price x number of images to label x scene\_complexity ## **Verified Auto Labeling estimator model** We use the benchmark data for object detection from our [Verified Auto Labeling research paper](https://voxel51.com/reports/auto-labeling-data-for-object-detection). The paper covers auto-labeling benchmarks on datasets of varying classes, number of images, and complexities. We also use additional segmentation experiments conducted by the Voxel51 ML researchers to benchmark the auto-labeling time and costs. Here’s a summary of the numbers from the paper. ![](https://cdn.sanity.io/images/h6toihm1/production/7a7cd78fcfb352c74d92516240a07800e695039c-1024x284.png?auto=format&dpr=2&fit=max&q=75&w=1024) Using this as a baseline, we can now compute time and costs given user inputs. ### **For classification and detection tasks** We observed that for classification and detection tasks, the cost per image scales _almost_ linearly with dataset size. We fit a simple least-squares fit across the four test sets and summarized that into an equation to calculate cost and time for these tasks. Auto-Labeling cost (USD) = 3.8541 \*10^-6 x number of images to label + 0.0011187 ### **For segmentation tasks** For segmentation tasks, using the data above, we use a similar estimation and bucketize it for smaller and larger datasets with a threshold of 20,000 images. This is done so that costs for a smaller number of user images (< 20,000) can be accurately estimated. For small datasets <= 20,000 images Auto-Labeling cost (USD) = 0.309\*10^-5 x number of images to label For larger datasets > 20,000 images Auto-Labeling cost (USD) = 0.309\*10^-5 x number of images to label ## **Putting it together: an example** Let’s walk through an example so you can see the comparison of human labeling versus auto-labeling costs side by side. Say your task is to annotate a driving dataset by drawing bounding boxes around several city street objects consisting of classes such as cars, trucks, bicycles, stop signs, pedestrians, traffic signs, … **User inputs:** Number of images to label: 100,000 Task type: Detection Task Complexity: Custom Avg num of obj/img: 12 (moderate task complexity) Total number of objects = number of objects per image x number of images to label Total number of objects to label = 1,200,000 Human labeling cost = total number of objects x base\_cost = 1,200,000 x $0.036 Auto-labeling cost = 3.8541 \*10^-6 x number of images to label + 0.0011187 = $0.3865 **Human labeling cost = $43,200** **Verified Auto Labeling cost = $0.3865** **Savings factor = 111,781x** ## **The hidden costs of human annotation** Human annotation carries several hidden costs beyond the visible per-label pricing. These include onboarding annotators, designing detailed labeling guidelines, implementing quality assurance (QA) and rework cycles, managing communication and oversight, and building or maintaining annotation tools and infrastructure. Especially in complex tasks, these overheads can double or even triple the base cost of labeling. To keep it simple, our estimator does not include these hidden costs, but know that they are there and can add another factor to the true cost of human data annotation. ## **Takeaways** - Manual annotation becomes disproportionately expensive at scale and complexity. - Verified Auto Labeling provides substantial time and cost savings—up to 100,000x lower cost and 5,000x lower time, depending on the number of labels. Whether you’re labeling 1,000 images or 1 million, the estimator can help you get ballpark ROI numbers to justify investment in automation. [Try it out](https://voxel51.com/annotation) and see how much time and money you could save with Verified Auto Labeling. [annotation](https://voxel51.com/blog/tag/annotation) [auto-labeling](https://voxel51.com/blog/tag/auto-labeling) [verified-auto-labeling](https://voxel51.com/blog/tag/verified-auto-labeling) [annotation-savings](https://voxel51.com/blog/tag/annotation-savings) [cost-estimation](https://voxel51.com/blog/tag/cost-estimation) ![](https://cdn.sanity.io/images/h6toihm1/production/217576d5aa661fd7bb82f46798a5f1b9902d637d-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Kirti Joshi Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/e60eea36edcff16d65c62e2c3fea99a66d36a8d5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 @CVPR 2025: Smarter, Faster Visual AI\\ \\ Event Recaps, Product & News\\ \\ • \\ \\ Jun 3, 2025](https://voxel51.com/blog/cvpr-2025) [![Your data, your advantage - the hidden cost of outsourced data annotation](https://cdn.sanity.io/images/h6toihm1/production/9733b62da57bf722c6ab7ec2dee89ec9353158a4-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Your data, your advantage: the hidden cost of outsourced data annotation \\ \\ Product & News\\ \\ • \\ \\ Jul 14, 2025](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation) [![](https://cdn.sanity.io/images/h6toihm1/production/c845648a6e64e00a3cac3d875280cb462c63633d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Complete Guide to Auto Labeling\\ \\ Learn\\ \\ • \\ \\ Jun 16, 2025](https://voxel51.com/blog/the-complete-guide-to-auto-labeling) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-58-lllmstxt|> ## Transforming Computer Vision [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) How Image Embeddings Transform Computer Vision Capabilities Nov 25, 2024 • 10 min read Article content In this article [How Image Embeddings Transform Computer Vision Capabilities](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#19f09c4aa792) [What are image embeddings?](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#10e39217c740) [How to compute embeddings](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#dbd177e80a92) [Three powerful embeddings use cases in ML workflows](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#f22a1e2d3acc) [The future of image embeddings](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#621bcb185dd9) [How to use embedding capabilities in FiftyOne](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#495bd331df12) [Conclusion](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#f34f80e9c32d) [Next steps](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#d8724b9b1d44) In this article [How Image Embeddings Transform Computer Vision Capabilities](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#19f09c4aa792) [What are image embeddings?](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#10e39217c740) [How to compute embeddings](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#dbd177e80a92) [Three powerful embeddings use cases in ML workflows](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#f22a1e2d3acc) [The future of image embeddings](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#621bcb185dd9) [How to use embedding capabilities in FiftyOne](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#495bd331df12) [Conclusion](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#f34f80e9c32d) [Next steps](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities#d8724b9b1d44) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) # How Image Embeddings Transform Computer Vision Capabilities Advances in Computer Vision (CV) have significantly transformed the way machines learn and infer complex visual data with human-like accuracy. Tasks such as image classification, object detection, and video analysis drive multiple applications ranging from detecting abnormalities in medical scans to powering self-driving car features. Innovative techniques such as embeddings help machine learning engineers simplify their data analysis by being able to generate image embeddings that extract relevant and meaningful information from their data. Several approaches have been developed in the field to analyze and interpret visual data. Classical approaches rely on hand-crafted features and algorithms such as edge and contour detection to extract and interpret visual information from images or video. However, these approaches are limited in capabilities since they are unable to generalize unseen data. These techniques are also highly sensitive to variability in light, orientation, and are quite noisy – making them inefficient for complex recognition scenarios. In contrast, deep learning-based approaches to analyzing visual data are better able to generalize to a variety of complex use cases and environments as the features with this approach are learned and not hand-crafted. Image embeddings are a popular deep learning-based approach and a powerful alternative to traditional or classical CV methods. The embedding approach extracts essential features from input data, enabling models to better understand, categorize, and retrieve images accurately and efficiently. In this article, we’ll dive deeper into the topic of embeddings. We’ll cover guidelines and best practices for generating embeddings and discuss typical ML workflow examples that can greatly benefit from using embeddings during data and model analysis. ## What are image embeddings? Image embeddings are compact, numerical representations that encode essential visual features and patterns in a lower-dimensional vector space. Unlike raw pixel data which primarily provides color and intensity information of the pixels, embeddings capture more abstract and meaningful attributes, such as the shape of the object, orientation, and overall semantic context. Embeddings are typically generated by computer vision models (such as CLIP or transformer-based models) which process the image through multiple layers to identify patterns and relationships within the data. Over the last decade, machine learning has further advanced the use of embeddings to capture spatial relationships and contextual information. Transferring the information of the image into an embedding makes it easy to analyze and understand patterns in data, e.g. visualizing data clusters to identify relationships or using it for image retrieval and to locate areas of poor model performance. ## How to compute embeddings There are several methods for generating lower-dimensional representations of visual data. Some examples of such methods are: \[bullet-checkmarks class="color-primary"\] - **Using pre-trained models** such as a ViT (Vision Transformer) or CLIP (Contrastive Language-Image Pretraining) which are trained on large datasets that enable them to recognize complex patterns and structures. - **Fine-tuning a large vision model** on a specific task so the embeddings extracted are better descriptive for a specific task or dataset. - **Training an autoencoder** to learn a lower dimensional representation of your data. This is useful when your data belongs to a specific domain and you don’t need the complex features learned by a pre-trained vision model. - **Using dimensionality reduction techniques** like [UMAP](https://github.com/lmcinnes/umap) or [t-SNE](https://lvdmaaten.github.io/tsne) on less complex images to generate embeddings. \[/bullet-checkmarks\] The method used often depends on the complexity of the domain and the variation in the source data. Once embeddings are generated, they can be consumed by visualization tools to further analyze the relationship between samples. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## Three powerful embeddings use cases in ML workflows Embeddings can be used in various ways to gain a deep understanding of your data and model performance. We’ll go over some typical use cases with examples. These use cases are tool-agnostic. However, here we showcase them using FiftyOne – a powerful tool that helps AI builders develop high-performing models by providing data insights at every phase of the ML workflow, from data exploration and analysis to model evaluation and testing. ### Using embeddings to understand structure in raw data Embeddings can reveal the underlying structure of the data by capturing hidden relationships and distributions in the original data. You can visualize embeddings by plotting the data in 2D space to better understand patterns in the data. This can be achieved by running dimensionality reduction methods like UMAP or PCA on your data to transform it into 2D or 3D space. You can also understand if the data naturally forms any particular clusters and what information those clusters convey. Let’s take a look at the [CIFAR-10 dataset](https://www.cs.toronto.edu/~kriz/cifar.html). This open dataset contains about 60K images spread across various categories. Visualizing embeddings on the dataset reveals several distinct distributions. After further drill-down into a few clusters, you’ll notice that every cluster is somewhat of a distinct category: images of ships, automobiles, trucks, horses, etc. This type of analysis can provide a first level of understanding of your image data.Visualizing image embeddings to understand clusters of images For unlabelled data, natural clustering is also used for pre-annotation and tagging workflows; where images are classified and categorized by certain attributes before they are queued up for annotation. For example, we can select the leftmost cluster. These images aren’t labeled but through visual inspection, we can see that this cluster contains car images. Now we can batch-select these images, tag them as “modern cars”, and queue them for annotation. This type of workflow significantly speeds up annotation as the tags already indicate the ground truth labels for human annotators and they can instead focus their time on drawing bounding boxes for those objects. Interested in learning more about image clustering? Check out this tutorial that teaches you [how to cluster images](https://docs.voxel51.com/tutorials/clustering.html?_gl=1*1gzzklk*_gcl_au*MTczMTA4OTI0OS4xNzM4MjYyNjc1) using scikit-learn and feature embeddings. ### Using embeddings to QA annotation quality Embeddings can be used to find annotation errors in your dataset. By coloring embedding vectors by fields, you can easily check if the samples match the clustering intuition and identify samples that could be incorrectly labeled. Let’s take a look at the [Berkeley Driving Dataset (BDD100)](https://bair.berkeley.edu/blog/2018/05/30/bdd/) – a large-scale diverse driving dataset that contains annotated data of multiple cities, weather conditions, times of day, and scene types. Coloring embeddings by timeofday.label you can see two distinct clusters. A quick visualization of them tells you that the right side cluster is “daytime” images and the left side is “nighttime” images, with dawn/dusk samples scattered over the 2 sides. Image embeddings on Berkeley Driving Dataset (BDD 100) showing 2 distinct clusters of nighttime and daytime samples After closer observation, you might notice that certain “green” night samples appear in the daytime cluster on the right. By visualizing these specific night samples, we can see that they are images taken during the day. The ground\_truth does not match what the model predicted, indicating that the model is accurate, but the samples have a labeling issue and are incorrectly classified. Through this process, labeling issues can be easily identified and corrected.Coloring image embeddings by "time of day" to identify mislabeled imagesUsing image embeddings for QA analysis. In this sample “daytime” is incorrectly labeled “nighttime” ### Using embeddings to find similar samples Another interesting scenario where embeddings play a role is to find similar or unique objects from your dataset. This type of workflow is particularly useful for cases where you want to get an understanding of a certain category of images to possibly augment your dataset with more of those similar image types (perhaps with variations) for improving model performance. For example, if you are building an e-commerce visual recommendation system, you’ll need to see if your models contain enough images of the product photographed under different conditions (angles, backgrounds, lighting) so it can do a good job of generalizing and identifying true visual similarities. Similarity searches by text or images use embeddings to find and display the appropriate images. Text and image embeddings can be used to identify relevant images because visually or semantically similar objects are mapped closer to each other in the 2D space. Let’s take a look at a dataset that has random images of people, animals, transportation, food, and other categories. Let’s select the image of an airplane and then find similar samples. You can do this by sorting samples by similarity to visualize all images in your dataset that look like airplanes.Similarity search uses embeddings to find samples that are mapped closer to each other Another way to look for similar samples is to use natural language prompts so you can “talk to your data”. Here we are trying to find samples of “pedestrians using a crosswalk”. A quick similarity image search yields the corresponding samples that closely relate to that prompt. ### Summary: Embedding use cases As we saw from the use cases, embeddings are transformative techniques that enable a deeper understanding of the underlying data through clustering and visualization. They reveal hidden structures in raw data, help identify patterns, speed up annotation QA workflows, and aid in finding visually or semantically related samples to enhance the diversity of the dataset and improve model performance. ## The future of image embeddings The growing use and specialization of embeddings are driving the development of accurate, reliable, and robust datasets and models. ### Embeddings for generative tasks A popular area of active innovation is the accurate representation of text-to-image conversion. Over the last few years, we’ve seen significant progress with models like DALL-E and CLIP that generate images from textual descriptions by learning joint embeddings of text and images. Such models enable the generation of textual context-aware embeddings and enable workflows like text similarity search. Advances in these techniques will further improve the accuracy of text-to-image conversion, create realistic AI-generated art, and perhaps even greatly assist with the scary side of deepfakes where images and videos are manipulated by encoding facial features into embeddings and reconstructing them onto other faces. ### Efficient and lightweight embedding models There is also considerable innovation ( [here](https://arxiv.org/abs/2009.07409) and [here](https://arxiv.org/pdf/2302.08387)) in smaller, faster models (e.g. MobileNet) that generate embeddings for use in real-time applications such as edge computing and mobile devices to improve the performance of ML systems, particularly in resource-constrained environments. ## How to use embedding capabilities in FiftyOne [FiftyOne](https://voxel51.com/), from Voxel51, is a solution designed to make it easy for AI builders to develop high-quality data and robust models by providing data insights through every critical phase of the ML workflow. FiftyOne simplifies the computation, visualization, and application of image embeddings, making it easier for AI builders and organizations to unlock the full potential of their visual data. The image embedding capabilities in FiftyOne analyze datasets in a low-dimensional space to reveal interesting patterns and clusters that can answer important questions about your dataset and model performance. Manual solutions to compute, visualize, and use embeddings in workflows can be time-consuming and harder to scale. FiftyOne supports out-of-the-box advanced techniques (powered by [FiftyOne Brain](https://docs.voxel51.com/brain.html?_gl=1*1v0oxwh*_gcl_au*MTczMTA4OTI0OS4xNzM4MjYyNjc1)) that make it easy to compute and visualize embeddings. By incorporating embedding-based workflows into the ML pipeline, FiftyOne equips teams with the data-centric capabilities they need to get the data insights for developing robust models. Here are some resources that can help you start using embeddings in your workflow \[bullet-checkmarks class="color-primary"\] - Tutorial on [How to use embeddings in FiftyOne](https://docs.voxel51.com/tutorials/image_embeddings.html?_gl=1*yeowg3*_gcl_au*MTczMTA4OTI0OS4xNzM4MjYyNjc1); follow along using this [Google Colab notebook](https://colab.research.google.com/github/voxel51/fiftyone/blob/v1.0.2/docs/source/tutorials/image_embeddings.ipynb). - Blog post: [Visualizing data with dimensionality reduction techniques](https://voxel51.com/blog/how-to-visualize-your-data-with-dimension-reduction-techniques/) - Practical pointers for [using embeddings with these tips-and-tricks](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023/) \[/bullet-checkmarks\] ## Conclusion Image embeddings have transformed how ML engineers analyze data to solve complex visual tasks, from object detection to anomaly detection and beyond. ML workflows can greatly benefit from using embeddings to improve the efficiency and reliability of processing and analyzing large datasets and visual AI models. ## Next steps Building visual AI successfully is possible with the right solution. Refer to FiftyOne docs on [how to get started](https://docs.voxel51.com/getting_started/install.html?_gl=1*1sblmkt*_gcl_au*MTczMTA4OTI0OS4xNzM4MjYyNjc1). Looking for a scalable solution for your visual AI projects? Check out [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) and [connect with an expert to see it in action](https://voxel51.com/book-a-demo/). [Visual AI](https://voxel51.com/blog/tag/visual-ai) [embeddings](https://voxel51.com/blog/tag/embeddings) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [images](https://voxel51.com/blog/tag/images) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/04eff439f21847f1068ea3f38b72f8bc7ec77f0e-2340x1308.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Image Similarity Search: Unlocking Pattern Detection in Visual Data\\ \\ Learn\\ \\ • \\ \\ Apr 16, 2025](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data) [![](https://cdn.sanity.io/images/h6toihm1/production/ca4f84addae3f8e97daad00cf856754301cbb5e0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Are Image Segmentation Maps Superior to Bounding Boxes?\\ \\ Learn\\ \\ • \\ \\ Feb 26, 2025](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes) [![](https://cdn.sanity.io/images/h6toihm1/production/9252e8bb5db5c4805f4a6f315b51527ee5c22072-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Guide to AI Image Segmentation\\ \\ Learn\\ \\ • \\ \\ Dec 19, 2024](https://voxel51.com/blog/a-guide-to-ai-image-segmentation) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-59-lllmstxt|> ## Computer Vision Research [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Computer Vision Research World-class computer vision research teams rely on the FiftyOne data platform to assess model failures, debug samples, and curate datasets for AI development. Check out our original research and citations from researchers at Google, Porsche, Microsoft, and more. [Get started for free](https://github.com/voxel51) [Learn about FiftyOne](https://voxel51.com/fiftyone) ![](https://cdn.sanity.io/images/h6toihm1/production/623afea741dc8a90fe89a4b4f1d61ca7276e17b0-1600x794.jpg?auto=format&dpr=2&fit=max&q=75&rect=0,42,1600,752&w=800) ![Computer vision embeddings](https://cdn.sanity.io/images/h6toihm1/production/eda6b1f41ad09ea94312fe8902f8837fa89cb74a-1200x1200.png?auto=format&dpr=2&fit=max&q=75&w=600) ![](https://cdn.sanity.io/images/h6toihm1/production/e0aab8017929a10411f1bd5f2da0f71de5417767-1163x1209.png?auto=format&dpr=2&fit=max&q=75&w=582) ![Model comparisons with samples](https://cdn.sanity.io/images/h6toihm1/production/2f8a83b43fce5df618bf04f71c60a75137c32c73-2560x1200.png?auto=format&dpr=2&fit=max&q=75&rect=765,211,1795,989&w=898) By the numbers ## Leading CV researchers rely on FiftyOne 0M Downloads of FiftyOne 0+ Enterprise Customers 0K+ Community Members > "FiftyOne enables researchers to analyze and improve the quality of their datasets rapidly, replacing the weeks of manual labor that would otherwise be required without this technology. High-quality data is critical to the success of machine learning systems. Without the right tools to analyze and curate datasets, machine learning development can be inefficient and ineffective.” > > **Jordi Pont-Tuset** > > Research Scientist at Google ![](https://cdn.sanity.io/images/h6toihm1/production/66938eaaa4ff7c21ee6b7c3c5fefbc004ee6d7c9-272x92.svg) > “FiftyOne is our primary resource for machine learning research. Thanks to FiftyOne's convenient field visualizations and filtering capabilities, we can easily distinguish incorrect labels and predictions, and therefore iterate on models faster than ever. As a result, we've achieved a 77% reduction in images sent for manual verification.” > > **Ryan Szeto** > > Senior Computer Vision Engineer at SafelyYou ![](https://cdn.sanity.io/images/h6toihm1/production/2b86407c4eee830da8897a6d94e08ad7f776c698-2550x834.png?auto=format&dpr=2&fit=max&q=75&rect=42,117,2439,582&w=100) > “As we developed our Florence-2 model, FiftyOne proved invaluable for data management and visualization. Its powerful capabilities helped streamline our workflow, ensuring we built a robust foundation for our models. Now, as we dive into the development of Florence-5B, we're relying on FiftyOne more than ever. The tool's intuitive interface and rich feature set are essential for effectively managing our large datasets and gaining critical insights.” > > **Bin Xiao** > > AI Researcher, Meta (formerly Principal Research Manager, Microsoft GenAI) ![](https://cdn.sanity.io/images/h6toihm1/production/d9b4bdb75662f7f850d7abbf017918438f9f4885-3333x713.png?auto=format&dpr=2&fit=max&q=75&w=100) > “FiftyOne has helped us speed up investigations by 3x. For example, if we see a wrong suction cup grasping an item, we can quickly visualize the issue across all data sources and identify what went wrong.” > > **Dimitry Pechyoni** > > Senior Principal Machine Learning Engineer at Berkshire Grey ![](https://cdn.sanity.io/images/h6toihm1/production/30b0acc56eeba0fd43bf86f4a390dd985d1916cd-640x152.png?auto=format&dpr=2&fit=max&q=75&w=100) > "Data visualization is crucial for understanding model performance during training and validation. FiftyOne provides built-in features, plugins, and automated pipelines that help you analyze inference outputs effectively." > > **Tin Stribor Sohn** > > Technical Lead Vehicle Data Analytics Automated Driving at Porsche AG > > Doctoral Candidate at Karlsruhe Institute of Technology ![](https://cdn.sanity.io/images/h6toihm1/production/0e2f149ba6e28a4e307e55a8ad52f81f790571df-921x96.svg) > “We use FiftyOne to organize large research datasets. My favorite feature is the ability to view distributions over image attributes in the dataset, and filter the dataset by those attributes.” > > **Brett Israelsen** > > Principal Research Scientist, AI, Raytheon ![](https://cdn.sanity.io/images/h6toihm1/production/293dc34ac3ccdd7973ac04b5a27e66a0c0862260-158x61.svg) ML Research ## Auto-labeling rivals human performance The latest paper from Voxel51's ML researchers, _Auto-Labeling Data for Object Detection_, benchmarks auto-labeling against human annotation. We reveal how foundation models can deliver labels at near-human accuracy, while reducing annotation costs by up to 100,000×. [Read the research](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance) ![](https://cdn.sanity.io/images/h6toihm1/production/6ef6221c55258d5132bdb6deaa8c5494cbbd7cda-3840x2161.png?auto=format&dpr=2&fit=max&q=75&rect=11,11,3829,2150&w=1915) Papers ## Publications citing FiftyOne Download and visualize Open Images dataset [Google Open Images](https://storage.googleapis.com/openimages/web/download_v7.html#download-fiftyone) Multimodal AI for Efficient Medical Imaging Dataset Curation [Booz Allen Hamilton](https://youtu.be/1FwSbj_DccI) Auto Labeling for Object Detection [arXiv](https://arxiv.org/abs/2506.02359) Drive4C: A Closed-Loop Benchmark on What Foundation Models Really Need to Be Capable of for Language-Guided Autonomous Driving [Porsche CVF](https://openaccess.thecvf.com/content/CVPR2025W/WDFM-AD/papers/Sohn_Drive4C_A_Closed-Loop_Benchmark_on_What_Foundation_Models_Really_Need_CVPRW_2025_paper.pdf) TryOffDiff: Virtual-Try-Off via High-Fidelity Garment Reconstruction using Diffusion Models [CVPR 2025](https://rizavelioglu.github.io/tryoffdiff/) Continuous Patient Monitoring with AI: Real-Time Analysis of Video in Hospital Care Settings [Frontiers in Imaging](https://lookdeep.github.io/ai-norms-2024/) Class-wise Autoencoders Measure Classification Difficulty And Detect Label Mistakes [arXiv](https://arxiv.org/abs/2412.02596) A Framework for a Capability-driven Evaluation of Scenario Understanding for Multimodal Large Language Models in Autonomous Driving [Porsche IEEE IAVVC 2025](https://arxiv.org/abs/2503.11400) Zero-Shot Coreset Selection: Efficient Pruning for Unlabeled Data [arXiv](https://arxiv.org/abs/2411.15349) Citations ## Cite FiftyOne in your research FiftyOne is open source and free for reasearchers. Sample citation: Moore, B. E., & Corso, J. J. (2025). Voxel51's FiftyOne (Version 1.7.0) \[Software\]. Available from [https://github.com/voxel51](https://github.com/voxel51). Computer vision. > “FiftyOne completely streamlined my workflow. Instead of manually clicking through images and masks, I can now view, compare, and tag everything in one place. It’s made data evaluation so much easier and more efficient.” > > **Navjot Singh** > > Precision Ag Tech, Texas A&M University ![](https://cdn.sanity.io/images/h6toihm1/production/71577ee4e9d9d72e9b9b8c85222e0b16e6fafae8-484x104.png?auto=format&dpr=2&fit=max&q=75&w=100) > "FiftyOne has been a game-changer for our fieldwork. We built a local user interface for data analysis using FiftyOne, which empowers our scientists to better understand automatically sorted insect data and efficiently retrain machine learning models. One of the standout benefits is its ability to run completely offline, without the need for internet access or cloud storage. This is absolutely critical for field scientists working in remote locations. It’s a tool that truly meets us where we are.” > > **Dr. Andrew Quitmeyer** > > Digital Naturalism Lab ![](https://cdn.sanity.io/images/h6toihm1/production/8bf8022994ecb609b285a0753e10a2f44764cf22-548x296.png?auto=format&dpr=2&fit=max&q=75&w=100) > “FiftyOne has allowed me to efficiently organize and visualize data, which is key in research. I have used it in several projects to analyze datasets before training. In a recent person re-identification project, I’m creating a new dataset and utilized the FiftyOne embedding visualization feature to perform clustering on my data and optimize labeling. Its simplicity and power make it possible to explore data with just a few lines of code. It’s a tool that saves time and helps focus on strategic tasks. I recommend it to any researcher looking to optimize their workflow.” > > **Carlos Hinojosa** > > Post-Doctoral Fellow at KAUST ![](https://cdn.sanity.io/images/h6toihm1/production/618f2c6696342caa43107d87b9fc5285bb3da21f-600x176.png?auto=format&dpr=2&fit=max&q=75&rect=10,25,574,112&w=100) > “FiftyOne streamlines the process of visualizing, exploring, and filtering data, enabling quick identification of duplicates and mislabeled samples.” > > **Rıza Velioğlu** > > PhD Researcher at Bielefeld University ![](https://cdn.sanity.io/images/h6toihm1/production/42c7d09cbb2cbb2395feca1a98cb7658a698ffcf-460x109.png?auto=format&dpr=2&fit=max&q=75&rect=8,0,452,109&w=100) > “FiftyOne helped me drastically reduce the time needed to validate inferences in a large-scale aerial imagery project. I was identifying solar panels across North Carolina using high-resolution aerial imagery. By saving inference outputs as image tiles and leveraging FiftyOne's clustering capabilities, I could visually explore data cluster by cluster instead of image by image. This streamlined the validation process and saved hours of manual work. Having all the tools—visualization, high-dimensional clustering, and model inspection—in one web-based interface made a huge difference in simplifying my workflow.” > > **Kshitiz Khanal** > > Research Associate at the Institute for Transportation Research and Education at NC State University ![](https://cdn.sanity.io/images/h6toihm1/production/f66466f6cbf45fb52e1d99df9df3f66d34daadb4-121x11.svg) ## Get started with FiftyOne for your computer vision research [Download open source](https://github.com/voxel51) [Talk to sales](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ea42e9b26f49f1cb54bb8aca31dc10e7f74fe11f-3024x960.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-60-lllmstxt|> ## Autonomous Vehicle Solutions [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Visual AI in Autonomous vehicles & solutions FiftyOne accelerates development of autonomous and driver assistance solutions, helping teams get their AI on the road. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/a703d2ec88281308526c4245efa390b520ed6d62-628x395.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/1bf13d3e5ab0e74206193b5f6daba34db30f68ed-628x817.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/bbfabe2f9f88ea60cfb232015276d58138faa30c-628x813.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/225fce214a9cdd950c6a704132485d3157d4ee39-628x393.png?auto=format&dpr=2&fit=max&q=75&w=314) Benefits ## Systematically analyze and refine datasets and models with FiftyOne Unlock your AI’s full potential. From ADAS to driver monitoring and more, data quality makes or breaks AI model performance. FiftyOne helps you curate high-quality datasets, analyze data and model results, and streamline development—turning R&D into real-world success. ### Increase productivity and efficiency Automate dozens of computer vision workflows with FiftyOne so you can free up valuable engineering time and focus on what matters most. ### Deliver features and innovations faster Shave months of development time off your vision-based AI projects by using FiftyOne to accelerate data curation and model evaluation. ### Save money Avoid the unnecessary costs of manual data wrangling and overpaying for annotation and data collection by using FiftyOne to streamline and automate how you work with data. Use cases ## Power any visual AI use case in automotive Computer vision plays a critical role in a wide variety of automotive use cases. That’s why leaders and innovators building solutions rely on FiftyOne. ![](https://cdn.sanity.io/images/h6toihm1/production/4699ca5463ad5be11715eb3c7c4663e99eb5e810-768x550.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Lane and object detection Enhance the ability to detect and follow traffic lanes in tough conditions by quickly spotting gaps and outliers in training data and model results. ![](https://cdn.sanity.io/images/h6toihm1/production/0ebeb59422774ba93e38795f5d3997f10c672633-768x497.jpg?auto=format&dpr=2&fit=max&q=75&w=384) In-cabin analytics Deliver new safety and improved driver experiences by using data insights to build reliable detection and analysis models. ![](https://cdn.sanity.io/images/h6toihm1/production/71750a23bdd539c71ee888ad4d5010b5f3b7c792-768x420.png?auto=format&dpr=2&fit=max&q=75&w=384) Multi-camera sensor fusion Group together visual data samples from multiple sensors and locations to create a better understanding of dynamic environments. ![](https://cdn.sanity.io/images/h6toihm1/production/ec8449933ece119d748172cc66ea7daf8d0cb825-715x521.png?auto=format&dpr=2&fit=max&q=75&w=358) Lidar and radar point clouds Build advanced solutions by being able to easily explore and visualize 3D data from multiple angles and sources. ![](https://cdn.sanity.io/images/h6toihm1/production/cb77cda69dc3379d2673b6fc216bf9e7fc505bbe-600x400.png?auto=format&dpr=2&fit=max&q=75&w=300) Semantic segmentation Analyze and improve perception of roads, obstacles, pedestrians, signs, and more to ensure safe navigation. ![](https://cdn.sanity.io/images/h6toihm1/production/be0c73dede6d0f3caffc39015a63b5a5f4e4110a-768x432.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Video object tracking Visualize videos, ground truth detections, and predicted trajectories to gain insights into your data and to evaluate model accuracy. Features ## How visual AI can help you ### Easily organize millions of samples in multiple formats Hands-free driving systems and vehicles rely on massive amounts of data. FiftyOne makes it easy to manage your samples across dozens of formats. Multimodal datasets: images, videos, clips, frames, geolocation, and 3D lidar and radar point clouds Any metadata you need: time of day, camera or device ID, location information, weather conditions, and anything else you need in your AI workflows Any model you’re working with: lane detection, object detection, semantic segmentation, and many more ![](https://cdn.sanity.io/images/h6toihm1/production/efc2dd080cda4cfab1e59377f015f2b984d9c703-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Quickly find the subsets of data you want Sifting through massive amounts of data is like searching for a needle in a haystack. Pinpoint samples of interest in seconds using FiftyOne. Create meaningful, balanced datasets: query samples by metadata to correct for imbalances Accelerate training data selection: quickly find unique scenarios and anomalies in your data streams Cover the edge cases: identify hard samples to strengthen your datasets and model performance ![](https://cdn.sanity.io/images/h6toihm1/production/22885a78a7c957147b1e93fc91cd4bb1736a1b2e-2560x1806.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Only annotate what you need Generating annotations can be complex, cumbersome, and costly. FiftyOne integrates with your favorite annotation tools to become your mission control for annotation workflows. Stop passing data around: collaborate with teammates and vendors on a single source of truth Stop overpaying for annotations: identify your most valuable samples to annotate, then automatically send them to your annotation vendor Mitigate annotation mistakes: assess the quality of your annotations to improve both your datasets and models ![](https://cdn.sanity.io/images/h6toihm1/production/ee90c73a52d84027e4b0eb1f7fbce730cc514644-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Continuously build models that perform Models don’t always perform on new, unseen data. FiftyOne gives you the ability to visualize and compare model performance so you can deploy into production with peace of mind. Understand your model’s failure modes: browse model performance at the sample level so you can take the right steps to address failures Embrace continuous evaluation: integrate FiftyOne into your training pipeline to evaluate and improve model performance and datasets with every model update Manage dataset versions: track revisions to your datasets so you can view or rollback to previous versions of your datasets at any time ![](https://cdn.sanity.io/images/h6toihm1/production/4fb68f28115173c89811e57ddde635273c8c8e99-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) features ## Designed for AI builders FiftyOne natively supports and enables the computer vision building blocks needed to develop robust automotive AI solutions. ![](https://cdn.sanity.io/images/h6toihm1/production/b65da79f5913eb1e20235f7bc09363e94b9dca14-2560x960.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,2560,960&w=1280) - Classification - Detection - Segmentation - Polygons and polylines - Keypoints - Pointclouds - Heatmaps - Geolocation - Embeddings - Multiview datasets - Images, videos, and 3D data > “From vehicle safety and autonomy to security systems to robotics, Bosch is a leader in artificial intelligence solutions utilizing computer vision. Voxel51’s solutions help us organize, evaluate and refine our data and models, enabling us to develop robust, reliable AI applications across multiple teams and projects. ” > > **Arvind Kumar Shekar** > > Lead Expert AI Validation, Bosch ![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) Resources ## Learn more about visual AI in automotive Learn how visual AI is increasing safety, autonomy, and efficiency in automotive and mobility applications. ### Solving the AI Blindspot: Using Data to Drive Models in Automotive Too many automotive AI projects get stuck in neutral because of blindspots that lead to failures in the real world. The solution starts with changing how you work with your data. [Read more](https://voxel51.com/blog/solving-the-ai-blindspot-using-data-to-drive-models-in-automotive) ![](https://cdn.sanity.io/images/h6toihm1/production/1c0fb79c49e38efb3d479589e0ee43daa842bbf3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=960) ### How to Make the Best Self-Driving Dataset In the race to safely release fully autonomous vehicles, understanding self-driving datasets is key. Join Daniel Gural, Machine Learning and Developer Relations expert at Voxel51, as he dives deep into the tools and techniques shaping the future of self-driving technology. [Read more](https://voxel51.com/blog/how-to-make-the-best-self-driving-dataset) ![](https://cdn.sanity.io/images/h6toihm1/production/aa202a26af141cab6f1c7785026a9edf71ef8101-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=960) ## Data eats models for lunch Talk to our computer vision experts to start building better datasets and models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-61-lllmstxt|> ## Point Cloud Data Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) A Comprehensive Guide to Working with Point Cloud Data Jul 25, 2025 • 6 min read Article content In this article [What is point cloud data?](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#be8a383bb4d3) [Why point clouds matter in AI](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#e936b78d3ea3) [Common challenges when working with 3D point cloud data](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#93d29525d245) [How FiftyOne helps with point cloud workflows](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#6cd4db4d54fd) [The future of point clouds and computer vision](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#2c5456d33402) [Getting started with point clouds in FiftyOne](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#deb465a3ea8f) In this article [What is point cloud data?](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#be8a383bb4d3) [Why point clouds matter in AI](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#e936b78d3ea3) [Common challenges when working with 3D point cloud data](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#93d29525d245) [How FiftyOne helps with point cloud workflows](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#6cd4db4d54fd) [The future of point clouds and computer vision](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#2c5456d33402) [Getting started with point clouds in FiftyOne](https://voxel51.com/blog/comprehensive-guide-point-cloud-data#deb465a3ea8f) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) When we talk about visual data, most people picture 2D images and videos. But for an increasing number of machine learning (ML) workflows—especially in autonomous driving, robotics, and industrial inspection—point clouds are where the real action is. A [point cloud](https://en.wikipedia.org/wiki/Point_cloud) is a 3D representation of the world: a set of data points in space, each representing a precise location captured by LiDAR, radar, depth cameras, or photogrammetry. Rich with geometric detail, point clouds are indispensable for tasks like obstacle detection, surface inspection, object tracking, and 3D scene reconstruction. ![](https://cdn.sanity.io/images/h6toihm1/production/876a170fe7f62c5464c8d83ccb41de344333f691-1600x797.png?auto=format&dpr=2&fit=max&q=75&w=1600) But while point cloud data unlocks powerful capabilities in machine learning and computer vision, they also introduce significant complexity. This guide walks through what point clouds are, why they matter, the challenges they pose, and how tools like FiftyOne from Voxel51 help make point cloud workflows practical, scalable, and production-ready. ## **What is point cloud data?** A **point cloud** is a set of data points in 3D space. Each point contains coordinates (X, Y, Z), and may include additional attributes like color, intensity, time, or classification labels. ![](https://cdn.sanity.io/images/h6toihm1/production/2517b5b7cd93d336a1a8af86079ea8c5f85c2ab4-300x300.gif?auto=format&dpr=2&fit=max&q=75&w=300) Modern sensors generate point clouds through various methods. LiDAR sensors emit laser pulses and measure return times to create highly accurate 3D maps, essential for autonomous vehicle perception systems. Photogrammetry reconstructs 3D scenes from multiple 2D photographs, while depth cameras like Microsoft's Kinect combine infrared patterns with traditional imaging. Each capture method produces point clouds with distinct characteristics - LiDAR excels at long-range accuracy but generates sparser data, while photogrammetry creates dense colorful reconstructions but requires good lighting and textured surfaces. Unlike images, which are organized in 2D grids of pixels, 3D point cloud data is unstructured. It captures 3D geometry without enforcing a fixed spatial layout—providing greater flexibility while maintaining high precision. This fundamental difference makes point clouds both incredibly powerful for capturing real-world geometry and challenging to process with conventional computer vision techniques. ## **Why point clouds matter in AI** Point clouds offer critical advantages in computer vision tasks where understanding depth, shape, and spatial relationships is essential. Key domains include: - Autonomous vehicles: Detecting and tracking pedestrians, vehicles, and road infrastructure. - Robotics: Navigating cluttered environments, performing pick-and-place tasks, mapping indoor spaces. - Manufacturing: Inspecting parts and assemblies in 3D for defects or misalignments. - Construction and infrastructure: Surveying sites, monitoring structural integrity, modeling buildings. - Geospatial analysis: Mapping terrain, forests, urban environments. - Medical imaging: Guiding orthopedic implants during surgery, fitting custom dental appliances These applications demand robust, real-time understanding of complex environments—and point cloud data is often the most reliable foundation. ## **Common challenges when working with 3D point cloud data** Despite their value, point clouds are notoriously difficult to work with. Key pain points include: ### Lack of structure Point clouds are unordered and sparse, making them difficult to process using conventional 2D computer vision pipelines. Tasks like segmentation, classification, or registration require specialized 3D algorithms and libraries (e.g., Open3D, PCL). ### Annotation complexity Labeling point cloud data is time-consuming and expensive. 3D bounding boxes and semantic segmentations require specialized tools and expertise, especially when fusing data from multiple sensors. ### High volume and storage overhead Point clouds often contain millions of points per frame. Managing, storing, and loading this data—especially in multimodal datasets that also include video or images—can become a major bottleneck. ### Limited visibility into model performance Without effective visualization and filtering, it's hard to spot failure cases, debug issues, or understand what the model is missing—especially when data is captured in 3D and evaluated in aggregate metrics only. ## **How FiftyOne helps with point cloud workflows** FiftyOne brings clarity to point cloud data management by treating 3D data as a first-class citizen alongside images and videos. ### Point cloud data visualization FiftyOne offers [native point cloud visualization](https://docs.voxel51.com/user_guide/using_datasets.html#point-cloud-datasets) with an interactive 3D viewer that lets users explore data from any angle with full 3D scene support with meshes and arbitrary geometries. Working with point clouds presents unique technical challenges that FiftyOne helps address. **Large point cloud datasets** can contain billions of points, creating storage and processing bottlenecks. FiftyOne's efficient data structures ensure smooth performance even with massive datasets. The platform's `compute_orthographic_projection_images()` function generates 2D bird's eye views for quick dataset navigation, solving the common problem of point clouds lacking natural thumbnails for grid view browsing. ![](https://cdn.sanity.io/images/h6toihm1/production/ce57554e044e29672c472c126b905c0092652c7b-2290x2036.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### Point cloud data curation FiftyOne’s approach to **point cloud dataset curation** addresses a critical pain point for ML teams. Traditional tools often treat point clouds in isolation, but real-world applications frequently combine 3D data with images, videos, and other sensor modalities. FiftyOne's [grouped datasets](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023) enable seamless integration of point clouds with 2D data, essential for autonomous driving applications where LiDAR scans complement camera feeds. Teams can filter, query, and analyze their 3D data using the same powerful interfaces they use for images, dramatically reducing the learning curve. ### Point cloud annotation and labeling **Point cloud annotation and labeling** becomes significantly more efficient with FiftyOne's integrated workflow. The platform supports [3D bounding box annotations](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023) with arbitrary rotation angles, crucial for detecting vehicles, pedestrians, and other objects in autonomous driving scenarios. Beyond basic labeling, FiftyOne [enables semantic segmentation visualization](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e) through dynamic point coloring based on any attribute in the point cloud data. This flexibility means teams can visualize classification results, confidence scores, or custom attributes without switching between multiple tools. Voxel51’s FiftyOne is the industry-standard toolkit for visualizing, curating, and evaluating datasets across 2D and 3D modalities. Here’s how it supports point cloud use cases: ### Point cloud model evaluation **Point cloud model evaluation** requires specialized metrics beyond traditional 2D computer vision. FiftyOne implements 3D intersection over union (IoU) calculations for [bounding box evaluation](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#36a73634e281), essential for autonomous driving applications. Teams can identify failure modes and improve model performance by computing precision, recall, and F1-scores for 3D object detection tasks. ![](https://cdn.sanity.io/images/h6toihm1/production/bed52614c96455cbf9163a72f66df926bef4a5fd-3438x1794.png?auto=format&dpr=2&fit=max&q=75&w=1600) Data quality issues [plague point cloud projects](https://ieeexplore.ieee.org/document/9616784/), from sensor noise to annotation errors. FiftyOne's [Brain module](https://docs.voxel51.com/brain.html) applies ML-powered analysis to identify problematic samples, duplicate data, and annotation mistakes that could limit model performance. For teams struggling with **point cloud data labeling**, the platform's [integration with annotation services](https://docs.voxel51.com/integrations/cvat.html#annotating-3d-data) and [support for AI-assisted labeling](https://voxel51.com/annotation) can reduce annotation time by up to 82x compared to manual methods. ## **The future of point clouds and computer vision** As sensors become more affordable and algorithms more sophisticated, point clouds will play an [increasingly central role](https://www.metastatinsight.com/report/3d-point-cloud-processing-software-market) in computer vision applications. The convergence of point cloud AI with foundation models promises even more powerful capabilities, from zero-shot 3D understanding to automated scene generation. For organizations building the next generation of 3D computer vision applications, FiftyOne provides both the immediate tools needed today and the flexibility to adapt as the field evolves. Point clouds represent more than just another data modality - they capture the richness of our 3D world in ways 2D images cannot. By making point cloud data processing accessible through familiar interfaces and powerful abstractions, FiftyOne empowers teams to unlock insights from their data and build more capable computer vision systems. ## **Getting started with point clouds in FiftyOne** Beginning your 3D point cloud visualization journey with FiftyOne requires just a few lines of Python code. The platform supports the widely-used PCD (Point Cloud Data) format natively, with easy conversion from other formats through Open3D integration. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3import fiftyone.utils.utils3d as fou3d 4 5dataset = foz.load_zoo_dataset("quickstart-groups") 6 7min_bound = (0, -15, -2.73) 8max_bound = (20, 15, 1.27) 9size = (608, -1) 10 11fou3d.compute_orthographic_projection_images( 12 dataset, 13 size, 14 "bev_images", 15 shading_mode="height", 16 bounds=(min_bound, max_bound) 17) 18 19session = fo.launch_app(dataset) ``` [3D](https://voxel51.com/blog/tag/3d) [point cloud](https://voxel51.com/blog/tag/point-cloud) ![](https://cdn.sanity.io/images/h6toihm1/production/c98c65f91ce1141210966b7a7b090b446e51c7bb-512x512.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Nick Lotz Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/53c3b307e202d573a6fccf19494ba35c43a3dc7e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Best Practices for Evaluating AI Models Accurately\\ \\ Learn\\ \\ • \\ \\ Dec 17, 2024](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately) [![](https://cdn.sanity.io/images/h6toihm1/production/b83b51550d5f3f3fc2f98dbacb2996313896e147-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Quality Dataset Annotation Is Key to Machine Learning\\ \\ Learn\\ \\ • \\ \\ Feb 17, 2025](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning) [![](https://cdn.sanity.io/images/h6toihm1/production/d2e24d0a14de508f8ccd36eaffd3b909c9f193b9-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Image Embeddings Transform Computer Vision Capabilities\\ \\ Learn\\ \\ • \\ \\ Nov 25, 2024](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-62-lllmstxt|> ## Scaling Computer Vision AI ## Description Computer vision is transforming industries—from detecting defects in manufacturing to analyzing behavior in retail environments—but turning these technologies into enterprise-ready solutions takes more than just a good model. Success depends on having the right workflows, evaluation methods, and implementation strategy. In this webinar, Nick Lotz, Technical Marketing Engineer at Voxel 51, will walk you through the end-to-end process of adopting computer vision AI in the enterprise. You’ll explore real use cases across industries, evaluate model performance, and get a look at cutting-edge tools powering the next generation of visual intelligence. ## Presenter Bio ![Nick Lotz Headshot](https://media.datacamp.com/cms/nicklotz_headshot.jpeg) Nick LotzTechnical Marketing Engineer at Voxel 51 Nick is an expert in technical education for AI and software engineering. He has a decade of experience creating educational content for computer vision AI and DevSecOps. Previously, Nick has worked as a systems engineer and in technical enablement roles for companies like SUSE, Harness, and GitLab. [View More Webinars](https://www.datacamp.com/resources/webinars) <|firecrawl-page-63-lllmstxt|> ## Data-Centric Visual AI [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Stop guessing what’s wrong with your data​ Your vision model is only as good as your data. But without a data-centric approach you’re likely to overlook critical data issues even while your datasets are sabotaging model performance. Want to master a data-centric approach? Find out how in this free hands-on Coursera course. [Go to the course](https://www.coursera.org/learn/hands-on-data-centric-visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/f3f56271d2ae2789218e6062d5f415a6442c9463-642x642.png?auto=format&dpr=2&fit=max&q=75&w=321) ![](https://cdn.sanity.io/images/h6toihm1/production/05422f06ff6ec8718dc4970bdbbdae5a7202b991-1610x909.png?auto=format&dpr=2&fit=max&q=75&w=805) ![](https://cdn.sanity.io/images/h6toihm1/production/84f989ba78adc032275da770e72621e97d5c99a4-1917x909.png?auto=format&dpr=2&fit=max&q=75&w=959) ![](https://cdn.sanity.io/images/h6toihm1/production/4b11ac1529f6fd395f12053d7f1da33c43aeb416-451x524.png?auto=format&dpr=2&fit=max&q=75&w=226) The next revolution in computer vision isn’t about bigger models or better architectures. It’s about something far more fundamental: data quality. Yet here’s the dirty little secret about visual data quality: most teams are wasting precious time fine-tuning hyperparameters and tweaking architectures when it’s their data that is the problem. We’ve partnered with [Coursera](https://www.coursera.org/) and [UC Davis](https://www.coursera.org/partners/ucdavis) to create a hands-on course about a systematic, data-centric approach that transforms how you use data to evaluate and improve your models. Gain in-depth knowledge and practical skills on how to: Build systematic processes for data completeness and consistency checks Implement robust data labeling workflows with built-in quality control Create robust data pre-processing pipelines [Start learning now](https://www.coursera.org/learn/hands-on-data-centric-visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/ee8a535dd86eed23b6029ad9ddf07c7abd0e6fde-1024x642.png?auto=format&dpr=2&fit=max&q=75&w=600) ## Take the next step Want to put a data-centric approach to [visual A](https://voxel51.com/blog/what-is-visual-ai-going-beyond-computer-vision) I to use? FiftyOne was designed from the start to help you put your data at the center of your development. Powered by open source, FiftyOne helps you create high-quality datasets and models easily and reliably. ### Understand and analyze visual data Easily explore data of any scale and modality to discover hidden patterns and distributions and gain a richer understanding of your data. ### Refine and curate high-quality datasets Systematically identify data inaccuracies, annotation errors, and shortcomings to create high-quality datasets while reducing inefficiencies and costs. ### Evaluate and improve models Experiment with models, uncover failure modes and find data gaps as you test and refine models to maximize performance and accuracy. ### Fuel high-impact, efficient collaboration Collaborate on datasets and models. Share datasets, request reviews, or tag mistakes to improve the quality and efficiency of development. ## Learn how to get started with FiftyOne today Get connected with a visual AI expert [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-64-lllmstxt|> ## NVIDIA AI Podcast Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) NVIDIA AI Podcast: ADAS, Visual AI, and the Road Ahead with Porsche and Voxel51 Jul 30, 2025 • 1 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Most AV companies are solving the wrong problem. They're obsessing over model architectures when the data is the real differentiating factor. On the latest [NVIDIA AI Podcast](https://ai-podcast.nvidia.com/) episode, I sat down with host Noah Kravitz and Porsche's Tech Lead Tin Sohn to discuss what it takes to build safe, intelligent autonomous vehicles — and why data quality, simulation, and curation are the real bottlenecks. [Listen to the podcast](https://open.spotify.com/episode/01z4OB6WSJ6Pti0dRPSAWX) 📊 Data, not models, is the real differentiator: As foundation models become commoditized, competitive advantage lies in how teams curate training sets, evaluate performance, and systematically address failure modes. 🛠️ Simulation and synthetic data are now essential: You can’t capture every rare edge case in the real world — generating diverse, high-fidelity scenes is how AV systems learn to handle the unexpected. Platforms like @ [NVIDIA Drive](https://www.linkedin.com/showcase/nvidia-drive/posts/?feedView=all&viewAsMember=true) are powering end-to-end solutions for AVs, from data collection to model testing and validation in simulation environments. 🔍 Trust requires explainability, not just accuracy: AVs must reason like humans, understanding spatial relationships, object dynamics, and physical constraints — then communicate their decisions clearly to drivers. Black-box accuracy isn't enough when lives are at stake. 🚗 Top automakers are becoming software companies: Leading teams are moving away from off-the-shelf solutions and building internal data loops to continuously refine models with their unique domain expertise and fleet data. At Voxel51, we're proud to help companies like Porsche turn raw visual data into trustworthy, high-performance AI systems. If you’re looking to leverage your data to build industry-leading AI, [this one's worth a listen](https://open.spotify.com/episode/01z4OB6WSJ6Pti0dRPSAWX). [partnership](https://voxel51.com/blog/tag/partnership) ![](https://cdn.sanity.io/images/h6toihm1/production/8d61ff90b31d151405f9e21a33c2802509f34651-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Brian Moore Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![Your data, your advantage - the hidden cost of outsourced data annotation](https://cdn.sanity.io/images/h6toihm1/production/9733b62da57bf722c6ab7ec2dee89ec9353158a4-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Your data, your advantage: the hidden cost of outsourced data annotation \\ \\ Product & News\\ \\ • \\ \\ Jul 14, 2025](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation) [![](https://cdn.sanity.io/images/h6toihm1/production/14713df0d4dec67cd3bb5e9c292c607820df061d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Databricks and Voxel51: Scaling Data-Centric Visual AI on the Data Intelligence Platform\\ \\ Product & News\\ \\ • \\ \\ Jul 22, 2025](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/a4c2bee9ed053c5be2a1c161e5abf758c9a12ff8-1400x923.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.18 with App Performance Improvements, Sidebar Modes, and Custom Attributes\\ \\ Product & News\\ \\ • \\ \\ Nov 15, 2022](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-65-lllmstxt|> ## Secury360 Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/9090ca83d439b677604b11896dbe83589b6a9e98-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Secury360 Secury360 uses FiftyOne to balance computer vision datasets & improve model performance Apr 23, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) The [Secury360](https://www.secury-360.com/) box transforms CCTV setups into proactive perimeter detection solutions that use AI to eliminate false alarms and guarantee only human detection, with 99.998% accuracy. The Secury360 box features edge AI to gradually learn the terrain through deep learning to recognize behavior and discern if someone has bad intentions. Because only human detections get through the filter, operators have more time to respond to and prevent real threats. Secury360 has to manage very large image and video datasets that are constantly being fed by devices. FiftyOne helps Secury360 constantly improve their machine learning model by enabling them to compare similar data, and in turn, deliver a more distributed dataset to train their surveillance model on. ![](https://cdn.sanity.io/images/h6toihm1/production/65962dd5ba60a9b4e9e078144b7398b41a3e90d6-1024x726.jpg?auto=format&dpr=2&fit=max&q=75&w=1024) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-66-lllmstxt|> ## Computer Vision in Healthcare [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Industry Solutions](https://voxel51.com/blog/category/industry-solutions) Computer vision in healthcare: 12 breakthrough case studies Jul 15, 2025 • 14 min read Article content In this article [1\. Vision-based behavior monitoring in autism care: solving the wearable challenge](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#41b04a054aec) [2\. PRISM: Counterfactual medial imaging, explaining AI for medical diagnosis](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#2bb07a4ea872) [3\. Building your medical digital twin: Real-world evaluation of medical LLMs](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#cfe3f5b8fb57) [4\. Medical imaging models: Benchmarking MedGemma, VISTA-3D, and MedSAM-2](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#dd5539ebea57) [5\. Dataset curation at scale: Automating medical imaging workflows](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#3200abf9b9dc) [6\. Real-time cardiac assessment on mobile devices: Democratizing cardiac care](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#9ed7e04342e0) [7\. Continuous visual monitoring: Transforming inpatient care](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#c388afabbf41) [8\. Auto-RECIST development: AI-enabled oncology](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#a9752fa6bd1c) [9\. Foundation models in pathology: Navigating benefits and biases](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#dd3848464ddb) [10\. MedVAE: High-fidelity compression for medical images](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#a561146bb9dc) [11\. LesionLocator: Zero-shot universal tumor segmentation and tracking](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#b92a33d53472) [12\. LLMs for smarter diagnosis: Benchmarking real-world performance](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#a0c06b2530ca) [The future of computer vision in healthcare](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#fb648e663e67) In this article [1\. Vision-based behavior monitoring in autism care: solving the wearable challenge](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#41b04a054aec) [2\. PRISM: Counterfactual medial imaging, explaining AI for medical diagnosis](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#2bb07a4ea872) [3\. Building your medical digital twin: Real-world evaluation of medical LLMs](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#cfe3f5b8fb57) [4\. Medical imaging models: Benchmarking MedGemma, VISTA-3D, and MedSAM-2](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#dd5539ebea57) [5\. Dataset curation at scale: Automating medical imaging workflows](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#3200abf9b9dc) [6\. Real-time cardiac assessment on mobile devices: Democratizing cardiac care](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#9ed7e04342e0) [7\. Continuous visual monitoring: Transforming inpatient care](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#c388afabbf41) [8\. Auto-RECIST development: AI-enabled oncology](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#a9752fa6bd1c) [9\. Foundation models in pathology: Navigating benefits and biases](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#dd3848464ddb) [10\. MedVAE: High-fidelity compression for medical images](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#a561146bb9dc) [11\. LesionLocator: Zero-shot universal tumor segmentation and tracking](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#b92a33d53472) [12\. LLMs for smarter diagnosis: Benchmarking real-world performance](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#a0c06b2530ca) [The future of computer vision in healthcare](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies#fb648e663e67) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) From automating diagnostics to curating massive datasets, computer vision in healthcare is emerging as a core enabler of smarter, faster, and more equitable care. As diagnostic imaging, video, and patient monitoring data continue to grow exponentially, the ability to extract actionable insight from visual information is becoming a strategic imperative across the healthcare industry. In this post, we showcase 12 real-world case studies drawn from our recent webinar series, “Visual AI in Healthcare,” led by leading researchers and practitioners. Each one highlights a practical application of computer vision — from research labs to clinical deployments — that’s shaping the future of medicine. Whether you're building medical AI tools, leading imaging research, or deploying clinical workflows, these case studies offer a front-row seat to where visual AI is delivering impact today. ## 1\. Vision-based behavior monitoring in autism care: solving the wearable challenge **Speaker**: Somaieh Amraee - Northeastern University ### The clinical challenge: Traditional behavior monitoring for autism spectrum disorder (ASD) relies heavily on wearable biosensors, but up to 10% of patients reject these devices due to sensory sensitivity, leaving caregivers without critical data for intervention planning. ### The visual AI solution: Researchers at Northeastern University developed a [computer-vision-driven behavior analysis system](https://openaccess.thecvf.com/content/WACV2025W/CV4Small/papers/Amraee_Advancing_Multi-Person_Tracking_for_Autism_Behavior_Analysis_Challenges_Opportunities_and_WACVW_2025_paper.pdf) that uses RGB cameras and multi-object tracking (MOT) to monitor aggressive or self-injurious behaviors without requiring any wearable devices. **How vision-based monitoring system works:** - **Multi-object tracking (MOT)** to simultaneously follow patients and staff - **Pose estimation** to extract skeletal keypoints to detect non-standard postures like crawling or curling - **Action recognition** to identify dysregulated behaviors - **Behavior classification** models trained specifically on ASD patterns ### Key insights: Standard MOT algorithms like ByteTrack and DeepSORT significantly underperform in clinical environments compared to standard benchmarks. The challenges are uniquely medical, such as staff in uniforms, frequent occlusion, and non-standard patient poses. Custom clinical datasets are critical to building robust models that perform in the real environment. Vision-based behavior monitoring can fill a critical gap in autism care technology, enabling continuous monitoring for previously excluded patient populations. This requires prioritizing clinical-grade data collection and evaluation methods that reflect the complexities of real-world care environments. [Watch the full webinar](https://www.youtube.com/watch?v=iAcEDug4y_M). ## **2\. PRISM: Counterfactual medial imaging, explaining AI for medical diagnosis** **Speaker**: Amar Kumar - McGill University ### The clinical challenge: Radiologists increasingly rely on AI as a medical diagnosis tool, but black-box models provide little insight into their reasoning, especially when medical devices or imaging artifacts cloud the image. ### The visual AI solution: [PRISM](https://amarkr1.github.io/PRISM/), developed at McGill University in collaboration with [Google Research](https://research.google/), leverages Stable Diffusion 1.5 to generate counterfactual medical images that visually explain AI decisions. **Clinical applications** - **Artifact removal**: Automatically clean chest X-rays by removing pacemakers, wires, and other artifacts - **Counterfactual image generation:** Visualize the healthy version of a patient image, providing radiologists with visual evidence of model focus areas - **Synthetic data augmentation**: Create synthetic training examples that improve classifier performance by 10% Counterfactual imaging transforms AI from a black box to a transparent tool, enabling radiologists to validate model reasoning and identify potential failures before they reach patients. [Watch the full webinar](https://www.youtube.com/watch?v=e4jV5LGMIjA). ## **3\. Building your medical digital twin: Real-world evaluation of medical LLMs** **Speaker**: Ekaterina Kondrateva - Maastricht University ### The clinical question: LLMs are increasingly used for clinical triage and decision support, but their reliability varies dramatically based on input format, prompt design, and demographic factors. Kondrateva conducted a live experiment using her blood work to find how reliable LLMs are when applied to real medical data. ### Critical findings: Testing GPT-4, Claude, and DeepSeek on real patient symptoms and laboratory data revealed significant gaps: - **Cost variability:** Recommendations for follow-up testing varied wildly per model, ranging from $100-$500 - **Bias amplification:** Including demographic data led to systematically different diagnostic pathways - **Context-seeking remains broken**: Even advanced models fail at follow-up questions - **Confidence calibration varies**: Models with higher confidence scores generally produce more accurate diagnoses, but this relationship isn't consistent across model families ### Key insights: - **GPT-4 O3** leads comprehensive medical benchmarks across 35 different test categories - **Data format matters:** Converting lab results to structured JSON reduced hallucinations significantly - **Prompt engineering impact:** Chain-of-thought reasoning and context-seeking prompts improved safety measures Kondrateva’s experiment demonstrates that LLMs require human oversight and cannot function as standalone diagnostic tools in clinical settings yet. However, they can serve as a useful triage aid or second opinion when deployed with appropriate safeguards. Model success depends heavily on rigorous data preparation and sophisticated prompt design, making it essential for healthcare organizations to develop deep expertise in data quality management, input standardization protocols, and model validation frameworks before clinical deployment. [Watch the full webinar](https://www.youtube.com/watch?v=5UaYhzXayYg). ## **4\. Medical imaging models: Benchmarking MedGemma, VISTA-3D, and MedSAM-2** **Speaker:** Dan Gural - Voxel51 ### The clinical challenge: With the proliferation of vision foundation models, healthcare developers face a growing need to assess which models perform best for specific tasks (e.g., segmentation, classification, visual QA) across different imaging modalities. ### The comparative analysis: Gural presented a comprehensive demo evaluating three leading medical imaging models across clinical workflows, using [FiftyOne](https://voxel51.com/evaluation) to systematically compare model outputs, visualize failure cases, and assess performance consistency. - **Google’s MedGemma** excelled at multimodal visual question answering but struggled with precise spatial tasks - **NVIDIA’s VISTA-3D** proved most effective for large-scale 3D organ segmentation workflows - **Meta’s MedSAM-2** demonstrated superior performance in promptable segmentation and label propagation tasks ### Key insights: Model selection should prioritize task-specific performance over general capabilities. Multimodal models excel at interpretation tasks while specialized architectures deliver superior results for structured spatial analysis, making targeted deployment more effective than one-size-fits-all approaches. [Watch the full webinar](https://www.youtube.com/watch?v=h6on-EM5axw). ## **5\. Dataset curation at scale: Automating medical imaging workflows** **Speaker:** Brandon Konkel, PhD - Booz Allen Hamilton ### The clinical challenge: Creating high-quality medical imaging datasets for AI development typically requires weeks of manual work from radiologists and data engineers, creating bottlenecks in research and model development cycles. ### The visual AI solution: Booz Allen Hamilton developed a multimodal pipeline that automatically identifies relevant scans, extracts clinical context, and flags quality issues across the hospital picture archiving and communication system (PACS). This automation transformed what was previously a weeks-long manual curation process into an automated workflow that completes in just minutes. The technical stack integrates: - **BioMedCLIP** multimodal AI model for image-text embedding and protocol classification - **RadLLama** for radiology report parsing and patient assessment - **TotalSegmentator** for automated segmentation, artifact detection, and 3D version generation All outputs were embedded back into DICOM metadata and made searchable using [FiftyOne’s visual interface](https://voxel51.com/curation), while maintaining regulatory traceability, essential for both research and compliance. ### Key insights: Automated curation reduces clinical burden while improving dataset consistency. Organizations should prioritize tools that integrate with existing PACS infrastructure and regulatory frameworks. [Watch the full webinar.](https://www.youtube.com/watch?v=1FwSbj_DccI) ## **6\. Real-time cardiac assessment on mobile devices: Democratizing cardiac care** **Speaker**: Jeffrey Gao - Caltech ### The clinical challenge: Heart failure affects millions globally, yet routine echocardiograms remain inaccessible due to cost, time constraints, and sonographer shortages, leading to delayed diagnoses and preventable hospitalizations. ### The technical breakthrough: Caltech researchers developed a deep learning system that estimates right atrial pressure from handheld ultrasound devices. The model identifies the inferior vena cava (IVC), guides proper scan acquisition, and calculates pressure while running entirely on consumer iPads in real-time. The solution combines: - **Largest RAP dataset ever:** 16,000+ labeled examples from 45+ cardiologists - **Dual-model architecture:** X3D for IVC detection, SlowFast for pressure estimation - **iOS-native deployment:** CoreML and Metal shaders for edge inference - **Real-time guidance:** Interactive feedback for proper probe positioning ### Key insights: Edge deployment using consumer hardware (iPad + $4,000 probe) makes sophisticated cardiac screening accessible in urgent care settings, dramatically expanding the scope of point-of-care diagnostics. [Watch the full webinar](https://www.youtube.com/watch?v=MFI5JNMYHlE). ## **7\. Continuous visual monitoring: Transforming inpatient care** **Speaker**: Paolo Gabriel, PhD - LookDeep Health ### The clinical challenge: Hospital patient monitoring faces a fundamental gap: traditional monitoring through nurse check-ins or call buttons can be intermittent and reactive. Nurses spend only 37% of their shift in direct patient care and physicians average just 10 visits per hospital stay, leaving patients unmonitored for the majority of their admission. This creates vulnerability windows where critical events like falls, rapid clinical deterioration, and adverse reactions can occur without immediate detection. ### The visual AI solution: LookDeep Health deployed [continuous visual monitoring](https://lookdeep.github.io/ai-norms-2024/) to track patient activity 24/7 across 11 hospital systems for almost 3 years, processing over 30,000 hours of video monthly. The system features: - **YOLOv4 object detection** running at 1 FPS on embedded devices - **Hybrid deployment:** Edge interface with cloud analytics - **Privacy-first design:** On-device processing with automated blurring - **Continuous learning:** Every fourth week of data collection is held out as a test set for comparing model performance over time - **Human-in-the-loop workflows**: Seamless integration of expert labeling by using [FiftyOne](https://voxel51.com/curation) to curate data sent for labeling ### Key insights: Continuous monitoring enables proactive care coordination with minimal additional infrastructure requirements. Successful clinical deployment requires robust MLOps practices, including continuous test set evolution and metadata tagging for model reliability across diverse hospital environments. [Watch the full webinar.](https://www.youtube.com/watch?v=P6rSdAL5BMI) ## **8\. Auto-RECIST development: AI-enabled oncology** **Speaker**: Asba Tasneem, PhD ### The clinical challenge: RECIST (Response Evaluation Criteria in Solid Tumors) is a standardized protocol for evaluating cancer treatment effectiveness. However, the manual measurement process is labor-intensive and variable across radiologists, creating a bottleneck in cancer drug development. ### The visual AI solution: A comprehensive Auto-RECIST development program demonstrates the complete lifecycle of an AI-enabled medical device: from automated detection, tracking, to measurement of tumors. The development process included: - **Gold-standard dataset:** Built with expert annotations of 6,000+ patient studies - **Multi-phase validation:** Training, testing, and regulatory QA with blinded datasets - **Post-deployment monitoring:** Continuous performance tracking and data drift detection ### Key insights: Clinical AI deployment requires fundamentally different validation approaches than academic research, with emphasis on regulatory compliance, data drift monitoring, and long-term performance tracking. Organizations should plan for post-deployment monitoring as a core system requirement, not an afterthought. [Watch the full webinar.](https://www.youtube.com/watch?v=QNCh690q9ng) ## **9\. Foundation models in pathology: Navigating benefits and biases** **Speaker**: Heather (Dunlop) Couture - PixelScientia ### The clinical challenge: Pathology datasets are massive yet weakly labeled, making training from scratch slow and expensive. Meanwhile, off-the-shelf embeddings often introduce site-specific biases that limit generalizability. ### The visual AI solution: Comprehensive evaluation of how DINOv2, CLIP, and other foundation models are being used for histopathology. While these models learn rich representations from unlabeled tiles and can be fine-tuned, there are some hidden biases to look out for. - **Domain specificity matters:** Models pre-trained on histology significantly outperform ImageNet-based approaches - **Site bias persistence:** All foundation models encode hospital-specific features that enable shortcut learning - **Stain normalization limitations:** Traditional preprocessing doesn't eliminate site-specific artifacts - **Medical center robustness index:** New metrics needed to evaluate cross-site generalization ### Key insights: Foundation models enable faster prototyping and better tile embeddings, but still require rigorous validation. Pathology labs must prioritize bias auditing and multi-site validation alongside traditional performance metrics to successfully deploy models that generalize across different hospitals, scanners, and patient populations without compromising diagnostic accuracy. [Watch the full webinar.](https://www.youtube.com/watch?v=uA78AHSb81Q) ## **10\. MedVAE: High-fidelity compression for medical images** **Speakers:** Aswin Kumar & Maya Varma - Stanford ### The clinical challenge: High-resolution medical imaging — especially 3D modalities like MRI and CT — creates storage and computational bottlenecks that limit AI development, particularly for resource-constrained hospitals and academic research centers. ### The visual AI solution: [MedVAE](https://arxiv.org/abs/2502.14753) is a family of six generalizable 2D and 3D variational autoencoders that downsizes high-dimensional medical images, reducing storage by up to 512x and accelerating downstream tasks by up to 70x while preserving clinically relevant features. MedVAE was trained on over one million images from 19 open-source medical imaging datasets using a novel two-stage training strategy. **Strategic implications:** Carefully designed medical image compression can preserve diagnostic information while enabling AI development in resource-constrained environments. Medical image compression is not just a technical optimization, but a critical factor in making healthcare accessible. [Watch the full webinar](https://www.youtube.com/watch?v=5zoxHz71ZgY&ab_channel=Voxel51). ## **11\. LesionLocator: Zero-shot universal tumor segmentation and tracking** **Speaker**: Maximilian Rokuss - DKFZ German Cancer Research Center ### The clinical challenge: Tumor tracking across time series represents one of the most challenging problems in medical imaging. Current clinical practice, following the RECIST criteria (Response Evaluation Criteria in Solid Tumors), only tracks 5 target lesions maximum, missing the complete disease picture. ### The visual AI solution: [LesionLocator](https://arxiv.org/abs/2502.20985) from DKFZ combines promptable 3D segmentation with longitudinal tracking, enabling radiologists to follow lesions across entire treatment cycles with point-and-click simplicity. This is the first system to achieve near-human performance in 3D tumor segmentation and tracking. **The technical framework includes:** - **3D U-Net architecture** trained on 50+ public datasets - **Synthetic longitudinal data** for robust temporal modeling - **Multimodal support:** CT, MRI, and PET imaging compatibility - **Joint segmentation + tracking training** for consistent lesion identity - **Multi-prompt support:** Points, boxes, and mask propagation **Performance milestones:** - **Near-human accuracy:** Performance approaching human inter-rater variability on multiple tumor types, 5-10 dice point improvement over SAM and MedSAM - **Tracking consistency:** Accuracy maintained across multiple follow-up time points - **Real-world validation**: 5-fold cross-validation on real patient data ### Key insights: Joint training of segmentation and tracking models produces more robust results than sequential processing, while synthetic longitudinal data enables training where real temporal datasets are unavailable. [Watch the full webinar.](https://www.youtube.com/watch?v=Bh8tqpHFQF0) ## **12\. LLMs for smarter diagnosis: Benchmarking real-world performance** **Speaker**: Gaurav K Gupta - Lake County Health Department ### The clinical challenge: With healthcare providers increasingly using LLMs for clinical support, understanding their diagnostic accuracy across different conditions and model types is crucial for safe implementation. Gupta's systematic evaluation provides essential benchmarking data for technical teams. ### The Benchmarking Study: Systematic evaluation of leading LLMs (GPT-4, Claude, DeepSeek) across 100+ symptom-based scenarios, measuring diagnostic accuracy, confidence calibration, and bias patterns. **Performance and implementation insights:** - **DeepSeek R1** outperformed GPT-4 O3 Mini on chronic diseases (88% vs 78% category-level accuracy) despite requiring significantly fewer computational resources - **GPT-4** exceeded the newer O1 Preview model across all metrics, challenging assumptions about model progression - **Respiratory diseases** proved challenging for all models due to overlapping symptom presentations - **Prompt engineering impact:** Structured inputs with chain-of-thought reasoning significantly improved safety - **Confidence calibration:** Models with higher confidence scores didn't necessarily provide more accurate diagnoses - **Demographic bias:** Including patient demographics led to systematically different recommendations ### Key insights: Healthcare organizations should implement LLMs as augmentation tools with robust safety frameworks rather than autonomous diagnostic systems. While bias and hallucinations remain challenges, LLMs show strong potential in medical diagnostics and triage. Prompt engineering and bias auditing are as critical as model selection. [Watch the full webinar.](https://www.youtube.com/watch?v=xbFXNFMrlyA) ## **The future of computer vision in healthcare** These 12 case studies illustrate that computer vision has moved beyond proof-of-concept demonstrations to become an operational infrastructure across healthcare settings. The technology is enabling new care models, accelerating research pipelines, and democratizing access to sophisticated diagnostic capabilities. The question isn't whether to adopt computer vision in healthcare, but how efficiently your organization can build the technical capabilities, operational frameworks, and clinical partnerships needed to deploy these technologies safely and quickly. Success in computer vision in healthcare depends fundamentally on data quality and model understanding. Building robust models that perform reliably in clinical environments requires deep insight into failure modes and performance patterns. With [FiftyOne](https://voxel51.com/curation), healthcare teams can visualize exactly what their models detect, identify systematic failure patterns across different patient populations and imaging conditions, and iterate based on concrete performance data rather than abstract metrics. This data-centric approach eliminates guesswork, surfaces actionable insights about model behavior, and enables teams to confidently develop and deploy production-ready visual AI systems that meet the rigorous standards required for clinical care. [Learn more about FiftyOne](https://voxel51.com/sales). [healthcare](https://voxel51.com/blog/tag/healthcare) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [artificial intelligence](https://voxel51.com/blog/tag/artificial-intelligence) [machine learning](https://voxel51.com/blog/tag/machine-learning) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/6445eab4dfdba3548381f187231e576c97c3c5a5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Computer Vision Is Changing Healthcare\\ \\ Industry Solutions, Product & News\\ \\ • \\ \\ Aug 31, 2023](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare) [![](https://cdn.sanity.io/images/h6toihm1/production/08ee8607c02188240496e20d6f5ba968ff9e5c5b-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Visual AI in Healthcare: 2025 Landscape\\ \\ May 12, 2025](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape) [![](https://cdn.sanity.io/images/h6toihm1/production/b83b51550d5f3f3fc2f98dbacb2996313896e147-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Quality Dataset Annotation Is Key to Machine Learning\\ \\ Learn\\ \\ • \\ \\ Feb 17, 2025](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-67-lllmstxt|> ## AV Datasets Integration [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Enabling the AV Datasets of the Future with NVIDIA NuRec and FiftyOne Aug 11, 2025 • 3 min read Article content In this article [From raw data to future-ready outputs](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone#6dc5eaf0798d) [Creating validation-ready datasets at scale](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone#639da18c43b7) [Building the foundation for tomorrow’s AV data standards](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone#5961d7026fcd) In this article [From raw data to future-ready outputs](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone#6dc5eaf0798d) [Creating validation-ready datasets at scale](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone#639da18c43b7) [Building the foundation for tomorrow’s AV data standards](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone#5961d7026fcd) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) As autonomous vehicles (AV) and advanced driver assistance systems (ADAS) rapidly move from R&D to real-world deployment, the bar for dataset quality is higher than ever. Yet, teams struggle with fragmented, incomplete, and unvalidated data, which creates significant bottlenecks in the development pipeline. In addition, datasets of the future need to evolve from fixed events frozen in time to reactive, replayable scenes that can simulate completely new events. As AV and ADAS teams often work with multi-sensor datasets, e.g., LiDAR, radar, issues such as misaligned sensor calibrations, drifting ego-poses, and timestamps often introduce downstream errors in simulation workflows. By integrating [NVIDIA Omniverse NuRec](https://developer.nvidia.com/blog/accelerating-av-simulation-with-neural-reconstruction-and-world-foundation-models/) neural reconstruction libraries with [Voxel51](https://voxel51.com/)’s data engine, we can address this bottleneck with a data ingestion and validation pipeline that delivers high-quality data ready for reconstruction and simulation workflows. ## **From raw data to future-ready outputs** Voxel51’s data engine for visual and multimodal AI, [FiftyOne](https://voxel51.com/fiftyone), provides the data preparation and evaluation capabilities for understanding visual data, identifying issues and outliers, and creating clean datasets that maximize model performance. [Omniverse NuRec](https://www.nvidia.com/en-us/glossary/3d-reconstruction/) is a set of libraries and AI models that enable AV developers to bring in their fleet sensor data and generate reconstructed interactive simulation environments that can be used for AV testing, replay, and closed-loop simulation. The new integration builds a pipeline that converts users’ data on FiftyOne to NuRec’s data format to accelerate neural reconstruction and simulation. With NuRec and FiftyOne, developers can: - Ingest datasets from various AV pipelines as well as from public datasets, such as [NuScenes](https://www.nuscenes.org/) or [Waymo Open Dataset](https://waymo.com/open/). - Validate their datasets for consistency by visually inspecting the data - Catch and correct misalignments before they propagate into model training - Maintain full traceability of validated datasets across the entire dataset lifecycle Once data is ingested and inspected, the pipeline delivers a high-integrity complete dataset ready for NuRec reconstruction, without any rework necessary. ![Sensor Calibration Check In FiftyOne](https://cdn.sanity.io/images/h6toihm1/production/8ac5061d0be3dc8d833b2374dc7a68b70253f18e-1600x908.png?auto=format&dpr=2&fit=max&q=75&w=1600)Sensor Calibration Check In FiftyOne ## **Creating validation-ready datasets at scale** Instead of retrofitting fixes later, technical teams can now build quality into their datasets from the start. But this integration isn’t just about fixing data quality problems; it’s about creating datasets that improve and accelerate AV development. The reconstruction-ready data delivered by the pipeline can be further amplified with NVIDIA Cosmos Transfer, a diffusion-based world foundation model that can add new variations, such as weather, lighting, and terrain, to driving scenarios. In addition, NuRec-reconstructed scenes can be used for simulation in the CARLA open-source AV simulator for seamless and scalable workflows. ![Final export preview from FiftyOne](https://cdn.sanity.io/images/h6toihm1/production/94bb89dfc03bb6357bd881ddd08bed1f6a0aaa85-640x363.webp?auto=format&dpr=2&fit=max&q=75&w=640)Final export preview from FiftyOne Listen to the latest [NVIDIA AI Podcast](https://ai-podcast.nvidia.com/) episode, where Porsche and Voxel51 discuss what it takes to build safe, intelligent autonomous vehicles — and why data quality, simulation, and curation are the real bottlenecks. ## **Building the foundation for tomorrow’s AV data standards** AV and ADAS development demands uncompromising data quality and scalability. As the industry moves toward autonomous systems, the datasets powering these decisions will need to meet increasingly stringent validation standards across initial training, continuous learning, and regulatory compliance. The AV datasets of the future not only need to be comprehensive, they’ll also need to be clean, consistent, and export-ready by design. By combining NuRec with FiftyOne’s visual data engine, we’re helping AV developers speed up their neural reconstruction and simulation pipelines with frictionless data preparation, while maintaining the rigor required for safety-critical applications. Join our upcoming webinar on **Sept 4, 2025, @9am PT** to learn how companies such as Porsche are advancing AV development by leveraging tools like FiftyOne to build next-generation datasets and models. [Register for the webinar →](https://events.databricks.com/FY260904-WB-EngineeringRD/registration?scid=701Vp00000U6EaCIAV&utm_medium=Partner&utm_source=n/a) [dataset curation](https://voxel51.com/blog/tag/dataset-curation) [integrations](https://voxel51.com/blog/tag/integrations) [NVIDIA](https://voxel51.com/blog/tag/nvidia) ![](https://cdn.sanity.io/images/h6toihm1/production/3b39056326e925c10b46da1324bc3c5840a1629c-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Dan Gural Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/35e10bff3f7e49854806cbbd163022932bf07548-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How Voxel51 is Powering Physical AI with Databricks\\ \\ Product & News\\ \\ • \\ \\ Aug 20, 2025](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks) [![](https://cdn.sanity.io/images/h6toihm1/production/663dd6a3e6f3e57a03425f932440b5d242133451-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Search and curate video data with FiftyOne, Twelve Labs, and Databricks Vector Search\\ \\ Product & News, Integrations\\ \\ • \\ \\ Jun 5, 2025](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs) [![](https://cdn.sanity.io/images/h6toihm1/production/14713df0d4dec67cd3bb5e9c292c607820df061d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Databricks and Voxel51: Scaling Data-Centric Visual AI on the Data Intelligence Platform\\ \\ Product & News\\ \\ • \\ \\ Jul 22, 2025](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-68-lllmstxt|> ## AI & ML Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/0b65abff4c08a6ad58713d835a3418f3a7e4a38d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/0b65abff4c08a6ad58713d835a3418f3a7e4a38d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Meetups AI, ML and Computer Vision Meetup - July 17, 2025 This event has ended, but you can still catch up! Watch the on-demand recordings and register for our [future events.](https://voxel51.com/events) Jul 17, 2025 10:00 AM Pacific Virtually over Zoom! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/e0c1ed4388878d2195a81fbbaa692b46056f5146-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Daniel Fortunato SEA.AI Bio ![](https://cdn.sanity.io/images/h6toihm1/production/90a760575266716d22284c63674b9d07b606dfc2-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Sage Elliot Union AI Bio ![](https://cdn.sanity.io/images/h6toihm1/production/6dded2eca1d9e7e15e9b263de4e4423aabfd897f-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Claudia Cuttano Politecnico di Torino Bio ![](https://cdn.sanity.io/images/h6toihm1/production/408ccb09b8146d7d589ea44df0a9e4607b212880-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos, PhD Voxel51 Bio About this event Join the Meetup to hear talks from experts on cutting-edge topics across AI, ML, and computer vision. Schedule Using VLMs to Navigate the Sea of Data ![](https://cdn.sanity.io/images/h6toihm1/production/e0c1ed4388878d2195a81fbbaa692b46056f5146-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Daniel Fortunato SEA.AI Bio At SEA.AI, we aim to make ocean navigation safer by enhancing situational awareness with AI. To develop our technology, we process huge amounts of maritime video from onboard cameras. In this talk, we’ll show how we use Vision-Language Models (VLMs) to streamline our data workflows; from semantic search using embeddings to automatically surfacing rare or high-interest events like whale spouts or drifting containers. The goal: smarter data curation with minimal manual effort. Building Efficient and Reliable Workflows for Object Detection ![](https://cdn.sanity.io/images/h6toihm1/production/90a760575266716d22284c63674b9d07b606dfc2-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Sage Elliot Union AI Bio Training complex AI models at scale requires orchestrating multiple steps into a reproducible workflow and understanding how to optimize resource utilization for efficient pipelines. Modern MLOps practices help streamline these processes, improving the efficiency and reliability of your AI pipelines. SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation ![](https://cdn.sanity.io/images/h6toihm1/production/6dded2eca1d9e7e15e9b263de4e4423aabfd897f-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Claudia Cuttano Politecnico di Torino Bio Referring Video Object Segmentation (RVOS) involves segmenting objects in video based on natural language descriptions. SAMWISE builds on Segment Anything 2 (SAM2) to support RVOS in streaming settings, without fine-tuning and without relying on external large Vision-Language Models. We introduce a novel adapter that injects temporal cues and multi-modal reasoning directly into the feature extraction process, enabling both language understanding and motion modeling. We also unveil a phenomenon we denote tracking bias, where SAM2 may persistently follow an object that only loosely matches the query, and propose a learnable module to mitigate it. SAMWISE achieves state-of-the-art performance across multiple benchmarks with less than 5M additional parameters. Your Data Is Lying to You: How Semantic Search Helps You Find the Truth in Visual Datasets ![](https://cdn.sanity.io/images/h6toihm1/production/408ccb09b8146d7d589ea44df0a9e4607b212880-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Paula Ramos, PhD Voxel51 Bio High-performing models start with high-quality data—but finding noisy, mislabeled, or edge-case samples across massive datasets remains a significant bottleneck. In this session, we’ll explore a scalable approach to curating and refining large-scale visual datasets using semantic search powered by transformer-based embeddings. By leveraging similarity search and multimodal representation learning, you’ll learn to surface hidden patterns, detect inconsistencies, and uncover edge cases. We’ll also discuss how these techniques can be integrated into data lakes and large-scale pipelines to streamline model debugging, dataset optimization, and the development of more robust foundation models in computer vision. Join us to discover how semantic search reshapes how we build and refine AI systems. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-69-lllmstxt|> ## Visual Agents Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Visual Agents [![](https://cdn.sanity.io/images/h6toihm1/production/97cf3d887735ab9574b2b3e2d3825146016d875e-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Embodied Computer Vision at CVPR 2025: The Next AI Frontier\\ \\ Event Recaps\\ \\ • \\ \\ Jun 30, 2025](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier) ## Enough data wrangling.
 Request a demo. [Get started](https://voxel51.com/link-catcher) [Explore the Demo](https://voxel51.com/link-catcher) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-70-lllmstxt|> ## Auto Labeling Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Auto-Labeling [![Your data, your advantage - the hidden cost of outsourced data annotation](https://cdn.sanity.io/images/h6toihm1/production/9733b62da57bf722c6ab7ec2dee89ec9353158a4-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Your data, your advantage: the hidden cost of outsourced data annotation \\ \\ Product & News\\ \\ • \\ \\ Jul 14, 2025](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation) [![](https://cdn.sanity.io/images/h6toihm1/production/c845648a6e64e00a3cac3d875280cb462c63633d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Complete Guide to Auto Labeling\\ \\ Learn\\ \\ • \\ \\ Jun 16, 2025](https://voxel51.com/blog/the-complete-guide-to-auto-labeling) [![](https://cdn.sanity.io/images/h6toihm1/production/990a31a85d850d41fb482ae29960a8fa10b18ecd-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Behind the math: how we built the annotation savings calculator\\ \\ Product & News\\ \\ • \\ \\ Jun 4, 2025](https://voxel51.com/blog/how-we-built-annotation-savings-estimator) [![](https://cdn.sanity.io/images/h6toihm1/production/e60eea36edcff16d65c62e2c3fea99a66d36a8d5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 @CVPR 2025: Smarter, Faster Visual AI\\ \\ Event Recaps, Product & News\\ \\ • \\ \\ Jun 3, 2025](https://voxel51.com/blog/cvpr-2025) ## Enough data wrangling.
 Request a demo. [Get started](https://voxel51.com/link-catcher) [Explore the Demo](https://voxel51.com/link-catcher) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-71-lllmstxt|> ## Building GUI Agents Workshop [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/2915719445e25047720991a84be3ca110f25d8b6-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/2915719445e25047720991a84be3ca110f25d8b6-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Webinars & Workshops From Research to Reality: Building GUI Agents That Actually Work - August 15, 2025 Aug 15, 2025 9 AM Pacific Online. Register for the Zoom! About this event Welcome to the Visual Agents Workshop Series, your virtual pass to learn about visual agents - how they work, how to develop them and how to fine-tune them. Host ![](https://cdn.sanity.io/images/h6toihm1/production/c96544cfe7c8fc1e23601a34ca6a5fc11ccd6aa5-320x320.png?auto=format&dpr=2&fit=max&q=75&w=96) Harpreet Sahota Voxel51 Bio ### Part 1: Navigating the GUI Agent Landscape Understanding the Foundation Before Building The GUI agent field is evolving rapidly, but success requires an understanding of what came before. In this opening session, we'll map the terrain of GUI agent research—from the early days of MiniWoB's simplified environments to today's complex, multimodal systems tackling real-world applications. You'll discover why standard vision models fail catastrophically on GUI tasks, explore the annotation bottlenecks that make GUI datasets so expensive to create, and understand the platform fragmentation that makes "click a button" mean twenty different things across datasets. We'll dissect the most influential datasets (Mind2Web, AITW, Rico) and models that have shaped the field, examining their strengths, limitations, and the research gaps they reveal. By the end, you'll have a clear picture of where GUI agents excel, where they struggle, and, most importantly, where the opportunities lie for your own contributions. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-72-lllmstxt|> ## FiftyOne Healthcare Workshop [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/56af9060688bf339888e57b7a508690dacb75b11-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/56af9060688bf339888e57b7a508690dacb75b11-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Webinars & Workshops Healthcare Getting Started with FiftyOne for Healthcare Use Cases - July 23, 2025 This event has ended, but you can still catch up! Watch the on-demand recordings and register for our [future events.](https://voxel51.com/events) Jul 23, 2025 9:00 - 10:30 AM Pacific Online. Register for the Zoom! About this event Visual AI is revolutionizing healthcare by enabling more accurate diagnoses, streamlining medical workflows, and uncovering valuable insights across various imaging modalities. Yet, building trustworthy AI in healthcare demands more than powerful models — it requires clean, curated data, strong visualizations, and human-in-the-loop understanding. Host ![](https://cdn.sanity.io/images/h6toihm1/production/186a455951c44bed78580fe4ce87c4538bb0a0e8-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Paula Ramos Voxel51 Bio Join us for a free, 90-minute, hands-on workshop built for healthcare researchers, medical data scientists, and AI engineers working with real-world imaging data. Whether you're analyzing CT scans, radiology images, or multi-modal patient datasets, this session will equip you with the tools to design robust, transparent, and insight-driven computer vision pipelines — powered by FiftyOne, the open-source platform for Visual AI. **By the end of the workshop, you'll be able to:** - Load and organize complex medical datasets (e.g., ARCADE, DeepLesion) with FiftyOne. - Explore medical imaging data using embeddings, patches, and metadata filters. - Curate balanced datasets and fine-tune models using Ultralytics YOLOv8 for tasks like stenosis detection. - Analyze segment CT scans using MedSAM2. - Analyze results from VLMs and foundation models like MedGEMMA, NVIDIA VISTA, and NVIDIA CRADIO. - Evaluate model predictions and uncover failure cases using real-world clinical examples. **Why Attend?** This healthcare edition of our "Getting Started with FiftyOne" workshop connects foundational tools with real-world impact. Through curated datasets and clinical use cases, you'll see how to harness Visual AI responsibly, building data-centric pipelines that promote accuracy, interpretability, and trust in medical AI systems. **Prerequisites:** Basic knowledge of Python and computer vision is recommended. No prior experience in healthcare is required — just curiosity and a commitment to building meaningful AI. All participants will receive access to workshop notebooks, code examples, and extended resources to continue their journey in healthcare AI. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-73-lllmstxt|> ## Visual AI Workshop [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/1df3365d49d3baf42d81c3a983d77bd43bbd6154-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/1df3365d49d3baf42d81c3a983d77bd43bbd6154-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Webinars & Workshops Building Visual AI in the Enterprise Workshop – June 4, 2025 This event has ended, but you can still catch up! Watch the on-demand recordings and register for our [future events.](https://voxel51.com/events) Jun 4, 2025 9:00 AM to 10:30 AM Pacific Time Virtually over Zoom! About this event Want to take your computer vision workflows to the next level? Then this workshop is for you! Join us for 90 minutes as we demonstrate the full power of the enterprise FiftyOne platform so your teams can deploy visual AI applications faster, with visibility and control. Host ![](https://cdn.sanity.io/images/h6toihm1/production/6c22c5a4d06f3c1a1826318ab35446770654f8a2-1342x1343.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Nick Lotz Technical Marketing Engineer Bio ## About the Workshop Want to take your computer vision workflows to the next level? Then this workshop is for you! Join us for 90 minutes as we demonstrate the full power of the enterprise FiftyOne platform so your teams can deploy visual AI applications faster, with comprehensive visibility and control. You will learn how to: - Visualize, manage, and augment datasets no matter where they reside - Streamline development of datasets and models with out-of-the-box workflows - Collaborate easily with team members across your organization - Orchestrate and automate compute-heavy tasks directly within FiftyOne - Extend FiftyOne’s capabilities with plugins and operators - Enforce data governance and security best practices across your AI toolchain All attendees will get access to the tutorials, videos, and code examples used in the workshop. ## Prerequisites Familiarity with basic computer vision concepts and enterprise AI/ML development workflows. ## FiftyOne Teams Features ### Team Collaboration Collaborate easily and securely with your teams by sharing entire datasets, individual samples, dataset views, and even analysis results. Configure user groups and roles to control access to datasets for individuals and teams. ### Data Lens Get fast data discovery from external data sources. If you're looking to augment your dataset with specific samples, Data Lens unlocks direct access to your data source (e.g. Databricks, PostgreSQL) to search through billions of visual data within seconds. Preview samples, select interesting ones, and import them directly into FiftyOne. ### Evaluate Model Performance Execute model evaluation and analyze performance. Compare models, examine poor-performing samples, and uncover data inadequacies to improve model performance. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-74-lllmstxt|> ## FiftyOne Workshop Overview [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/f15572465b6cd96530e85d1fc3dd5f1e7f2675e7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/f15572465b6cd96530e85d1fc3dd5f1e7f2675e7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Webinars & Workshops Getting Started with FiftyOne Workshop – June 18, 2025 This event has ended, but you can still catch up! Watch the on-demand recordings and register for our [future events.](https://voxel51.com/events) Jun 18, 2025 9:00 – 10:30 AM Pacific Virtually over Zoom! About this event Join us for the free 90-minute, hands-on workshop to learn how to get started with open source FiftyOne. Host ![](https://cdn.sanity.io/images/h6toihm1/production/9a33597e3617aacb10dc014889e3c06b3716367b-320x320.png?auto=format&dpr=2&fit=max&q=75&w=96) Antonio Rueda-Toicen Voxel51 Bio ## About the Workshop Want greater visibility into the quality of your computer vision datasets and models? Then join us for this free 90-minute, hands-on workshop to learn how to leverage the open source FiftyOne computer vision toolset. At the end of the workshop you’ll be able to: - Object detection - Embeddings - Mistakenness - Deduplication This workshop will explore the importance of taking a data-centric approach to computer vision workflows. We will start with importing and exploring visual data, then move to querying and filtering. Next, we’ll look at ways to extend FiftyOne’s functionality and simplify tasks using plugins and native integrations. We’ll generate candidate ground truth labels, and then wrap things up by evaluating the results of fine tuning a foundational model. **Prerequisites:** working knowledge of Python and basic computer vision concepts. All attendees will get access to the tutorials, videos, and code examples used in the workshop [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-75-lllmstxt|> ## AI Image Segmentation Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) A Guide to AI Image Segmentation Dec 19, 2024 • 11 min read Article content In this article [Understanding Image Segmentation](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#9dd5cb24074b) [Popular Techniques for Image Segmentation](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#3f4108a63654) [Image Segmentation with FiftyOne](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#98d8da8d5a75) [Effective Segmentation Task Workflows with FiftyOne](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#ac496c17cc1f) [Conclusion](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#8fe1f62fd998) [Next Steps](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#6c233ac48d56) In this article [Understanding Image Segmentation](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#9dd5cb24074b) [Popular Techniques for Image Segmentation](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#3f4108a63654) [Image Segmentation with FiftyOne](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#98d8da8d5a75) [Effective Segmentation Task Workflows with FiftyOne](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#ac496c17cc1f) [Conclusion](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#8fe1f62fd998) [Next Steps](https://voxel51.com/blog/a-guide-to-ai-image-segmentation#6c233ac48d56) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Image segmentation is a widely used technique in Computer Vision (CV) that divides an image into more meaningful and distinguishable objects. Image segmentation is commonly used in object detection, recognition, and tracking. It finds relevance in a variety of use cases such as healthcare for medical imaging, automotive, robotics, and many others. In this guide, we’ll go through the basics of image segmentation, its benefits and techniques as well as explore effective workflows that you can easily implement in your CV work. ## Understanding Image Segmentation ### What Is Image Segmentation? Image segmentation involves assigning a specific label to each pixel in an image, resulting in a label map where every pixel corresponds to a predicted category. This classification is referred to as pixel-level classification. Semantic image segmentation is a specific type of image segmentation where each pixel is assigned a semantic class label, identifying what the pixel represents without differentiating between individual objects of the same class. Instance segmentation is another type of image segmentation that differentiates between individual instances of the same object. Consider an image example containing multiple cars. In semantic segmentation, all image pixels belonging to the car category would be labeled as a “car”. On the other hand, instance segmentation would label an individual car object with a unique label. The accuracy of image segmentation models is highly dependent on the quality of image data. Poor image quality, diverse anatomical structures, and noise present major challenges when preparing data for image segmentation models. Specialized tools like [FiftyOne](https://voxel51.com/) can help development teams develop high-performing models backed by high-quality datasets. We’ll discuss more on that topic as well as outline workflows using FiftyOne that you can implement in a bit, so read on. **Benefits and Use Cases of Image Segmentation** Image segmentation is useful in extracting detailed information about objects, shapes, and boundaries. Analyzing shapes is useful in imaging, while boundary identification helps identify edges in images. For example, in use cases such as robotics, surveillance, or self-driving cars, the object-tracking capabilities of image segmentation help follow objects of interest over a given period. Let’s explore this further across different use cases. **Scene understanding:** Image segmentation helps to categorize different regions of an image so AI systems can understand complex scenes and be more accurate in tasks such as image captioning and scene classification. **Content manipulation:** In tasks such as photo editing, image segmentation enables the enhancement of specific parts of an image without affecting the rest of the image. The most common use case we see is in augmented reality applications to overlay virtual objects onto real-world scenes. **Autonomous vehicles:** In automotive use cases such as self-driving features, image segmentation enables vehicles to identify lanes, pedestrians, obstacles, and traffic signs for safe navigation. **Robotics and automation:** Image segmentation enables robots to perform highly specialized tasks at high precision. For example, they can interact with objects and navigate effectively while avoiding obstacles. **Medical Imaging:** Image segmentation is useful in isolating and analyzing anatomical structures and tumors to aid medical professionals in the diagnosis of disease. Traditional image segmentation methods, such as thresholding, edge detection, and region-based algorithms, use specified parameters and heuristics. While they are useful for specialized tasks, they have limited adaptability and may not perform well across different or complicated images. Modern segmentation algorithms, particularly those that use deep learning, provide greater versatility and customization. These methods can be customized for a variety of applications by training models on specific datasets, allowing them to understand complicated patterns and features specific to each task. For example, in photo editing, better segmentation models may reliably extract complicated objects or regions, allowing for more precise and creative modifications. By fine-tuning models on smaller, task-specific datasets, practitioners can improve accuracy and efficiency. Overall, image segmentation is beneficial because: - It is faster and more accurate than traditional methods such as region-based segmentation, edge detection, and clustering algorithms. - It is scalable and adaptable to various domains. ## Popular Techniques for Image Segmentation Semantic segmentation and instance segmentation are two common techniques used in image segmentation. Depending on the use case and goal, you can decide which one might be most appropriate. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ### Semantic Segmentation In semantic segmentation, each pixel is assigned a class label, and all objects of the same type are given the same label. Let’s take a look at the image above that contains multiple people. With semantic segmentation, every pixel corresponding to a person is labeled identically, distinguishing them from the background and other objects as shown. This approach offers several advantages: - **Enhanced Image Understanding:** By categorizing each pixel, semantic segmentation provides a comprehensive understanding of the scene’s content, facilitating tasks like object recognition and scene interpretation. - **Improved Object Localization:** Assigning consistent labels to objects of the same class allows for precise localization within the image, which is crucial for applications such as autonomous driving and robotic navigation. - **Simplified Data Analysis:** Uniform labeling of similar objects streamlines the analysis process, making it easier to quantify and assess specific elements within an image. By applying the same label to all objects of a particular class, semantic segmentation enables machines to process and interpret visual information more effectively. ### Instance Segmentation Instance segmentation, on the other hand, assigns a unique mask to each object instance within an image, even when multiple objects belong to the same class. This approach offers several advantages: - **Precise Object Differentiation:** By generating distinct masks for each object, instance segmentation enables the identification and differentiation of individual instances, which is crucial in scenarios where understanding the exact number and location of objects is essential. - **Accurate Object Counting:** The ability to distinguish between instances allows for the precise counting of objects, which is beneficial in applications like crowd analysis, inventory management, and even wildlife monitoring. - **Enhanced Object Tracking:** In dynamic environments, such as video surveillance or autonomous driving, unique masks facilitate the tracking of specific objects over time, improving the system’s ability to monitor movements and interactions. Instance segmentation provides detailed information about each object instance, enhancing the machine’s understanding of complex scenes and leading to more informed decision-making across various applications. ## Image Segmentation with FiftyOne - [FiftyOne](https://voxel51.com/) is a tool that enables CV ML development teams to build high-performing models by getting a deeper understanding of the data, exploring and visualizing it, and analyzing model strengths and weaknesses down to the data sample level.FiftyOne supports popular image segmentation techniques and makes it possible to visualize [semantic](https://docs.voxel51.com/user_guide/using_datasets.html#semantic-segmentation) and [instance segmentation](https://docs.voxel51.com/user_guide/using_datasets.html#instance-segmentations) outputs through a powerful graphical interface.FiftyOne enhances the visualization and interpretation of image segmentation datasets and models and offers comprehensive features tailored for image segmentation workflows: - **Dataset Visualization:** FiftyOne allows users to examine datasets interactively by presenting images alongside their segmentation masks, so you can get a clear understanding of image segmentation and the discovery of patterns or discrepancies in the data. - **Model Evaluation:** With FiftyOne, you can evaluate segmentation model performance using metrics like Intersection Over Union (IoU). Users can compare predicted masks to ground truth annotations to see where the model thrives and where it needs to be improved. - **Error Analysis:** FiftyOne helps identify specific failure modes by identifying disparities between expected and actual segmentations. This targeted study is critical for refining models and improving segmentation accuracy. - **Data Curation:** FiftyOne helps to curate datasets by detecting duplicate or mislabeled images, resulting in a high-quality dataset that contributes to better segmentation models. ### Integration with Segmentation Libraries One of the features that sets FiftyOne apart from other tools is its built-in integration with [popular image segmentation models and algorithms](https://docs.voxel51.com/model_zoo/models.html) like YOLO, DINOv2, ResNet, and Meta AI’s SAM2. Check out the natively available models supported in FiftyOne and available in [FiftyOne Model Zoo](https://docs.voxel51.com/model_zoo/models.html). Click on “segmentation” to filter the Model Zoo listings. You can use these models or bring your own. ### Data Augmentation One of the main challenges of training image segmentation models is the lack of enough data. This can lead to the model overfitting on the little data that is available. Augmenting the data using FiftyOne can help to alleviate this problem. Augmenting image datasets enhances the generalization and robustness of segmentation models by exposing them to a diverse range of variations during the model training process. This process helps models perform effectively on new, unseen data. FiftyOne facilitates this by integrating with the [Albumentations](https://pypi.org/project/albumentations/) library, which offers a wide array of image augmentation techniques. Check out this [tutorial on how to augment datasets in FiftyOne with Albumentations](https://docs.voxel51.com/tutorials/data_augmentation.html). By applying transformations such as rotations, blurring, and noise addition, models learn to recognize objects under different conditions. This exposure reduces overfitting to the original dataset and improves the model’s ability to handle real-world variations. ### Version Control and Experiment Tracking With FiftyOne you can track different iterations of segmentation models along with their respective masks, giving full visibility into the development process. For example, FiftyOne has [an integration with MLflow](https://github.com/voxel51/fiftyone_mlflow_plugin) which enables you to track and register your segmentation models. It enables you to create stages for your model such as Staging, Production, and Archived. ### Integration with Evaluation Metrics FiftyOne simplifies the evaluation of segmentation models by providing tools to assess performance using metrics like Intersection over Union (IoU). IoU measures the overlap between predicted and ground truth segmentation masks, offering a pixel-level accuracy assessment. Beyond standard IoU, FiftyOne supports alternative evaluation strategies, such as focusing on boundary pixels, to provide a more nuanced understanding of model performance. These capabilities enable comprehensive analysis, helping to identify specific areas where a model excels or requires improvement. ## Effective Segmentation Task Workflows with FiftyOne ### Using Segmentation-Specific Integrations FiftyOne supports integration with custom models and libraries tailored for specific segmentation tasks. For example, with the [Segments AI integration](https://github.com/segments-ai/segments-voxel51-plugin), you can label 3D points cloud faster. The integration supports pointcloud-cuboid and pointcloud-vector. ### Extracting Features to Improve Model Accuracy Advanced ML techniques powered by [FiftyOne Brain](https://docs.voxel51.com/brain.html) enable you to extract segmentation-specific features for better performance and model analysis. For example, you can visualize model embeddings to identify patterns and clusters in your dataset that can help improve the model’s accuracy. Advanced techniques allow for identifying similar images and help analyze how the images affect performance. By examining clusters of similar images, you can detect groups where the model underperforms, indicating potential weaknesses in handling specific features or patterns. This facilitates the training of segmentation models with unique data through the computation of a uniqueness measure of an image with all the images in the dataset. ### Leveraging Active Learning for Efficient Annotation FiftyOne supports Active Learning, a strategy that makes data annotation faster by identifying the most _informative_ or _ambiguous_ examples for labeling. This is important because your segmentation model can improve even with less data. It’s made possible using FiftyOne’s [Active Learning plugin](https://github.com/jacobmarks/active-learning-plugin), which you can integrate directly into your annotation workflow. Learn more about using this plugin on the [_Supercharge Your Annotation Workflow with Active Learning_](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning/) blog post. ### Autonomous Vehicle Development Another critical application of segmentation models is autonomous vehicle development. The use of [FiftyOne in autonomous vehicles](https://voxel51.com/computer-vision-use-cases/driving/) is quite similar to its applications in medical image segmentation with a few nuances. Like the previous application, FiftyOne can be used for data preparation, visualization, active learning, and model evaluation. You can use FiftyOne to review masks for pedestrians, vehicles, and road signs. You can use active learning to label difficult images to make the model more robust. FiftyOne integrates with popular libraries such as [YOLO](https://docs.voxel51.com/tutorials/yolov8.html), TensorFlow, and PyTorch to enable the development of _real-time applications._ ### Refining Existing Segmentation Models FiftyOne allows you to refine models for object detection by improving segmentation mask accuracy through error analysis and correction. This can be done by overlying segmentation masks on images, making it faster to spot errors. You can perform error analysis to identify situations where predictions differ from the ground truth. FiftyOne supports the computation of segmentation accuracy metrics such as Intersection over Union (IoU). After computation, you can identify samples with low scores and focus on improving their score. Furthermore, you can use FiftyOne’s Brain to analyze [_mistakenness_](https://docs.voxel51.com/tutorials/detection_mistakes.html#Finding-Detection-Mistakes-with-FiftyOne) and detect misclassifications and inconsistent masks. Once you have identified challenging samples, correct them manually or use FiftyOne’s Active Learning Feature. Finally, re-train the segmentation model using the new dataset. ## Conclusion Image segmentation has made it possible to detect the object, shape, and boundary of any object within images and video datasets. Semantic segmentation and instance segmentation have found great benefits in use cases ranging from robotics, and surveillance, to medical imaging and self-driving cars. ## Next Steps FiftyOne makes it easy to achieve accurate image segmentation. You can [get started with FiftyOne](https://docs.voxel51.com/getting_started/install.html) in just a few minutes. Looking for a scalable solution for your ML team as you collaborate on visual AI projects? Check out [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) and [connect with an expert](https://voxel51.com/book-a-demo/) to see the collaborative, enterprise features of FiftyOne in action. [segmentations](https://voxel51.com/blog/tag/segmentations) [AI](https://voxel51.com/blog/tag/ai) [images](https://voxel51.com/blog/tag/images) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/ca4f84addae3f8e97daad00cf856754301cbb5e0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Are Image Segmentation Maps Superior to Bounding Boxes?\\ \\ Learn\\ \\ • \\ \\ Feb 26, 2025](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes) [![](https://cdn.sanity.io/images/h6toihm1/production/d2e24d0a14de508f8ccd36eaffd3b909c9f193b9-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Image Embeddings Transform Computer Vision Capabilities\\ \\ Learn\\ \\ • \\ \\ Nov 25, 2024](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities) [![](https://cdn.sanity.io/images/h6toihm1/production/e54b1e5afaaede40db246a7681a611856ebd8782-2260x1268.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Enhancing YOLOv8 Segmentation: Precision, Efficiency, and Robustness\\ \\ Learn\\ \\ • \\ \\ Apr 3, 2025](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-76-lllmstxt|> ## Voxel51 Press Releases [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Press [![](https://cdn.sanity.io/images/h6toihm1/production/3a033eb2f111561b2fb00f3afc1e5820293ba6f0-1200x675.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ New Data and Model Workflows from Voxel51 Accelerate Visual AI Development for Enterprises\\ \\ Press\\ \\ • \\ \\ Mar 25, 2025](https://voxel51.com/blog/new-data-and-model-workflows-from-voxel51-accelerate-visual-ai-development-for-enterprises) [![](https://cdn.sanity.io/images/h6toihm1/production/77fb65cdcbfbbc42c805d116b804179f581516bd-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 and ORI Partner to Accelerate Visual AI Innovation in U.S. Government\\ \\ Press\\ \\ • \\ \\ Feb 6, 2025](https://voxel51.com/blog/voxel51-and-ori-partner-to-accelerate-visual-ai-innovation-in-u-s-government) [![](https://cdn.sanity.io/images/h6toihm1/production/c54c9ac3306b0a424046c3b80d8d0fa8351e2b76-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 Launches FiftyOne Open Source 1.0, Accelerating the Creation of Production-Ready Visual AI Applications\\ \\ Press\\ \\ • \\ \\ Oct 1, 2024](https://voxel51.com/blog/voxel51-launches-fiftyone-open-source-1-0-accelerating-the-creation-of-production-ready-visual-ai-applications) [![](https://cdn.sanity.io/images/h6toihm1/production/c727ebde95a4d6135d0bf2740a99aa8931220c33-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 Raises $30M Series B Funding to Make Visual AI a Reality\\ \\ Press\\ \\ • \\ \\ May 16, 2024](https://voxel51.com/blog/voxel51-raises-30m-series-b-funding-to-make-visual-ai-a-reality) [![](https://cdn.sanity.io/images/h6toihm1/production/563c1d2568cdc2346048e80d18526574ae408669-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Introducing VoxelGPT: AI-Powered Computer Vision Insights Delivered Through Chat\\ \\ Press\\ \\ • \\ \\ Jun 8, 2023](https://voxel51.com/blog/introducing-voxelgpt) [![](https://cdn.sanity.io/images/h6toihm1/production/0a21c6ca2af5253f72f6b88f9d8f6dafa5fb4ab3-2560x1390.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51 & Google Collaborate to Make Downloading & Visualizing Open Images A Breeze\\ \\ Press\\ \\ • \\ \\ May 13, 2021](https://voxel51.com/blog/fiftyone-open-images-collaboration) [![](https://cdn.sanity.io/images/h6toihm1/production/62b79d1d13bfc9926b5f560049adc33fd5d2cbf3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 Launches Computer Vision Industry’s First Open-Source Rapid Dataset Experimentation Tool\\ \\ Press\\ \\ • \\ \\ Aug 12, 2020](https://voxel51.com/blog/fiftyone-open-source-launch) [![](https://cdn.sanity.io/images/h6toihm1/production/b2822f56ac527e57cfa5a5e6dbdd2fea96ae2817-1934x1110.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51’s Coronavirus Physical Distancing Index Tracks Reaction to Social Distancing Around the World\\ \\ Press\\ \\ • \\ \\ Apr 1, 2020](https://voxel51.com/blog/voxel51-physical-distancing-index) [Voxel51 Launches Image-To-Video Tool Enabling Models Trained on Images to Automatically Process Video\\ \\ Press\\ \\ • \\ \\ Oct 3, 2019](https://voxel51.com/blog/voxel51-launches-image-to-video-tool) [Voxel51 Raises $2 Million To Advance Video Understanding\\ \\ Press\\ \\ • \\ \\ Aug 8, 2019](https://voxel51.com/blog/voxel51-raises-2-million-to-advance-video-understanding) ## Enough data wrangling.
 Request a demo. [Get started](https://voxel51.com/link-catcher) [Explore the Demo](https://voxel51.com/link-catcher) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-77-lllmstxt|> ## Aidence Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/1ed4c2b8a455ca7707714826c8a8006865e5d2fd-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Aidence Aidence gives lung cancer patients a fighting chance using AI Apr 15, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Aidence](https://www.aidence.com/) turns complex data science into practical and intuitive solutions that make physicians’ work easier. The company made it their mission to provide AI that empowers healthcare and pharmaceutical professionals to deliver faster, more precise diagnostics and treatments. > "I use FiftyOne on a daily basis to improve the quality of our data and visually inspect model predictions. I much prefer the interactive experience of FiftyOne over static notebooks with matplotlib figures." – Marijn Lems, Machine Learning Engineer at Aidence ![](https://cdn.sanity.io/images/h6toihm1/production/f73ca658b4118a09323c14f0207fca8fea4b66f7-877x580.png?auto=format&dpr=2&fit=max&q=75&w=877) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-78-lllmstxt|> ## Smart Automotive Dataset Selection [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision) Why the “Annotate Everything” Era in Automotive AI Is Over Jul 24, 2025 • 4 min read Article content In this article [The cost of doing it the old way](https://voxel51.com/blog/smarter-automotive-datasets-selection#1530db0d7743) [Annotation isn’t dead—but it’s definitely evolving](https://voxel51.com/blog/smarter-automotive-datasets-selection#ea7f86651a84) [Implications for automotive AI](https://voxel51.com/blog/smarter-automotive-datasets-selection#8ca09f28fd7c) [Looking ahead](https://voxel51.com/blog/smarter-automotive-datasets-selection#b5289fb13330) In this article [The cost of doing it the old way](https://voxel51.com/blog/smarter-automotive-datasets-selection#1530db0d7743) [Annotation isn’t dead—but it’s definitely evolving](https://voxel51.com/blog/smarter-automotive-datasets-selection#ea7f86651a84) [Implications for automotive AI](https://voxel51.com/blog/smarter-automotive-datasets-selection#8ca09f28fd7c) [Looking ahead](https://voxel51.com/blog/smarter-automotive-datasets-selection#b5289fb13330) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) For years, the dominant mindset in computer vision—especially in the automotive space—has been to label everything. Teams would collect massive automotive datasets, send them off to an annotation team, and hope that enough brute-force labeling would lead to better models. That approach made sense when our tools were immature and our only lever was volume. But things have changed. In 2025, the idea that we must annotate _all_ our data to build performant AI models is not only outdated—it’s counterproductive. [The “annotate everything” era is over](https://voxel51.com/blog/the-hidden-cost-of-outsourced-data-annotation). And it’s time we talk about why. ## The cost of doing it the old way In the past, automotive teams invested millions of dollars and months of time in large-scale annotation campaigns. It was the norm. But it quickly became clear that the payoff didn’t justify the cost. Most roadway footage contains little to no signage. If you randomly sample 10% of a petabyte-scale video dataset, what you’ll likely get is hours of open highway or crowded urban streets—scenarios that are already overrepresented. What you _won’t_ get are the rare cases that actually improve model performance: speed limit signs, temporary construction warnings, edge-case intersections. By annotating everything uniformly, teams were spending an extraordinary amount of effort to reinforce patterns the model had already mastered—while missing the data that would actually help it improve. ### More isn’t always better One of the most damaging myths in computer vision is that more labeled data automatically yields better models. This is simply not true as throwing more annotations at the problem is often a waste of money. The better question is: _What data actually matters for improving perception systems performance?_ And how can we find that data more efficiently? ### The shift to smarter automotive datasets selection This question led us to develop a different kind of tooling—focused not on labeling _more_, but on identifying _what to label_. Today, with the help of semantically rich foundation models like OpenAI’s [CLIP](https://docs.voxel51.com/integrations/openclip.html), we can embed entire datasets and visualize the structure of the data before touching a single label. Using these embeddings, we can automatically select representative samples, filter for outliers, and ensure we’re building a dataset that is both diverse and targeted. Instead of blindly sampling 10% of your data, you can sample the _right_ 10%—the subset that actually improves model accuracy. This isn’t just theoretical. I’ve seen teams use this approach to achieve better performance with fewer labels, lower cost, and faster turnaround. ## Annotation isn’t dead—but it’s definitely evolving A while back, I wrote a blog post titled [Annotation is dead](https://medium.com/@jasoncorso/annotation-is-dead-1e37259f1714). That headline ruffled a few feathers—but the core point holds. Annotation isn’t disappearing, but the way we think about it must change. Thanks to foundation models, we can now auto-label up to 40% of a dataset with reasonably high accuracy. With the right QA tools, we can accept, reject, or correct those labels without manual inspection of every sample. This shift drastically reduces the time and cost of building training sets. What used to take 3–4 months and cost millions can now be done in a few hours with significantly less human effort. Recent [research](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance) has demonstrated that [verified auto-labeling](https://voxel51.com/annotation) can achieve up to 95% of human-level performance while cutting labeling costs by up to 100,000x. Models trained solely on these auto-labels often match or even exceed the performance of those trained on traditional human-labeled data, particularly for challenging edge cases where foundation models can generalize better than human annotators working at scale. ## Implications for automotive AI In the automotive industry, where the bar for accuracy is extraordinarily high, this shift is a game changer. Whether you’re building perception systems for ADAS or full autonomy, you need performance at the edge, where rare events happen. By focusing annotation efforts where they matter most, you can get to that performance faster, with lower spend and more confidence in the result. You're no longer wasting time labeling redundant data. You're strategically building better models from the start. I recently joined AI in Automotive Podcast to discuss this in detail, [listen to the full episode](https://www.ai-in-automotive.com/aiia/504/jasoncorso?utm_campaign=Podcast&utm_medium=organic-social&utm_source=LinkedIn) if you’re dealing with video datasets that seem to have a mind of their own. ## Looking ahead At Voxel51, we’ve spent the last several years building tooling that reflects this new philosophy: putting data at the center of visual AI. I believe that in the next few years, we’ll see fewer teams asking “how many annotations do we need?” and more teams asking “how much of this work can we _avoid_?” This isn’t just a technical shift—it’s a strategic one. The annotate-everything era was a product of its time. But we’ve outgrown it. If we’re serious about deploying robust perception systems and real-world AI, we need smarter data pipelines—not bigger labeling budgets. The future of automotive datasets lies in intelligent selection, not exhaustive annotation. [autonomous vehicles](https://voxel51.com/blog/tag/autonomous-vehicles) [labeling](https://voxel51.com/blog/tag/labeling) [Visual AI](https://voxel51.com/blog/tag/visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/ad9fb967c5455e0f763411fb81956767d7f26482-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Jason Corso Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/db7358784a18a2ffa365704e7d941e73fbdf1fcd-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 1\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1) [![](https://cdn.sanity.io/images/h6toihm1/production/65dd49021e60c3c8d3f0982f888781871ba4ed26-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ From Prototype to Production: What it Really Takes to Deploy Computer Vision in Manufacturing\\ \\ Computer Vision\\ \\ • \\ \\ Aug 6, 2025](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing) [![](https://cdn.sanity.io/images/h6toihm1/production/c63b546423b6cec32ccccc7df6bc4e0fced0b1a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 2\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-79-lllmstxt|> ## Kitro Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/4c3b8e9d7736050517b8c81baf652a0566d6c7a2-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Kitro Kitro uses FiftyOne to train models to reduce food waste Apr 17, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [KITRO](https://www.kitro.ch/) is a Swiss company on a mission to harness the power of technology and use it for sustainable change. With artificial intelligence as the foundation, KITRO offers an automated food waste data collection and analysis solution that can be adopted by food and beverage outlets worldwide. > "FiftyOne is a key ingredient for Kitro in developing ML models and curating dataset for training our models. We use it extensively for data curation and it helps us a lot in investigating the mistakes our pipeline makes. The similarity search feature and looking at the data from the embedding view has helped us make changes in the dataset we couldn't otherwise see. We are also more and more integrating FiftyOne in the automated models testing. It's really a "Swiss army knife" for working with data." – Luka Posilović, Head of Machine Learning at KITRO ![](https://cdn.sanity.io/images/h6toihm1/production/d0248e1b223ac37daffa7f3aec2a4d919080dd31-800x800.png?auto=format&dpr=2&fit=max&q=75&w=800) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-80-lllmstxt|> ## Pricing Plans Overview [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # Simple plans for powerful data workflows [Team\\ \\ Streamline your team's visual AI development.\\ \\ - 8 user seats, 16 guest seats\\ \\ - 4 VPUs, 2,800 hours/month of compute\\ \\ - 1 production deployment\\ \\ - 1 dev/staging deployments\\ \\ \\ - Unlimited data\\ \\ - Unlimited model inference\\ \\ - SSO\\ \\ - Standard enterprise support\\ \\ \\ Contact sales](https://voxel51.com/pricing#contact) [Growth\\ \\ Scale your visual AI stack across projects.\\ \\ - 25 user seats, 100 guest seats\\ \\ - 20 VPUs, 14,000 hours/month of compute\\ \\ - 3 production deployments\\ \\ - 3 dev/staging deployments\\ \\ \\ - Everything in Starter, plus:\\ \\ - On-premise or air-gapped deployments\\ \\ - Premium enterprise support\\ \\ - Dedicated customer success engineer\\ \\ \\ Contact sales](https://voxel51.com/pricing#contact) [Custom\\ \\ Get visual AI into production across your organization.\\ \\ - Unlimited user seats, unlimited guest seats\\ \\ - Unlimited VPUs\\ \\ - Unlimited production deployments\\ \\ - Unlimited dev/staging deployments\\ \\ \\ - Everything in Growth, plus:\\ \\ - Dedicated customer success team, including engineering, solutions, and architect resources\\ \\ - Professional services\\ \\ \\ Get custom pricing](https://voxel51.com/pricing#contact) Compare features Team Growth Custom Compare features Team Growth Custom Overview Data volume Unlimited Unlimited Unlimited Model inference Unlimited Unlimited Unlimited User seats 8 25 Unlimited Guest seats 16 100 Unlimited VPUs 4 20 Unlimited Hours of compute 2,800 14,000 Unlimited Overview Data volume Unlimited Model inference Unlimited User seats 8 Guest seats 16 VPUs 4 Hours of compute 2,800 Deployment Production environments 1 3 Unlimited Staging environments 1 3 Unlimited Cloud (public, private, hybrid) On-premise Add-on Air-gapped Add-on Deployment Production environments 1 Staging environments 1 Cloud (public, private, hybrid) On-premise Add-on Air-gapped Add-on Data governance & security Single Sign-On (SSO) User and role-based access controls ISO 27001 certifications Support for Personally Identifiable Information (PII) Support for Protected Health Information (PHI) Add-on Secrets management Dataset versions Unlimited Unlimited Unlimited Audit logs Data encrypted at rest and in transit Data governance & security Single Sign-On (SSO) User and role-based access controls ISO 27001 certifications Support for Personally Identifiable Information (PII) Support for Protected Health Information (PHI) Add-on Secrets management Dataset versions Unlimited Audit logs Data encrypted at rest and in transit Customization & extensibility Integrations \- Annotation tools \- Cloud storage \- SDK & notebooks \- Vector search databases \- Models \- Experiment tracking \- Datasets Front-end customization Plugins framework for build-your-own capabilities Custom dashboards and analytics Customization & extensibility Integrations \- Annotation tools \- Cloud storage \- SDK & notebooks \- Vector search databases \- Models \- Experiment tracking \- Datasets Front-end customization Plugins framework for build-your-own capabilities Custom dashboards and analytics Dataset support Data volume Unlimited Unlimited Unlimited Data points at top performance Millions Tens of millions Billions Images Video 3D (point cloud, mesh) DICOM, NIfTI Audio Geospatial data Multimodal grouped datasets [Custom media types](https://docs.voxel51.com/user_guide/using_datasets.html#media-type) Dataset support Data volume Unlimited Data points at top performance Millions Images Video 3D (point cloud, mesh) DICOM, NIfTI Audio Geospatial data Multimodal grouped datasets [Custom media types](https://docs.voxel51.com/user_guide/using_datasets.html#media-type) Computer vision tasks Segmentation (instance, semantic) Classification Object detection Computer vision tasks Segmentation (instance, semantic) Classification Object detection Annotation types Polylines & polygons Keypoints Bounding boxes 3D (cuboids, arbitrary geometries) Annotation types Polylines & polygons Keypoints Bounding boxes 3D (cuboids, arbitrary geometries) Support Email support Onboarding support Solutions architect Custom MSA & SLA Support Email support Onboarding support Solutions architect Custom MSA & SLA Annotation Model & label prediction import Zero-shot prediction Active learning Auto-labeling Add-on Add-on \- Automated QA workflows Add-on Add-on \- Smart ranking for labels Add-on Add-on \- Human-in-the-loop workflows Add-on Add-on Computer vision tasks \- Segmentation (instance, semantic) \- Classification \- Object detection Annotation Model & label prediction import Zero-shot prediction Active learning Auto-labeling Add-on \- Automated QA workflows Add-on \- Smart ranking for labels Add-on \- Human-in-the-loop workflows Add-on Computer vision tasks \- Segmentation (instance, semantic) \- Classification \- Object detection Data management & curation Data exploration \- Dynamic data lake retrieval \- Dataset slicing, querying, filtering \- Natural language search \- Similarity search \- Metadata management Data visualization \- Interactive data visualization \- Embeddings \- Dashboard and analytics Data quality \- Outlier & anomaly detection \- Data issue detection \- Automated data quality scoring Customization \- Custom metadata fields & schema \- Custom workflows \- Custom data quality metrics \- Custom predictions \- Custom embeddings \- Custom dashboards Data management & curation Data exploration \- Dynamic data lake retrieval \- Dataset slicing, querying, filtering \- Natural language search \- Similarity search \- Metadata management Data visualization \- Interactive data visualization \- Embeddings \- Dashboard and analytics Data quality \- Outlier & anomaly detection \- Data issue detection \- Automated data quality scoring Customization \- Custom metadata fields & schema \- Custom workflows \- Custom data quality metrics \- Custom predictions \- Custom embeddings \- Custom dashboards Model evaluation Model comparison Scenario evaluation Dataset and model versioning Aggregate model analytics \- Metrics: Precision, recall, accuracy, F1 \- Confusion matrices \- False positives, negative \- Confidence Sample-level analysis \- Prediction vs. ground truth Model evaluation Model comparison Scenario evaluation Dataset and model versioning Aggregate model analytics \- Metrics: Precision, recall, accuracy, F1 \- Confusion matrices \- False positives, negative \- Confidence Sample-level analysis \- Prediction vs. ground truth Show all features > “From vehicle safety and autonomy to security systems to robotics, Bosch is a leader in artificial intelligence solutions utilizing computer vision. Voxel51’s solutions help us organize, evaluate and refine our data and models, enabling us to develop robust, reliable AI applications across multiple teams and projects. ” > > **Arvind Kumar Shekar** > > Lead Expert AI Validation, Bosch ![](https://cdn.sanity.io/images/h6toihm1/production/0d38aded4af6b111c971ef12cfeb409c9f2b7b44-324x72.png?auto=format&dpr=2&fit=max&q=75&w=100) > “The biggest benefit of FiftyOne has been the speed of development. What used to take weeks or even months can now be done in days, with fewer people and a 7% increase in model performance. It’s freed up our team to focus on what they do best while accelerating our computer vision pipeline.” > > **Kermal Eren** > > Lead Computer Vision Engineer, Ancera ![](https://cdn.sanity.io/images/h6toihm1/production/17832cc5c128034561711a5d43942252792462b4-479x101.webp?auto=format&dpr=2&fit=max&q=75&rect=0,4,479,95&w=100) > “As we dive into the development of Florence-5B, we’re relying on FiftyOne more than ever. The tool’s intuitive interface and rich feature set are essential for effectively managing our large datasets and gaining critical insights.” > > **Bin Xiao** > > AI Researcher, Florence Visual Language Model ![](https://cdn.sanity.io/images/h6toihm1/production/dbee9b7f84a55058499c98caa6944bb96f4e64f7-487x103.png?auto=format&dpr=2&fit=max&q=75&w=100) ## Questions? We have answers. ### What is a VPU? ### How much compute does one VPU provide? ## Not sure what plan to pick? Get in touch and our team will get back to you. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-81-lllmstxt|> ## Enhancing Retail Customer Experience [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) Using Computer Vision to Enhance Customer Experience in Retail Jan 23, 2025 • 10 min read Article content In this article [Using Computer Vision to Enhance Customer Experience in Retail](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#2b26f4209cca) [The Power and Potential of Computer Vision in Retail](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#797f3dc45cc0) [Transforming Customer Experience with Computer Vision: Real-World Applications](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#27d2935c2930) [Challenges and Considerations for Implementing Computer Vision in Retail](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#f48fae35159b) [Building Customer Experience Solutions for Retail Using FiftyOne](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#9774e464ec68) [The Future of Computer Vision in the Retail Customer Experience](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#0fd17348d600) In this article [Using Computer Vision to Enhance Customer Experience in Retail](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#2b26f4209cca) [The Power and Potential of Computer Vision in Retail](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#797f3dc45cc0) [Transforming Customer Experience with Computer Vision: Real-World Applications](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#27d2935c2930) [Challenges and Considerations for Implementing Computer Vision in Retail](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#f48fae35159b) [Building Customer Experience Solutions for Retail Using FiftyOne](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#9774e464ec68) [The Future of Computer Vision in the Retail Customer Experience](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail#0fd17348d600) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) # Using Computer Vision to Enhance Customer Experience in Retail Although the retail industry remains a pillar of the global economy, with the global retail sector reaching [$22 trillion in sales in 2023](https://www.forrester.com/blogs/global-retail-e-commerce-sales-will-reach-6-8-trillion-by-2028), it continues to face transformative disruptions. From the continuing growth and evolution of ecommerce to the long-term impacts of the COVID pandemic, the influence of social media and more, retailers need to adapt and evolve. Throughout any challenge or disruption, the ability to deliver superior customer experiences is critical to success–according to Zendesk’s [Customer Experience Trends Report 2024](https://www.zendesk.com/blog/customer-expectations-meet-rising-demands/), 57% of consumers say they would switch to a competitor following a single bad experience. As retailers look for new ways to innovate the customer experience, computer vision and artificial intelligence (AI) are emerging as transformative technologies with the potential to enhance customer experience to the benefit of both retailers and their customers. Across both physical stores and online shopping, leading retailers are turning to solutions powered by computer vision to deliver more personalized, engaging, and efficient experiences for their customers. ## **The Power and Potential of Computer Vision in Retail** Computer vision is a specialized field within AI that analyzes and interprets visual data with machine learning algorithms. Using computer vision, AI applications interpret and analyze visual information to provide insights and help make decisions based on that data. Computer vision includes core functionality essential to building applications for retail, such as: \[bullet-checkmarks class="one-column"\] - **Object detection:** computer vision algorithms can analyze picture, video, and imaging data collected from security cameras and sensors to identify objects, whether people or carts or items on store shelves. - **Object classification:** further analysis of the objects detected in visual data can be used to identify which people are shoppers vs. employees, which items were taken from a shelf, or damaged items. - **Object tracking:** computer vision algorithms can also track the movement of individual objects, from tracking where people walk in a store to what a shopper does with an item after picking it up from the shelf. - **Behavior analysis:** computer vision also makes it possible to analyze customer behavior and paths through a store, to distinguish dwell times assessing a product from time spent checking one’s phone, and to assess wait times at checkout. \[/bullet-checkmarks\] Retail is ideally suited for applications that utilize computer vision and AI (also referred to as “ [visual AI](https://voxel51.com/blog/what-is-visual-ai-going-beyond-computer-vision/)”). Visual information plays a critical role in consumer behavior and experience, from how shoppers navigate stores to how they assess the attractiveness and value of a product. There is a range of camera and sensor solutions designed for retail, increasingly built for use with computer vision, that capture and process picture, video, and imaging data from stores, warehouses, and delivery vehicles. The recent surge of investment in AI has led to rapid advances in computer vision, and has made those advances available to a broader range of organizations and applications than ever before. Not only is computer vision enabling innovations in customer experience, but there are also [transformative applications in other areas of retail](https://voxel51.com/blog/how-computer-vision-is-changing-retail/), from inventory management to supply chain management, loss prevention, and more. ## **Transforming Customer Experience with Computer Vision: Real-World Applications** Let’s take a closer look at a few examples of ways that retailers can use computer vision and visual AI to improve customer experience, both in-store and online. **Just-in-Time Customer Assistance** We’ve all had the experience of wandering up and down aisles of a store, frustrated by being unable to find a specific item or unable to find someone who can help us with an item that’s hard to reach or locked away. Computer vision technology and analysis can help, processing data from video cameras to identify shoppers who appear to be circling looking for something or struggling to reach an item on a high shelf, and notify staff so that an associate can be routed to offer the customer assistance. Particularly in large footprint stores, that can dramatically reduce customer frustration while also ensuring that staff are able to provide more timely and efficient service. **Store Layout** Computer vision technology also helps retailers organize their stores and aisles in ways that improve the customer experience. By analyzing customer movements throughout the store and correlating that information with customer shopping baskets, retailers can improve store layout to reduce the amount of time that customers spend walking to different parts of the store, making customers’ shopping more efficient. They can also use computer vision algorithms for object and pose detection to identify items that are most frequently difficult for customers to reach so that those items can be placed on more accessible shelves. **Virtual Try-Ons** It’s not only physical retail stores where computer vision can give customers a better experience. One of the challenges of online retail is not having the ability to try out a product–to see how that jacket looks on you, or if that pair of glasses works well on your face, or if that coffee table would complement your sofa. Using computer vision, retailers can now create a virtual try-on experience that online shoppers can use anywhere, including the comfort of their own homes. These augmented reality (AR) applications use computer vision algorithms such as pose estimation, image segmentation, and 3D modeling to analyze images and video from the user’s phone or computer to provide a realistic virtual representation of how a product would look in real life. For example, L’Oreal provides a [virtual makeup try-on](https://www.lorealparisusa.com/virtual-try-on-makeup) so that online customers can see how different lipstick shades and brands might look on their face. Offering virtual try-ons provides an improved customer experience for online shoppers. That increases online engagement, reduces returns, and offers a competitive edge in the ecommerce landscape. ### **Streamlined and Automated Checkout** Extended time waiting in the checkout line or dealing with errors during checkout has a significant impact on the customer experience–according to one study, [77% of customers](https://www.retailcustomerexperience.com/blogs/enhancing-customer-experience-through-faster-checkouts/) said that they would avoid returning to a store where they previously had a long wait at checkout.Rather than relying on staff to notice long lines and request help, retailers can use computer vision-powered cameras to continuously monitor checkout lines and automatically request additional staff and checkout lanes. In addition, real-time analysis of in-store foot traffic can be used to predict when checkout is going to become busy so that staffing levels at checkout can be proactively adjusted. Retailers are also beginning to deploy solutions that streamline and automate self-checkout with computer vision technology. We’ve all experienced the frustration of extended waits while other shoppers struggle with self-checkout systems. Self-checkout areas using computer vision cameras can help by monitoring and guiding shoppers through the process. More advanced systems can [instantly recognize and tally products](https://www.forbes.com/sites/laurendebter/2022/06/02/coming-to-stores-a-new-type-of-self-checkout-where-no-scanning-is-required/), eliminating the need to individually scan items. Shoppers can simply place their items in a designated area and the system automatically calculates the total. Not only do these systems facilitate a seamless and rapid checkout experience, they also help prevent accidental or intentional failure to scan all items that leave the store. For retailers, streamlined self-checkout increases accuracy and efficiency while reducing labor costs. For shoppers, contactless checkout systems offer a smooth grab-and-go shopping experience, reducing the time spent at checkout counters and enhancing overall convenience. ## **Challenges and Considerations for Implementing Computer Vision in Retail** Computer vision solutions can be transformative for the retail customer experience, but that doesn’t mean that there are not challenges and considerations that need to be taken into account when evaluating and adopting them. **Data privacy and security** are important to both consumers and regulators, with implications for computer vision systems and for how data captured and used by those systems is stored and managed. Privacy regulations, including the California Consumer Privacy Act (CCPA), [require explicit consent](https://www.clarip.com/data-privacy/ccpa-biometric-information/) from the consumer to allow use of facial recognition, and personal data protection and privacy regulations, including the [General Data Protection Regulation](https://eugdpr.org/) (GDPR), also apply to image and video data. When planning to deploy computer vision systems, retailers need to understand these requirements and chose vendors and partners who can help meet them. **Integration** with other systems is also crucial to success. Although some vendors offer fully-integrated solutions spanning from hardware to software and services, in many cases retailers can benefit from combining pieces from multiple vendors. When selecting tools for a computer vision stack, look for compatible data formats, storage options, and security measures to simplify deployment and make scaling across different projects and stores easier. **Data quality** is critical to the success of computer vision solutions. Without quality data, computer vision models and algorithms can fail to correctly identify and analyze objects and patterns. Video and image data collected in retail operations can be particularly challenging due to variable lighting conditions, dusty or smeared camera lenses, and placements that result in high and wide camera angles. It’s not just a question of having enough quantity and variety of real-world data to train your computer vision models, it’s also critically important to have tools that allow you to understand and manage that data to understand and address the edge cases, outliers, gaps, and scenarios that are challenging for your computer vision models. For example, if training data doesn’t include video samples from dirty cameras in poor lighting, computer vision models will struggle to accurately identify shoppers and their activity in similar scenarios. ## **Building Customer Experience Solutions for Retail Using FiftyOne** When it comes to building computer vision solutions tailored to deliver enhanced customer experiences in the retail sector, being able to understand and intelligently use visual data at every step of development is critical to success. The explosion in quantity and variety of visual data has created both opportunities and challenges for builders creating retail solutions. More data helps train algorithms to understand a greater variety of scenarios, but also exponentially increases the complexity of managing data and putting it to use effectively and efficiently, not to mention addressing privacy and security considerations. Those challenges are only magnified if approached with handbuilt patchworks of fragmented tools. [FiftyOne](https://voxel51.com/) from Voxel51 is a solution designed from the start to make it easy for computer vision developers to understand, refine, and use data throughout development. Powered by open source, FiftyOne helps builders: \[bullet-checkmarks class="one-column"\] - **Explore, visualize, and understand visual data and find valuable insights in it** - **Improve data quality and make it possible to curate optimal datasets for development use** - **Use data to evaluate and refine computer vision models to improve capability and accuracy** \[/bullet-checkmarks\] Whether used by retailers building computer vision solutions customized for their specific needs or by suppliers creating solutions for retailers, FiftyOne’s flexibility and extensibility enable builders to simplify and accelerate development of robust and reliable solutions. FiftyOne makes it easy for builders to leverage the latest advances in algorithms and models with their data, and integrates easily with other critical solutions including those for labeling and annotation, experiment tracking, and more. FiftyOne also helps ensure that data security and privacy requirements are met. Unlike handbuilt solutions, FiftyOne provides built-in access control and auditing, whether deployed on-premises or in the cloud, so that you can control how data is accessed and used. ## **The Future of Computer Vision in the Retail Customer Experience** New examples of ways to use computer vision to enhance the retail customer experience continue to emerge as adoption grows. Integration with other data sources creates possibilities for even more personalized customer experiences, for example enhancing virtual try-on experiences with visual suggestions for complementary accessories. It also creates new opportunities for targeted promotions, such as noticing increased traffic for certain products and delivering real-time promotions via digital signage to convert a surge in interest to shopping baskets and sales. As illustrated by these examples, computer vision is a transformative technology for the retail customer experience. By enabling personalized, efficient, and engaging experiences for shoppers across brick and mortar stores and digital retail, computer vision contributes to critical customer satisfaction and retention priorities. Powered by FiftyOne, AI builders are creating new and innovative solutions that computer vision makes possible. [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [retail](https://voxel51.com/blog/tag/retail) [retail use case](https://voxel51.com/blog/tag/retail-use-case) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Loading related posts... [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-82-lllmstxt|> ## Indian Institute of IT [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/5c5a66794fd02f9b3f3594f57be644e9f7f68ea7-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Indian Institute of IT & Management Cybersecurity research at the Indian Institute of IT & Management runs on FiftyOne Apr 4, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) The [Indian Institute of Information Technology and Management](http://www.iiitm.ac.in/index.php/en/) is part of a group of prestigious public technical institutes located across India. Known for excellence in education, they are under the ownership of the Ministry of Education of the Government of India. > "FiftyOne was integrated into a research project that required turning tabular data concerning zero-day attacks into image data for an intrusion detection system’s machine learning model. The goal of the research was to detect and minimize the intrusions occurring over the network through the use of machine learning models." – Shivam Agrawal, BTech + MTech at Information Technology ![](https://cdn.sanity.io/images/h6toihm1/production/d1bd615a0e27af7ad33f62584fba50a7dc365702-1079x667.png?auto=format&dpr=2&fit=max&q=75&w=1079) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-83-lllmstxt|> ## Best of CVPR 2025 Day 3 [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision) The Best of CVPR 2025 Series – Day 3 May 29, 2025 • 10 min read Article content In this article [Pushing AI to See, Segment, Detect, and Predict with Precision](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#e02a28e49534) [Making Every Pixel Count \[1\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#64a599d989b6) [OpenMIBOOD: Raising the Bar for Medical OOD Detection \[2\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#cc0945881b9a) [DyCON: Smarter Segmentation Under Uncertainty \[3\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#199ce5c7f0cd) [RANGE: Smarter Geo-Embeddings from Fewer Images \[4\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#2f6581966915) [Why These Papers Matter — A New Era in Computer Vision](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#a29d5ec8b40c) [What is next?](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#347485ae6c54) [References](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#aa3cb7970f00) In this article [Pushing AI to See, Segment, Detect, and Predict with Precision](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#e02a28e49534) [Making Every Pixel Count \[1\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#64a599d989b6) [OpenMIBOOD: Raising the Bar for Medical OOD Detection \[2\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#cc0945881b9a) [DyCON: Smarter Segmentation Under Uncertainty \[3\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#199ce5c7f0cd) [RANGE: Smarter Geo-Embeddings from Fewer Images \[4\]](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#2f6581966915) [Why These Papers Matter — A New Era in Computer Vision](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#a29d5ec8b40c) [What is next?](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#347485ae6c54) [References](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3#aa3cb7970f00) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Pushing AI to See, Segment, Detect, and Predict with Precision As we wrap up Day 3 of our Best of CVPR series, we spotlight four groundbreaking papers that challenge conventional boundaries in vision research. From detailed vision-language alignment to robust anomaly detection, adaptive medical segmentation, and scalable geospatial prediction, each work dives deep into precision, context, and adaptability. Register for the [virtual meetup](https://voxel51.com/events/best-of-cvpr-july-11-2025) to dive deeper. These papers improve model metrics and rethink how vision models interact with complexity, ambiguity, and real-world constraints. Whether tuning into subtle textual cues in an image or interpreting a brain scan under data scarcity, these contributions push forward what it means for AI to understand. Check out [Day 1](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1) and [Day 2](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2) blog posts if you haven't already! ![](https://cdn.sanity.io/images/h6toihm1/production/0220c664325f1b00d91fd65d91ad1155408e9ef8-1400x522.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ## Making Every Pixel Count \[1\] **_Paper title:_** _[FLAIR: VLM with Fine-grained Language-informed Image Representations](https://arxiv.org/abs/2412.03561)._ **_Paper Authors:_** _Rui Xiao, Sanghwan Kim, Mariana-Iuliana Georgescu, Zeynep Akata, Stephan Alaniz._ **_Institutions:_** _Technical University of Munich, Helmholtz Munich, MCML, MDSI_ ![](https://cdn.sanity.io/images/h6toihm1/production/84ae72d14ed2ecaee77688fb2e3faa0da5004231-866x828.webp?auto=format&dpr=2&fit=max&q=75&w=866) CLIP models are powerful at connecting images and text, but only globally. What if we could push this alignment to the fine details of an image, allowing AI to understand “a photo of a drink” and “a blurred parking lot in the background filled with cars”? ### What It’s About: The paper presents FLAIR (Fine-grained Language-informed Image Representations), a new vision-language model that enhances the fine-grained alignment between image regions and textual descriptions. Unlike CLIP, which aligns image-text pairs at the global level, FLAIR learns localized, token-level embeddings by leveraging detailed sub-captions about specific image features. ### Why It Matters: Current models like CLIP fail to pick up subtle yet important image-text associations, limiting performance in tasks that require partial image content understanding — such as localized retrieval or segmentation. FLAIR improves the granularity of visual-textual understanding, which is vital in real-world applications like medical imaging, robotics, and surveillance. ### How It Works: - FLAIR samples diverse, fine-grained sub-captions for each image, describing detailed visual elements. - It introduces a text-conditioned attention pooling mechanism over local image tokens, creating token-level embeddings that capture both global and fine-grained semantics. - Trained on 30M image-text pairs, FLAIR learns to match each text token with the relevant image region. - The model is evaluated on standard benchmarks and a newly proposed fine-grained retrieval task. The visual example in Figure 1 shows FLAIR’s superiority in token-level grounding: while other models like DreamLIP-30M and OpenCLIP-1B struggle to highlight relevant details (e.g., “cars” in the background or “frappuccino” in the foreground), FLAIR correctly aligns image regions with their corresponding textual cues, even when context is blurred. ### Key Result: FLAIR achieves state-of-the-art results on existing multimodal retrieval benchmarks and the proposed fine-grained retrieval task. Notably, it performs well even in zero-shot segmentation, surpassing CLIP-based models trained on larger datasets. ### Broader Impact: FLAIR redefines how vision-language models interpret image-text relationships, pushing beyond global embeddings toward localized semantic understanding. It enhances transparency, precision, and context-awareness in downstream applications, crucial for systems interacting with complex visual environments. ## OpenMIBOOD: Raising the Bar for Medical OOD Detection \[2\] **_Paper title:_** _OpenMIBOOD:_ _[Open Medical Imaging Benchmarks for Out-Of-Distribution Detection](https://arxiv.org/abs/2503.16247)._ **_Paper Authors:_** _Max Gutbrod, David Rauber, Danilo Weber Nunes, Christoph Palm._ **_Institutions:_** _Regensburg Medical Image Computing (ReMIC), OTH Regensburg, Regensburg Center of Health Sciences and Technology (RCHST), OTH Regensburg, Germany_ ![](https://cdn.sanity.io/images/h6toihm1/production/4d729d17c273e303771fa8f38973baad734f9546-1400x665.webp?auto=format&dpr=2&fit=max&q=75&w=1400) AI in healthcare must be reliable, even when it encounters the unexpected. OpenMIBOOD introduces a comprehensive benchmark suite to test and improve models’ detection of medical inputs that fall outside their training distribution. ### What It’s About: This paper presents OpenMIBOOD, the first standardized benchmark suite designed to evaluate Out-of-Distribution (OOD) detection in medical imaging. It comprises 14 datasets spanning multiple imaging modalities and scenarios, classified as: - In-distribution (ID) - Covariate-shifted ID (cs-ID) - Near-OOD - Far-OOD OpenMIBOOD evaluates 24 OOD detection methods, including post-hoc approaches, and reveals their limitations when applied to medical data. ### Why It Matters: While existing OOD detection benchmarks (like OpenOOD) focus on natural images, they fail to generalize to the high-stakes, low-variance world of medical data. This misalignment can lead to critical safety risks when AI systems face unexpected patient inputs. OpenMIBOOD fills this gap, offering a realistic and reproducible benchmark for healthcare AI development. ### How It Works: - Builds on the OpenOOD taxonomy and adapts it to medical imaging. - Introduces a multi-domain benchmark with a clear distinction between cs-ID, near-OOD, and far-OOD based on semantic and contextual differences. - Divides all datasets into validation and test sets for robust evaluation. - Evaluates and compares OOD detection methods based on AUROC, AUPR, and FPR@95 across categories like MIDOG, PhaKIR, and OASIS-3. ### Key Result: - Feature-based OOD methods (e.g. ViM) consistently outperform logit/probability-based methods in medical domains. - However, no single method works across all domains, and models optimized for natural images often fail in medical contexts. - Highlights the need for domain-specific solutions tailored to the statistical characteristics of medical imaging data. ### Broader Impact: OpenMIBOOD lays the foundation for trustworthy, safety-aware AI in medicine by offering an open, extensible benchmark. It challenges the assumption that natural-image benchmarks are sufficient and instead calls for purpose-built evaluation tools in clinical AI. This can directly inform the development of regulatory-grade models in real-world healthcare settings. ## DyCON: Smarter Segmentation Under Uncertainty \[3\] **_Paper title:_** _DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation._ **_Paper Authors:_** _Maregu Assefa, Muzammal Naseer, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Mohamed L Seghier, Naoufel Werghi._ **_Institutions:_** _Center for Cyber-Physical Systems (C2PS), Khalifa University of Science and Technology, Abu Dhabi, UAE_ ![](https://cdn.sanity.io/images/h6toihm1/production/233843e8cff0a856e8f53f28fd2ba6be7c2b58d2-1400x720.webp?auto=format&dpr=2&fit=max&q=75&w=1400) Medical image segmentation often struggles when labeled data is scarce and pathology is complex. DyCON takes a dynamic approach — teaching models when and where to trust their predictions, especially in the hardest regions. ### What It’s About: This paper introduces DyCON, a semi-supervised learning framework for 3D medical image segmentation that tackles two common problems in clinical settings: uncertainty in lesion boundaries and class imbalance. DyCON integrates two key modules into any consistency-learning-based framework: - Uncertainty-aware Consistency Loss (UnCL) - Focal Entropy-aware Contrastive Loss (FeCL) ### Why It Matters: In clinical practice, annotation is expensive, and lesions are often small, irregular, or difficult to distinguish. Standard semi-supervised methods discard uncertain voxels or treat all regions equally, causing segmentation to fail when precision is most needed. DyCON avoids this by using uncertainty as a signal, not a weakness. ### How It Works: - UnCL: Dynamically adjusts the consistency loss based on voxel-level uncertainty (entropy). Early in training, it prioritizes learning from uncertain voxels; later, it focuses on refining confident areas. - FeCL: This method applies contrastive learning with dual focal weights and entropy-aware adjustments, giving more weight to hard positives and hard negatives (e.g., visually similar but different regions). It also includes top-k hard negative mining to improve feature discrimination. - Built into a Mean-Teacher framework with 3D U-Net backbones and an ASPP-based projection head for feature embedding. ### Key Result: Across four medical datasets (ISLES’22, BraTS’19, LA, Pancreas CT), DyCON outperforms all state-of-the-art semi-supervised segmentation methods, particularly: - +11.6% Dice on ISLES’22 with only 10% labeled data (61.48% → 73.73%) - 88.75% Dice on BraTS’19 with 20% labels (↑ from 86.63%) - Better precision on small and scattered lesions, where other models produce false positives or miss targets ### Broader Impact: DyCON enables reliable lesion segmentation with minimal annotation, improving diagnostic AI in stroke, tumor, and organ segmentation. By integrating uncertainty into both global and local learning, it improves model trust and interpretability — key steps toward deployable medical AI in clinical workflows. ## RANGE: Smarter Geo-Embeddings from Fewer Images \[4\] **_Paper title:_** _[RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings](https://arxiv.org/pdf/2502.19781)._**_Paper Author:_** _Aayush Dhakal, Srikumar Sastry, Subash Khanal, Adeel Ahmad, Eric Xing, Nathan Jacobs._ **_Institutions:_** _Washington University in St. Louis; Taylor Geospatial Institute._ ![](https://cdn.sanity.io/images/h6toihm1/production/18c06f713324aa56faaed614f5948fe00993d90c-878x726.webp?auto=format&dpr=2&fit=max&q=75&w=878) Modern geospatial AI relies on location-image alignment, but what if your model could predict what a place looks like — without seeing it? RANGE taps into the power of retrieval to approximate rich visual features and boost geolocation-based predictions. ### What It’s About: The paper introduces RANGE, a framework for generating multi-resolution geo-embeddings by augmenting location representations with retrieved visual features. While methods like SatCLIP align images and geolocations contrastively, they discard high-frequency image-specific details. RANGE addresses this by retrieving relevant image features from a compact database using both semantic and spatial similarity, then injecting those into the location embedding. ### Why It Matters: Many geospatial tasks — like species classification, climate estimation, and land use prediction — depend on subtle image features at specific locations. Traditional contrastive models miss this detail, limiting downstream performance. RANGE recovers the lost signal efficiently, without storing or processing massive satellite imagery in real-time. ### How It Works: - Uses a pre-trained contrastive model (e.g., SatCLIP) to align satellite images with locations. - Builds a retrieval database of: - Low-resolution embeddings (shared image-location info) - High-resolution embeddings (pure image features via SatMAE) - At inference, retrieves a weighted combination of high-res features using cosine similarity and concatenates it with the original geo-embedding. - RANGE+ introduces β-controlled spatial smoothing, blending semantic and spatial similarity to generate embeddings at variable frequencies. ### Key Result: RANGE and RANGE+ achieve state-of-the-art performance on 7 geospatial tasks, with up to: - +21.5% accuracy in biome classification - +0.185 R² gain in population density prediction - 0.896 R² average across 8 climate variables from the ERA5 dataset They also outperform other methods in fine-grained species classification using iNaturalist 2018 (Top-1: 75.2%). ### Broader Impact: RANGE enables scalable, accurate geospatial inference without real-time access to satellite imagery. It enhances applications in ecology, climate science, and urban planning while lowering compute and data storage requirements. The framework supports multi-resolution analysis and robustness across database sizes (even with just 10% of data). ## Why These Papers Matter — A New Era in Computer Vision Today’s research highlights how much AI has matured — not by relying on more data, but by learning smarter from what it sees. FLAIR zooms into visual-text details that CLIP overlooked. OpenMIBOOD redefines reliability standards for medical anomaly detection. DyCON embraces uncertainty to segment better where others falter. RANGE retrieves meaningful context to make location-aware predictions — even in the absence of images. These efforts exemplify a new era of vision research: one where generalization is paired with domain awareness, and where performance is coupled with purpose. As these systems become part of critical real-world workflows, they promise not just better models — but more capable, trustworthy, and transparent ones. Thanks for following along with our CVPR 2025 coverage. Hope to see you all at the [virtual meetup](https://voxel51.com/events/best-of-cvpr-july-11-2025). ## What is next? If you’re interested in following along as I dive deeper into the world of AI and continue to grow professionally, feel free to connect or follow me on [LinkedIn](https://www.linkedin.com/in/paula-ramos-phd/). Let’s inspire each other to embrace change and reach new heights! You can find me at some [Voxel51 events](https://voxel51.com/events), or if you want to join this fantastic team, it’s worth taking a look at this page: [https://voxel51.app/careers](https://voxel51.com/careers) ![](https://cdn.sanity.io/images/h6toihm1/production/571af7476f44954ee67389529f23791f4d50ff97-990x990.png?auto=format&dpr=2&fit=max&q=75&w=990) ## References \[1\] R. Xiao, S. Kim, M.-I. Georgescu, Z. Akata, and S. Alaniz, “FLAIR: VLM with Fine-grained Language-informed Image Representations,” in _Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR)_, 2025\. Temporal link: [https://arxiv.org/abs/2412.03561](https://arxiv.org/abs/2412.03561) \[2\] M. Gutbrod, D. Rauber, D. W. Nunes, and C. Palm, “OpenMIBOOD: Open Medical Imaging Benchmarks for Out-Of-Distribution Detection,” _Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)_, 2025\. Temporal link: [https://arxiv.org/abs/2503.16247](https://arxiv.org/abs/2503.16247) \[3\] M. Assefa, M. Naseer, I. I. Ganapathi, S. S. Ali, M. L. Seghier, and N. Werghi, “DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation,” in _Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR)_, 2025. \[4\] A. Dhakal, S. Sastry, S. Khanal, A. Ahmad, E. Xing, and N. Jacobs, “RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings,” in _Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR)_, 2025\. Temporal link: [https://arxiv.org/pdf/2502.19781](https://arxiv.org/pdf/2502.19781) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [CVPR](https://voxel51.com/blog/tag/cvpr) ![](https://cdn.sanity.io/images/h6toihm1/production/e926c07c7d1426c0fde8fdefa637c528d47b16f4-512x512.webp?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/db7358784a18a2ffa365704e7d941e73fbdf1fcd-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 1\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1) [![](https://cdn.sanity.io/images/h6toihm1/production/c63b546423b6cec32ccccc7df6bc4e0fced0b1a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 2\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2) [![](https://cdn.sanity.io/images/h6toihm1/production/770b8cfdbd7944916b1195dc11e5b173dbab8e97-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 and the State of Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 18, 2023](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-84-lllmstxt|> ## Raleigh AI & ML Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/bea9d97e9c547b98238656db6a83c465784b02fd-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/bea9d97e9c547b98238656db6a83c465784b02fd-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) In-person Americas Meetups Raleigh AI, ML and Computer Vision Meetup - August 20, 2025 Aug 20, 2025 5:30 - 8:30 PM TEKsystems 4300 Edwards Mill Rd Raleigh, NC Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/47df4433ba2a5453886db35375eb5cf01c156e74-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Hanxue Gu Duke University Bio ![](https://cdn.sanity.io/images/h6toihm1/production/fe3c17f4f82391695a11d150b838cec1c17f6814-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Ricardo Henao Duke University Bio ![](https://cdn.sanity.io/images/h6toihm1/production/839e47cf184866af8158b4e833f3ceb83ac51072-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Heather (Dunlop) Couture, PhD Pixel Scientia Labs, LLC Bio ![](https://cdn.sanity.io/images/h6toihm1/production/3474fa592df5384527fc788eedb80aa551d813c8-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos, PhD Voxel51 Bio About this event Join the Meetup to hear talks from experts on AI, ML, and Computer Vision. Schedule Adapting Vision Foundation Models to Medical Imaging: Strategies and Clinical Applications ![](https://cdn.sanity.io/images/h6toihm1/production/47df4433ba2a5453886db35375eb5cf01c156e74-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Hanxue Gu Duke University Bio Foundation models like SAM and DINO-v2 have shown strong performance on natural image tasks. However, when applied directly to medical imaging, they often underperform due to domain shifts, limited labeled data, and modality-specific challenges. This raises an important question: how can we adapt foundation models to work reliably and meaningfully in medical images? In this talk, I will share our research efforts toward answering that question. I will begin by exploring several fine-tuning strategies for different data scenarios, ranging from few-shot labeled examples to large collections of unlabeled scans. These strategies aim to help identify the optimal adaptation framework under various data availability settings. I will then introduce a series of models we developed based on these insights. SegmentAnyBone and SegmentAnyMuscle are two SAM-based models designed for accurate bone and muscle segmentation across all body locations and a wide range of MRI sequences. MRI-Core is a self-supervised model that learns general-purpose MRI features from unlabeled data and can be easily adapted to multiple downstream tasks. Finally, I will present a clinical application where one of these models is used to support abdominal surgical risk prediction. This example shows how I have explored using these models to contribute to real-world clinical decision-making. I hope this talk can share some of my experiences in building foundation models that are both practical for research and adaptable to clinical settings and to spark new insights and discussions in this field! Learning with Small Datasets in Real-World Medical Imaging Applications ![](https://cdn.sanity.io/images/h6toihm1/production/fe3c17f4f82391695a11d150b838cec1c17f6814-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Ricardo Henao Duke University Bio The talk will explore approaches that are effective when developing computer vision models for real-world medical imaging applications in situations where available datasets are limited in size. This setting is of special interest because the most challenging prediction problems in the medical domain have usually low incidence rates, thus resulting in (relatively) small datasets. Specifically, it will consider data-efficient architectures, multi-task learning and data augmentation through pseudo interventions. For illustration, a use case in which volumetric ophthalmology images are used to predict geographic atrophy conversion will be discussed. Bias & Batch Effects in Medical Imaging ![](https://cdn.sanity.io/images/h6toihm1/production/839e47cf184866af8158b4e833f3ceb83ac51072-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Heather (Dunlop) Couture, PhD Pixel Scientia Labs, LLC Bio Medical AI models can exhibit concerning biases, such as the ability to predict race from radiology images, which is impossible for human experts. This talk will examine bias and batch effects in medical imaging, beginning with a histopathology case study to illustrate the origins of some of these biases. I'll cover detection methods, such as exploratory data analysis, and mitigation strategies, including careful cross-validation and model-level interventions. While research has shown that foundation models reduce some biases, they don't eliminate the problem entirely. Bias represents a fundamental challenge in medical AI requiring early detection, careful validation, and tailored mitigation approaches. Managing Medical Imaging Datasets: From Curation to Evaluation ![](https://cdn.sanity.io/images/h6toihm1/production/3474fa592df5384527fc788eedb80aa551d813c8-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Paula Ramos, PhD Voxel51 Bio High-quality data is the cornerstone of effective machine learning in healthcare. This talk presents practical strategies and emerging techniques for managing medical imaging datasets, from synthetic data generation and curation to evaluation and deployment. We’ll begin by highlighting real-world case studies from leading researchers and practitioners who are reshaping medical imaging workflows through data-centric practices. The session will then transition into a hands-on tutorial using FiftyOne, the open-source platform for visual dataset inspection and model evaluation. Attendees will learn how to load, visualize, curate, and evaluate medical datasets across various imaging modalities. Whether you're a researcher, clinician, or ML engineer, this talk will equip you with practical tools and insights to improve dataset quality, model reliability, and clinical impact. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-85-lllmstxt|> ## Best of CVPR Event [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/a40cbd0a71e7aa91b458ffe06d5eb6d1fc8bb43b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/a40cbd0a71e7aa91b458ffe06d5eb6d1fc8bb43b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Meetups Best of CVPR – July 9, 2025 This event has ended, but you can still catch up! Watch the on-demand recordings and register for our [future events.](https://voxel51.com/events) Jul 9, 2025 9 AM Pacific Online. Register for the Zoom! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/8244ed8cc90991f9872b860b86cee5e8136be5bc-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Tin Stribor Sohn Porsche AG Bio ![](https://cdn.sanity.io/images/h6toihm1/production/c52411dc6962a764f97f7494fbf3e60d399beacb-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Cecilia Curreli TUM Bio ![](https://cdn.sanity.io/images/h6toihm1/production/0828bbacd7c2c6f668caac870ee3db5964476f0f-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Sudhir Sornapudi, PhD Corteva Agriscience Bio ![](https://cdn.sanity.io/images/h6toihm1/production/b4ed57d4e561cd3b2cca585c427cbc5b59117c39-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Wang Benquan Nanyang Technological University Bio ![](https://cdn.sanity.io/images/h6toihm1/production/51787eff79f6547125a83983919a3b93a2e2928b-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Ruyi An University of Texas Bio About this event Welcome to the Best of CVPR series, your virtual pass to some of the groundbreaking research, insights, and innovations that defined this year’s conference. Live streaming from the authors to you. Schedule What Foundation Models really need to be capable of for Autonomous Driving - The Drive4C Benchmark ![](https://cdn.sanity.io/images/h6toihm1/production/8244ed8cc90991f9872b860b86cee5e8136be5bc-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Tin Stribor Sohn Porsche AG Bio Foundation models hold the potential to generalize the driving task and support language-based interaction in autonomous driving. However, they continue to struggle with specific reasoning tasks essential for robotic navigation. Current benchmarks typically provide only aggregate performance scores, making it difficult to assess the underlying capabilities these models require. Drive4C addresses this gap by introducing a closed-loop benchmark that evaluates semantic, spatial, temporal, and physical understanding—enabling more targeted improvements to advance foundation models for autonomous driving. Human Motion Prediction – Enhanced Realism via Nonisotropic Gaussian Diffusion ![](https://cdn.sanity.io/images/h6toihm1/production/c52411dc6962a764f97f7494fbf3e60d399beacb-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Cecilia Curreli TUM Bio Predicting future human motion is a key challenge in generative AI and computer vision, as generated motions should be realistic and diverse at the same time. This talk presents a novel approach that leverages top-performing latent generative diffusion models with a novel paradigm. Nonisotropic Gaussian diffusion leads to better performance, fewer parameters, and faster training at no additional computational cost. We will also discuss how such benefits can be obtained in other application domains. Efficient Few-Shot Adaptation of Open-Set Detection Models ![](https://cdn.sanity.io/images/h6toihm1/production/0828bbacd7c2c6f668caac870ee3db5964476f0f-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Sudhir Sornapudi, PhD Corteva Agriscience Bio We propose an efficient few-shot adaptation method for the Grounding-DINO open-set object detection model, designed to improve performance on domain-specific specialized datasets like agriculture, where extensive annotation is costly. The method circumvents the challenges of manual text prompt engineering by removing the standard text encoder and instead introduces randomly initialized, trainable text embeddings. These embeddings are optimized directly from a few labeled images, allowing the model to quickly adapt to new domains and object classes with minimal data. This approach demonstrates superior performance over zero-shot methods and competes favorably with other few-shot techniques, offering a promising solution for rapid model specialization. OpticalNet: An Optical Imaging Dataset and Benchmark Beyond the Diffraction Limit ![](https://cdn.sanity.io/images/h6toihm1/production/b4ed57d4e561cd3b2cca585c427cbc5b59117c39-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Wang Benquan Nanyang Technological University Bio ![](https://cdn.sanity.io/images/h6toihm1/production/51787eff79f6547125a83983919a3b93a2e2928b-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Ruyi An University of Texas Bio Optical imaging capable of resolving nanoscale features would revolutionize scientific research and engineering applications across biomedicine, smart manufacturing, and semiconductor quality control. However, due to the physical phenomenon of diffraction, the optical resolution is limited to approximately half the wavelength of light, which impedes the observation of subwavelength objects such as the native state coronavirus, typically smaller than 200 nm. Fortunately, deep learning methods have shown remarkable potential in uncovering underlying patterns within data, promising to overcome the diffraction limit by revealing the mapping pattern between diffraction images and their corresponding ground truth object images. However, the absence of suitable datasets has hindered progress in this field —— collecting high-quality optical data of subwavelength objects is highly difficult as these objects are inherently invisible under conventional microscopy, making it impossible to perform standard visual calibration and drift correction. Therefore, we provide the first general optical imaging dataset based on the “building block” concept for challenging the diffraction limit. Drawing an analogy to modular construction principles, we construct a comprehensive optical imaging dataset comprising subwavelength fundamental elements, i.e., small square units that can be assembled into larger and more complex objects. We then frame the task as an image-to-image translation task and evaluate various vision methods. Experimental results validate our “building block” concept, demonstrating that models trained on basic square units can effectively generalize to realistic, more complex unseen objects. Most importantly, by highlighting this underexplored AI-for-science area and its potential, we aspire to advance optical science by fostering collaboration with the vision and machine learning communities. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-86-lllmstxt|> ## Taranis Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/899b1fbd73a4f3061789a79484124b63b384d688-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Taranis FiftyOne helps Taranis provide farmers with leaf-level insights for healthier crops and fields Apr 19, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) Taranis crop intelligence solutions use deep agronomic expertise and the most advanced AI and submillimeter image technologies to automatically detect crop threats and provide farmers with leaf-level insights. Taranis serves more than 100 agribusinesses, thousands of farmers, and millions of acres around the world, and is growing rapidly. > "We’ve been using FiftyOne for over a year and it has drastically changed the way we work. The ability to easily display and analyze our images and their metadata, including experiment results, has been a refreshing change compared to the way we’ve worked before – mainly writing our own metrics and viewers. I’ve personally used FiftyOne for a segmentation model I’ve trained – trying to analyze the results and visually see what my model outputs has been really easy and fluid thanks to FiftyOne." – Ido Greenfeld, AI Team Lead at Taranis ![](https://cdn.sanity.io/images/h6toihm1/production/adaa1c5e7617b647bb1ce5ab36069f80bd402094-1033x660.png?auto=format&dpr=2&fit=max&q=75&w=1033) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-87-lllmstxt|> ## Automated Data Labeling Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) How Automated Data Labeling Enhances Computer Vision Efficiency and Accuracy Apr 17, 2025 • 8 min read Article content In this article [The Rise of Computer Vision (CV)](https://voxel51.com/blog/how-automated-data-labeling-enhances-computer-vision-efficiency-and-accuracy#ca5586e9c08b) [The Bottleneck of Manual Data Labeling](https://voxel51.com/blog/how-automated-data-labeling-enhances-computer-vision-efficiency-and-accuracy#2b1c5cf966c5) [Enter Automated Data Labeling: A Game Changer](https://voxel51.com/blog/how-automated-data-labeling-enhances-computer-vision-efficiency-and-accuracy#171f4053a5da) [The Benefits of Automated Data Labeling for Efficiency](https://voxel51.com/blog/how-automated-data-labeling-enhances-computer-vision-efficiency-and-accuracy#afe45f194347) In this article [The Rise of Computer Vision (CV)](https://voxel51.com/blog/how-automated-data-labeling-enhances-computer-vision-efficiency-and-accuracy#ca5586e9c08b) [The Bottleneck of Manual Data Labeling](https://voxel51.com/blog/how-automated-data-labeling-enhances-computer-vision-efficiency-and-accuracy#2b1c5cf966c5) [Enter Automated Data Labeling: A Game Changer](https://voxel51.com/blog/how-automated-data-labeling-enhances-computer-vision-efficiency-and-accuracy#171f4053a5da) [The Benefits of Automated Data Labeling for Efficiency](https://voxel51.com/blog/how-automated-data-labeling-enhances-computer-vision-efficiency-and-accuracy#afe45f194347) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/f3fb986c6ca0f0135a30f11a2efe01bb95e10817-1270x952.png?auto=format&dpr=2&fit=max&q=75&w=1270) The animation above demonstrates how computer vision algorithms detect, classify, and segment objects, converting raw visual data into structured training data. Industries today generate vast amounts of image and video content, creating demand for labeling methods that exceed the capabilities of traditional manual annotation. Automated data labeling techniques have emerged as effective solutions, providing scalable, precise, and efficient approaches for building high-quality datasets essential for machine learning and artificial intelligence applications. As machine learning and artificial intelligence drive innovation across industries like healthcare, automotive, retail, and security, the rapid growth in image and video data requires [efficient annotation](https://docs.voxel51.com/user_guide/annotation.html) methods. Traditionally, labeling this data has been manual, slowing projects and increasing costs. However, automated data labeling technologies now revolutionize dataset creation by improving both accuracy and efficiency. This article explores the transition from manual to automated labeling, highlights key enabling technologies, and demonstrates how [FiftyOne](https://voxel51.com/fiftyone/) simplifies annotation workflows, helping data scientists and AI practitioners rapidly scale their algorithms without compromising data quality. ## **The Rise of Computer Vision (CV)** Over the last decade, we've seen computer vision evolve from an obscure research field into a mission-critical technology. Automotive manufacturers employ CV-based machine learning for self-driving systems; healthcare providers rely on it for disease detection via medical imaging; retail giants use object detection to track items; and government agencies harness it for security surveillance. In short, computer vision pervades countless applications in modern life. ### **The Importance of Labeled Datasets** ![](https://cdn.sanity.io/images/h6toihm1/production/7bbb3157271613a63af1d61a7dea02c60ddbaeec-1950x428.png?auto=format&dpr=2&fit=max&q=75&w=1600) Behind every successful machine learning or CV application is a robust training dataset of images or videos annotated with relevant objects, classes, or attributes. Accurate labels enable machine learning models to interpret visual data effectively and generalize to new inputs, while incomplete or inconsistent datasets negatively impact a model’s performance. Annotations serve as reference points during both training and validation, helping models refine predictions by providing clear examples of correct object classification, segmentation, and localization. Accurate labels clearly delineate object boundaries and consistently represent object classes, whereas inaccurate labels often misrepresent object locations, sizes, or categories, leading to errors and reduced model reliability. ### **The Challenge of Manual Data Labeling** Although labeled data is critical, manual labeling is labor-intensive, prone to fatigue, bias, and human error, especially with thousands or millions of images. As datasets grow, manual annotation increasingly becomes a bottleneck, driving organizations toward automated labeling solutions that significantly reduce human effort while maintaining or enhancing labeling quality. ## **The Bottleneck of Manual Data Labeling** ### **Manual Data Labeling Workflow** Historically, the manual data labeling process for computer vision has looked like this: - **Identify objects/features**: An annotator locates items of interest in each image or video frame (e.g., bounding boxes around cars, tumors in scans). - **Assign labels**: Each object or region receives an appropriate label (e.g., "person," "car," "tumor"). - **Quality check**: A second round of review might spot inconsistencies or mistakes. Though direct, this becomes overwhelming with millions of data points, introducing several challenges associated with manual labeling, outlined in the table below: **Challenge** **Description** Time and Cost Employing human annotators or external services is expensive and slow. Large-scale labeling tasks can span weeks or months. Inconsistencies Even expert teams may label data differently, introducing subjective bias (e.g., varying interpretations of partial occlusions). Scaling Limits Manually annotating datasets that require continuous updates (like autonomous vehicle footage) quickly becomes cumbersome and impractical. These difficulties underscore the necessity of automated data labeling for modern, large-scale computer vision initiatives. ## **Enter Automated Data Labeling: A Game Changer** ### **What is Automated Data Labeling?** ![](https://cdn.sanity.io/images/h6toihm1/production/a7c1c71844479a13a97f329a9b66ff8160744697-800x450.png?auto=format&dpr=2&fit=max&q=75&w=800) Automated data labeling, or auto labeling, seeks to automate data labeling tasks, reducing human involvement by leveraging machine learning pipelines or pre-trained models to suggest labels, limiting human input primarily to validation or targeted corrections. Techniques include semi-supervised learning, active learning loops, and using pre-trained networks for rapid labeling of new datasets. ### **Techniques for Automated Data Labeling** - **Active Learning**: The system flags uncertain or highly informative samples for human review, while it auto-labels straightforward examples. - **Weakly Supervised Learning**: Partial or uncertain labels are refined using high-level rules. A bounding box guess might be adjusted by the system. - [**Leveraging Pre-Trained Models**](https://docs.voxel51.com/model_zoo/index.html): A model trained on broad data (e.g., ImageNet) annotates new samples, which humans then refine. Iterating these improvements can train the model further, creating a virtuous cycle. ![](https://cdn.sanity.io/images/h6toihm1/production/3e62887a2ed45f36c5c5fca2fea6d892676c4b2e-2379x1262.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## **The Benefits of Automated Data Labeling for Efficiency** One standout advantage of automated data labeling is the dramatic reduction in time and costs, achieved by using machine learning algorithms for bulk annotation tasks; this results in quicker dataset turnarounds and allows human effort to be reserved for more complex or ambiguous edge cases. ![](https://cdn.sanity.io/images/h6toihm1/production/ffeaf36126d305d82e2f9cd0ba9301c755dc12ad-3570x1274.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### **Improved Scalability** Manual efforts might suffice for modest projects, but fail to scale for massive data streams. Automated pipelines, however, improve with volume: - Effectively handle vast, diverse datasets - Adapt quickly to dynamic environments, like new road signs or store layouts ### **Consistency and Reduced Bias** Humans can inadvertently produce inconsistent labels, especially under pressure. Automated systems employ the same logic every time, generating uniform bounding boxes, categories, or segmentation masks, resulting in more consistent automated data. This uniformity reduces label noise and boosts overall [dataset quality](https://voxel51.com/blog/data-quality-the-hidden-driver-of-ai-success/). ### **Enhanced Accuracy with Automation: Beyond Efficiency** Automated data labeling enhances accuracy by enabling the annotation of a wider variety of images, including rare cases and complex features often tedious for humans, leading to better model generalization. Additionally, these systems continuously improve their own precision over time by learning from human feedback through confirmation and correction loops. ![](https://cdn.sanity.io/images/h6toihm1/production/bc8a508bd66dbfd891364b8f5eb81d2d5a84cd5b-2370x1780.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### **Leveraging FiftyOne for Advanced Labeling** Practical automated labeling requires strong infrastructure for dataset management, refinement, and quality control. Enter Voxel51’s open-source tool, FiftyOne, which supports every stage of your computer vision workflow. #### **FiftyOne's Role in Automated Data Labeling** FiftyOne integrates with various ML frameworks and active learning setups, acting as a central hub for: - Uploading and versioning unlabeled data - Merging auto-labeled results from models or weak supervision - Visualizing annotations in a user-friendly interface - Allowing rapid validation and edits By consolidating data and tools, FiftyOne helps you assess label consistency and overall dataset integrity in one place. #### Key FiftyOne Features for Data Labeling 1. **Integration with Active Learning Frameworks** FiftyOne pinpoints ambiguous data points that need human attention. Updated labels are fed back to retrain the system. 2. **Refinement Tools** If bounding boxes or categories are somewhat accurate, FiftyOne’s GUI allows quick corrections. This human-in-the-loop synergy maintains quality without starting from scratch. 3. **Collaboration Features** Teams often have multiple subject matter experts. FiftyOne lets users share datasets, track annotation changes, and unify feedback, ensuring consistent labeling and accountability. ### **Real-World Applications: Automation in Action** #### **Medical Imaging** In healthcare, MRI or X-ray images often require labeling to detect tumors or fractures. Automated labeling pinpoints potential anomalies, drastically reducing a radiologist’s effort. This shorter cycle propels machine learning-driven diagnostics and speeds up critical refinements. #### **Self-Driving Cars** Autonomous vehicle systems rely on huge datasets containing cars, pedestrians, and road elements. Automated labeling rapidly annotates camera feeds, letting self-driving companies adapt to changing conditions. As roads evolve, these pipelines keep the labeled data current without overtaxing human teams. ![](https://cdn.sanity.io/images/h6toihm1/production/d3dd95126e4b8d5256c803d25a6717eecd9ba55b-2560x1920.jpg?auto=format&dpr=2&fit=max&q=75&w=1600) #### **Retail Object Recognition** Retailers monitor inventory, warehouse workflows, and customer interactions through CV. Automated labeling supports swift recognition of new products, quick re-annotation for seasonal changes, and accurate tracking despite constantly shifting store layouts. This ensures models stay relevant without relentless manual interventions. #### **Future Trends in Automated Labeling** Future trends in automated labeling include emerging techniques like using Natural Language Processing (NLP) to create richer, context-aware semantic annotations by linking textual descriptions to images, and employing self-supervised learning models that train on unlabeled data to become proficient at auto-labeling with reduced human intervention. #### **Fully Automated Labeling Pipelines** While expert oversight may still be necessary, routine labeling tasks are rapidly approaching full automation, particularly in familiar domains. Feedback loops continually refine algorithms, enabling near real-time generation of labeled data. Ultimately, we're moving toward on-demand systems where prior knowledge automatically annotates datasets, humans provide essential corrections, and the process iterates seamlessly. ### **Final Insights** #### **The Power of Automated Data Labeling** Automated labeling marks a pivotal shift for machine learning and computer vision. By taking over repetitive annotation, teams reduce costs, expedite model development, and broaden project scopes. Crucially, automating labeling can also improve quality, ensuring thorough coverage of complex, real-world data. #### **How FiftyOne Empowers AI Builders:** Even top-tier auto-labeling gains from comprehensive data management. Voxel51’s FiftyOne helps data scientists: - Incorporate varied labeling methods (weak supervision, active learning, or pre-trained predictions). - Quickly refine results with human-in-the-loop edits. - Collaborate with domain experts. - Analyze label consistency and dataset balance to ensure thoroughness. ### **Accelerate Your Computer Vision Workflow Today** Whether you are building diagnostic tools, self-driving systems, or advanced retail analytics, embracing automated labeling can help you outpace manual bottlenecks. With the right pipelines and platforms like FiftyOne, you achieve scalable, high-quality results that adapt with incoming data. Consider integrating automated labeling into your CV workflow now to unlock faster, more accurate model outcomes. **Image Citations** - Lin, Tsung-Yi, et al. _"Microsoft COCO: Common Objects in Context."_ COCO Dataset 2017 Validation Split, cocodataset.org, 2017, [https://cocodataset.org/#home](https://cocodataset.org/#home). Accessed 24 Mar. 2025. [labeling](https://voxel51.com/blog/tag/labeling) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/d2e24d0a14de508f8ccd36eaffd3b909c9f193b9-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Image Embeddings Transform Computer Vision Capabilities\\ \\ Learn\\ \\ • \\ \\ Nov 25, 2024](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities) [![](https://cdn.sanity.io/images/h6toihm1/production/b4fd054bf7574e4060ffc7a9a8f201c31a66c327-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Using Computer Vision to Enhance Customer Experience in Retail\\ \\ Learn\\ \\ • \\ \\ Jan 23, 2025](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail) [![](https://cdn.sanity.io/images/h6toihm1/production/e0c50d51f1558ec65daf139f10ce4bb099e088ce-2338x1294.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ AI for Predictive Maintenance Using Computer Vision\\ \\ Learn\\ \\ • \\ \\ Apr 17, 2025](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-88-lllmstxt|> ## Video Search and Curation [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news), [Integrations](https://voxel51.com/blog/category/integrations) Search and curate video data with FiftyOne, Twelve Labs, and Databricks Vector Search Jun 5, 2025 • 5 min read Article content In this article [Find the exact video clip in seconds](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#85a3c976e539) [FiftyOne – visualize and curate your video data](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#bd1ea71817eb) [Twelve Labs – multimodal embeddings](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#8c4aa30ea7bd) [Databricks Vector Search – scalable indexing](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#92bde142968c) [Putting it together in a workflow](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#520624f56208) [Next steps](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#7244c44033ac) [Conclusion](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#28700f34de17) In this article [Find the exact video clip in seconds](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#85a3c976e539) [FiftyOne – visualize and curate your video data](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#bd1ea71817eb) [Twelve Labs – multimodal embeddings](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#8c4aa30ea7bd) [Databricks Vector Search – scalable indexing](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#92bde142968c) [Putting it together in a workflow](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#520624f56208) [Next steps](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#7244c44033ac) [Conclusion](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs#28700f34de17) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) TLDR; Unlabeled video search is notoriously hard. This post outlines a novel workflow that combines FiftyOne, Twelve Labs, and Databricks Vector Search to unlock fast, semantic video search and exploration, no manual labels required. ## Find the exact video clip in seconds Navigating massive video datasets to find the right clip for your Machine Learning problem can feel like an unattainable task. Most videos come without labels or metadata, making it nearly impossible to locate the samples you need, especially when dealing with thousands or even millions of hours of footage. And traditional tools fall short when it comes to video search and curation. What if there were a smarter way to search videos? Something that understands content without needing manual labels or tags and surfaces the exact clip you’re looking for? Recent advances in AI and data infrastructure are making this possible by combining: - Rich video embeddings from foundation models that capture semantic content - Fast similarity search over those embeddings to find matches across large datasets - User-friendly visualization for exploring and curating the results In this blog, we introduce a high-level workflow that addresses these challenges by combining three tools: [FiftyOne](https://voxel51.com/), [Twelve Labs](https://www.twelvelabs.io/), and [Databricks Vector Search](https://www.databricks.com/product/machine-learning/vector-search). This stack streamlines video understanding in ML workflows and enables users to index video content, search it by example or by natural language, and visualize the findings with ease. The result is a _single, smooth pipeline_ for working with video data at scale, from ingestion and curation to semantic search and analysis. ![A high-level overview of the video similarity search workflow, combining FiftyOne, Twelve Labs, and Databricks.](https://cdn.sanity.io/images/h6toihm1/production/3d5a2f20e076531f36e8855a4a020e7cbd4ca1f1-1242x698.png?auto=format&dpr=2&fit=max&q=75&w=1242)A high-level overview of the video similarity search workflow, combining FiftyOne, Twelve Labs, and Databricks. Before we go into the details, let’s get a bit more familiar with these tools. ## FiftyOne – visualize and curate your video data [FiftyOne](https://voxel51.com/) is a toolset for building high-quality computer vision datasets and models from Voxel51. It acts as a refinery for visual data so development teams can analyze model performance, uncover data gaps, find edge cases, and curate better datasets. With FiftyOne, users can interactively load and explore videos, inspect frames, and refine datasets with precision. It supports vector search to run semantic queries directly in the UI and integrates with embedding tools and vector DBs, making it the perfect hub for this workflow. ## Twelve Labs – multimodal embeddings [Twelve Labs](https://www.twelvelabs.io/) is a platform that provides state-of-the-art multimodal video understanding through its foundation models and APIs. It turns raw videos into rich vector embeddings that capture the full semantic essence, visuals, actions, audio, and even text of your data. With TwelveLabs, you can search videos with natural language or images. Looking for clips of “a car stopping at a red light”? You can just type in this query. Twelve Labs handles the understanding and returns the necessary output, without the need for a user to train anything. ## Databricks Vector Search – scalable indexing [**Databricks Vector Search**](https://www.databricks.com/product/machine-learning/vector-search), offers a managed, serverless way to store and query vector embeddings at scale. Fully integrated into the Databricks Data Intelligence Platform, it serves as the [**scalable embedding index**](https://docs.databricks.com/aws/en/generative-ai/create-query-vector-search) and lets you feed in embeddings (such as video vectors from TwelveLabs) to retrieve the most relevant matches by querying with a vector. ## Putting it together in a workflow In our workflow: - Twelve Labs provides the _multimodal intelligence_ needed to _understand_ video data and handles the heavy lifting of analyzing frames, audio, and text in videos to return embeddings. - Databricks Vector Search is where the Twelve Labs embeddings go and serves as the _scalable backend_, indexing those embeddings for fast similarity search. - FiftyOne acts as the _central hub_, orchestrating the entire workflow—from loading and curating video datasets to querying results to Databricks Vector Search and visualizing search matches. One of the most impressive aspects of this workflow is its setup simplicity, thanks to the tight integration of the three tools. In the past, implementing something like this might have required stitching together code from multiple libraries, setting up your own vector database, and dealing with a lot of glue logic. Here, much of the heavy lifting is abstracted away. The section below provides an overview of all the steps, but you can find detailed step-by-step instructions in [this example notebook.](https://github.com/danielgural/fiftyone_twelvelabs_databricks_demo) ![](https://cdn.sanity.io/images/h6toihm1/production/5b31a591083fcf36d94cc82f348ebce96059271d-1750x738.png?auto=format&dpr=2&fit=max&q=75&w=1600) - **Connect to Databricks and create a similarity index** Establish a connection from your environment (e.g. a notebook or script) to your Databricks workspace where vector search is enabled. In Databricks, you create a new vector search endpoint (index) that will store the video embeddings. [Refer to Databricks Vector Search integration docs](https://docs.voxel51.com/integrations/mosaic.html#basic-recipe). - **Load your video dataset into FiftyOne** Using FiftyOne’s dataset abstractions, you import the videos and any available metadata or labels. In FiftyOne, you can explore the videos, inspect frames, and do basic curation (filtering, viewing distributions of any labels etc) - **Generate video embeddings with Twelve Labs** Using the [Twelve Labs integration](https://github.com/danielgural/semantic_video_search), it only takes a single command or function call to compute embeddings for all samples in the FiftyOne dataset. They can later be compared with embeddings of text queries or images, enabling cross-modal search. It's super easy to retrieve embeddings in code, too. ```python 1from twelvelabs import TwelveLabs 2 3# Start client 4client = TwelveLabs(api_key=API_KEY) 5 6# Pass your video file in via API 7task = client.embed.task.create( 8model_name="Marengo-retrieval-2.7", 9video_file=file_path, 10) 11 12# Retrieve embeddings 13retrieved_task = task.retrieve( 14embedding_option=["visual-text", "audio"] ``` ![](https://cdn.sanity.io/images/h6toihm1/production/32d2d9d575ce0b8b24b2591b7f372f2193c3208e-1024x578.webp?auto=format&dpr=2&fit=max&q=75&w=1024) - **Index those vectors in Databricks** Create a **similarity index** in Databricks and upload the vectors through FiftyOne’s API. [See docs for more details.](https://docs.voxel51.com/integrations/mosaic.html#using-the-mosaic-backend) From here on, you don’t need to loop over all videos for search; you can simply query the index! - **Search by image or text** With the index in place, you can now perform semantic search queries against your video data by image or text.Pick a sample frame and find similar clips across your dataset or search using natural language like “person walking a dog at night.” ![](https://cdn.sanity.io/images/h6toihm1/production/b711bbae16ed4c23d9515715d1accac025e57e3a-1246x670.png?auto=format&dpr=2&fit=max&q=75&w=1246) - **Visualize and refine results in FiftyOne** The final step brings everything together in FiftyOne’s interactive UI, where you can review, refine, and iterate on search results. Extract visual data insights by flag false positives, adjust queries, and re-run all these steps in one place. Check out the details [in this notebook.](https://github.com/danielgural/fiftyone_twelvelabs_databricks_demo) ## Next steps If working with video data has been a challenge, try this integrated approach. Here’s how you can get started. 1. [**Clone the example notebook**](https://github.com/danielgural/fiftyone_twelvelabs_databricks_demo) 2. Get your free API keys from [Twelve Labs](https://www.twelvelabs.io/?utm_campaign=demo&utm_medium=twelvelabs&utm_source=logo) 3. Load your dataset into FiftyOne 4. Run your first query ## **Conclusion** Working with video at scale doesn’t have to be overwhelming. With the combined power of FiftyOne, Twelve Labs, and Databricks, you can easily search and explore video content. Try the workflow, check out the docs, and see how it can level up your video understanding. We’re excited to see what the community will build and discover! [video](https://voxel51.com/blog/tag/video) [data curation](https://voxel51.com/blog/tag/data-curation) [embeddings](https://voxel51.com/blog/tag/embeddings) [similarity search](https://voxel51.com/blog/tag/similarity-search) [vector search](https://voxel51.com/blog/tag/vector-search) [integrations](https://voxel51.com/blog/tag/integrations) ![](https://cdn.sanity.io/images/h6toihm1/production/3b39056326e925c10b46da1324bc3c5840a1629c-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Dan Gural Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/6aafb2b5fa699824c252fabfe2607eaeb820616a-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Enabling the AV Datasets of the Future with NVIDIA NuRec and FiftyOne\\ \\ Product & News\\ \\ • \\ \\ Aug 11, 2025](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/35e10bff3f7e49854806cbbd163022932bf07548-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How Voxel51 is Powering Physical AI with Databricks\\ \\ Product & News\\ \\ • \\ \\ Aug 20, 2025](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks) [![](https://cdn.sanity.io/images/h6toihm1/production/272733186f35c5c6aca68b421960cd41c570e24b-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Composed Image Retrieval at CVPR 2025\\ \\ Event Recaps\\ \\ • \\ \\ Jun 2, 2025](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-89-lllmstxt|> ## Forsight Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/af1426b965796134b254c4c1ff02be13740eebc5-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Forsight Forsight Finds a Centralized Dataset Management Solution in FiftyOne Teams May 1, 2025 - Company: Forsight - Industry: AL Software for safety & security on jobsites - Challenge: Scaling dataset management as the R&D team grows - Solution: Dataset management through FiftyOne Teams Daily managed visual data 1.5 TB Engineering time saved 1/2 FTE Article content In this article [Success Story at a Glance](https://voxel51.com/customers/forsight#178a31b85f6c) [How Forsight Uses AI](https://voxel51.com/customers/forsight#592195690e4d) [Scaling Dataset Management as Business Grows](https://voxel51.com/customers/forsight#20e938ec7c62) [FiftyOne Teams Makes Dataset Management Easy](https://voxel51.com/customers/forsight#e38738c3fd7b) [Better data, better visibility, better models](https://voxel51.com/customers/forsight#3571ad79587a) [Cost effective dataset management](https://voxel51.com/customers/forsight#f648084326ef) [Enabling quick iteration](https://voxel51.com/customers/forsight#4be443e14e9d) [Conclusion](https://voxel51.com/customers/forsight#83960d7111ab) In this article [Success Story at a Glance](https://voxel51.com/customers/forsight#178a31b85f6c) [How Forsight Uses AI](https://voxel51.com/customers/forsight#592195690e4d) [Scaling Dataset Management as Business Grows](https://voxel51.com/customers/forsight#20e938ec7c62) [FiftyOne Teams Makes Dataset Management Easy](https://voxel51.com/customers/forsight#e38738c3fd7b) [Better data, better visibility, better models](https://voxel51.com/customers/forsight#3571ad79587a) [Cost effective dataset management](https://voxel51.com/customers/forsight#f648084326ef) [Enabling quick iteration](https://voxel51.com/customers/forsight#4be443e14e9d) [Conclusion](https://voxel51.com/customers/forsight#83960d7111ab) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) # Success Story at a Glance ### Challenge - As their R&D team grew, manual dataset management was no longer going to work - [Forsight](https://forsight.ai/) needed an easier, more transparent way to share image and video datasets among team members ### Solution - Forsight chose [FiftyOne](https://voxel51.com/) Teams as their central computer vision dataset management system where all the data is aggregated and consumed by all of their machine learning pipelines ### Results - 1.5TB of visual data is managed daily through FiftyOne Teams - Half an engineer’s time is saved - Datasets and model performance are improved - Costs are straightforward, not complex ## How Forsight Uses AI [Forsight](https://forsight.ai/) is on a mission to help keep workers and the jobsites where they work safe and secure. Forsight does this by using cutting edge AI vision technology to build solutions for dynamic environments that address challenges regarding safety, security, management and more. “Missing safety equipment or not adhering to safety regulations are common but preventable occurrences on jobsites in the US today. We founded Forsight around the idea that if we could save just one person’s life with AI and software, then our efforts would be well worth it,” said Ivan Ralašić, CTO and co-founder of Forsight. ![Source: forsight.ai](https://cdn.sanity.io/images/h6toihm1/production/864c1d83e96595af304dbb609acc537558e3b6a0-1024x274.png?auto=format&dpr=2&fit=max&q=75&w=1024)Source: forsight.ai Although Forsight started with an initial focus on safety in the construction industry, the company now also works with industrial plant owners and operators, manufacturers, mining operations, and other customers to make their sites safer and day to day tasks easier through the use of real-time, vision-powered technology. Forsight develops AI software that uses CCTV cameras to detect and predict safety incidents, security threats, and management issues in real time. Their solution monitors jobsites for safety issues, like missing personal protective equipment, no-go zones for safety or security reasons, as well as vehicles, license plates, and more. It does this all in real time, which means processing camera feeds ten times a second, and many cameras in parallel. ## Scaling Dataset Management as Business Grows Forsight’s AI technology is now deployed at hundreds of customer sites and as Forsight’s business grew, so did their R&D team. Ivan noted, “initially, dataset management was a manual process, and the ingestion of data to the training instances was cumbersome with all the syncing. Luckily, it was only me training the models at the time, and I somehow managed to keep track of everything. As we added more team members, we needed a more transparent and easier way to share datasets among team members.” Ivan wanted to scale and streamline dataset management as the R&D team grew, as well as avoid the pitfalls of manual dataset versioning. ![Source: @petergyang on Twitter](https://cdn.sanity.io/images/h6toihm1/production/888bb9192238fb832f8b60731ab7498efccdbd52-659x352.jpg?auto=format&dpr=2&fit=max&q=75&w=659)Source: @petergyang on Twitter Forsight’s primary requirements for a dataset management solution were that it had to fit into their existing cloud infrastructure and be cost effective. The Forsight engineering team tested a few different products and then found open source FiftyOne on their way to finding FiftyOne Teams. >> [FiftyOne](https://github.com/voxel51/fiftyone) is the open source tool for building high-quality datasets and computer vision models. >> [FiftyOne Teams](https://voxel51.com/) inherits all the goodness of open source FiftyOne and adds collaborative features built specifically for teams, including cloud-backed media, dataset permissions, versioning, sharing, and more. “The open source FiftyOne version stood out from other solutions because it had much more powerful features, and the FiftyOne Brain was a big bonus. FiftyOne worked really great, but as a team we needed additional features like central permission management and access control, and that’s where FiftyOne Teams comes into the picture,” explained Ivan. ## FiftyOne Teams Makes Dataset Management Easy Forsight is using [FiftyOne Teams](https://voxel51.com/) as their central dataset management system where all the data is aggregated and consumed by all of their machine learning pipelines. “It’s a great thing to have centralized dataset management in the form of FiftyOne Teams, which really makes our lives easier when curating datasets and training new models. All the team members have the same view on the datasets, which ensures that everyone understands the data used to train the models. This really saves a few hours of back and forth between team members,” said Ivan. Forsight’s R&D team trains different models to perform tasks such as detection of personal protective equipment, intrusion detection, construction vehicle detection, fire detection, semantic segmentation, and more. ![Source: Forsight’s FiftyOne Teams App](https://cdn.sanity.io/images/h6toihm1/production/91df43514c86c9ec11b38d300029efd859b58fa7-1024x601.png?auto=format&dpr=2&fit=max&q=75&w=1024)Source: Forsight’s FiftyOne Teams App Their lightweight CV and ML algorithms run on-premises on edge devices that process the CCTV camera streams in real time. For typical deployments, the edge devices are powered by NVIDIA’s embedded Jetson technology, and for larger deployments, Forsight uses NVIDIA GPUs that can process 25+ camera streams in real time in parallel. Video feeds and inferred metadata are then sent into Forsight’s cloud provider, which is AWS. They’re using Amazon EC2 instances for training their ML models, and Amazon S3 buckets for media storage. Ivan added, “We wanted to have high throughput between our S3 buckets and the EC2 instances, which we are able to get from FiftyOne Teams.” In addition to integrating with AWS, Forsight also integrated FiftyOne Teams with CVAT for annotation, PyTorch for model training loops, and ClearML for experiment tracking. Overall, Forsight’s R&D team works with about 1.5 TB of data (mostly images, but some video datasets) on a daily basis, and all that data goes through FiftyOne Teams. ## Better data, better visibility, better models FiftyOne exists to give developers and scientists comprehensive visibility into their datasets, so they can c​urate better data and build better models. Because FiftyOne Teams is built on top of open source FiftyOne, it does all that and adds collaboration and access features that teams need. Forsight uses FiftyOne to automate dataset management, such as auto-tagging suspicious samples defined by rules set by mistakeness and embeddings. They also use FiftyOne to auto-balance their datasets and help reduce bias in their models. Adding the collaboration features of FiftyOne Teams extends dataset visibility team-wide. “The quality of our datasets is a key component of the performance of our models. FiftyOne helps us improve our datasets by identifying mislabeled data, adversarial samples, and unusually sized samples. FiftyOne Teams helps us track our datasets across multiple model runs and evaluate our models. We can see how our models are improving as we track model predictions on certain samples over multiple runs,” said Daniel Reiff, Machine Learning Engineer at Forsight. “We have seen performance improvements in our models directly due to using FiftyOne Teams for dataset management. FiftyOne Teams has greatly improved the visibility of our datasets across our entire R&D team, it has made it extremely easy for multiple team members to access and collaborate on datasets,“ explained Ivan. ![Source: Forsight’s FiftyOne Teams App](https://cdn.sanity.io/images/h6toihm1/production/85718ea9c44a8b722a2aef64c780755ba65690b6-1024x606.png?auto=format&dpr=2&fit=max&q=75&w=1024)Source: Forsight’s FiftyOne Teams App ## Cost effective dataset management The pricing model for other dataset management solutions can be complex, but with FiftyOne Teams, pricing is straightforward and includes unlimited data. “Some companies charge per gigabyte transferred, while others charge per image stored, per model trained, and per hour — it’s so complex. The costs are maybe manageable in the beginning to attract you to the product, but as you scale and have more and more data, then you run into problems. With FiftyOne Teams, pricing is straightforward, includes unlimited data, and that’s what we like,” explained Ivan. FiftyOne Teams also streamlines the process of working with data, freeing up valuable engineering time. Ivan estimates that the time saved with FiftyOne Teams is equivalent to half an engineer, which means the team can spend more time building new features and less time wrangling data. ## Enabling quick iteration The Forsight R&D team iterates quickly on their models, and uses FiftyOne Teams to help. “There are a ton of features in FiftyOne Teams that I use to iterate quickly on our models. For example, there are a lot of features of the [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html) that I find incredibly useful, including mistakenness and uniqueness that help me quickly identify suspicious samples to send to our labelers, get things turned around, and retrain the model.” Daniel gave a shout out to another one of his favorite features — [heatmaps](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#heatmaps). “My favorite FiftyOne Teams feature is heatmaps! At Forsight we love using Grad-CAM heatmaps to interpret what our models have learned. Intuitively, we use these heatmaps to build more adversarial samples into our datasets so we can rapidly improve our models. Being able to view these heatmaps in FiftyOne Teams is a great advantage because we can tag samples where the model has not learned the critical regions and build sets of adversarial samples,” said Daniel. Check out [this article](https://towardsdatascience.com/understand-your-algorithm-with-grad-cam-d3b62fce353) to learn more about how Forsight uses Grad-CAM heatmaps. ## Conclusion Forsight uses cutting edge AI vision technology to build solutions for jobsites that address challenges regarding safety, security, management, and more. The Forsight R&D team needed a dataset management solution and tested out a few products before finding open source FiftyOne, which ultimately led them to FiftyOne Teams. FiftyOne Teams meets all their requirements including integrating with their existing cloud infrastructure and being a cost effective solution. In addition, FiftyOne Teams has led to significant results in terms of improved speed, quality, and performance. See for yourself: - [Try FiftyOne](https://voxel51.com/docs/fiftyone/), it’s easy to get up and running in a few minutes - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help - Check out [FiftyOne Teams](https://voxel51.com/) to enable multiple users to securely collaborate on the same datasets and models ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-90-lllmstxt|> ## Amsterdam AI Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/f89e06b45ca827b6d50f8fc9cf0d1cb9a777204b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/f89e06b45ca827b6d50f8fc9cf0d1cb9a777204b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) In-person EMEA Meetups Amsterdam AI, ML and Computer Vision Meetup - July 14 Jul 14, 2025 5:30-8:30 PM Hotel Arena Gravesandestraat 55 Amsterdam Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/e9e6ac0aeee80a2921be316414fa2c05d7afb12c-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Tuana Celik LlamaIndex Bio ![](https://cdn.sanity.io/images/h6toihm1/production/8076a97931f85a61ec22a7df4704b1a4d39d6618-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Rafael Pierre Weet AI Bio ![](https://cdn.sanity.io/images/h6toihm1/production/8e77ff943b2ac0ced5fe9c512a85194ee3e9c5fb-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Santiago Ruiz Royal Schiphol Group Bio ![](https://cdn.sanity.io/images/h6toihm1/production/5b9770ebf879163fd519d57717c074eab4e49cca-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Voxel51 Bio About this event Hear talks from experts on cutting-edge topics in AI, ML, and computer vision on July 14 Schedule Build Event Driven Agentic Workflows with LlamaIndex ![](https://cdn.sanity.io/images/h6toihm1/production/e9e6ac0aeee80a2921be316414fa2c05d7afb12c-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Tuana Celik LlamaIndex Bio In this talk we'll be looking into a specific approach to building agentic applications, with custom event driven workflows. Build agentic AI applications that reflect, make decisions and branch out, have human-in-the-loop and more. Agentic OCR: The Next Frontier in Intelligent Document Processing ![](https://cdn.sanity.io/images/h6toihm1/production/8076a97931f85a61ec22a7df4704b1a4d39d6618-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Rafael Pierre Weet AI Bio The evolution of document processing has reached a new milestone with Agentic OCR—systems that not only read text but also act upon it. By combining OCR with AI agents, enterprises can achieve higher accuracy in data extraction, better handling of unstructured documents, and seamless integration into business workflows. This presentation will cover the architecture of Agentic OCR systems, their advantages over traditional methods, and case studies demonstrating their impact on efficiency and compliance in sectors like finance and healthcare. Deep Turnaround: Transforming Airport Operations with AI & Computer Vision ![](https://cdn.sanity.io/images/h6toihm1/production/8e77ff943b2ac0ced5fe9c512a85194ee3e9c5fb-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Santiago Ruiz Royal Schiphol Group Bio This talk explores how the critical flight preparation (turnaround) process, a major source of flight delays, can be transformed by cutting-edge AI. The "Deep Turnaround" solution leverages AI image-based processing and computer vision to capture over 70 unique events in real-time. This technology enables proactive management, significantly reducing turnaround delays and enhancing collaboration across the entire airport ecosystem. Visual Agents: What it takes to build an agent that can navigate GUIs like humans ![](https://cdn.sanity.io/images/h6toihm1/production/5b9770ebf879163fd519d57717c074eab4e49cca-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Harpreet Voxel51 Bio We’ll examine conceptual frameworks, potential applications, and future directions of technologies that can “see” and “act” with increasing independence. The discussion will touch on both current limitations and promising horizons in this evolving field. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-91-lllmstxt|> ## Visual AI in Manufacturing [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/c25840d7b086070df8290f52a85856ba1fd0874d-2881x1620.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/c25840d7b086070df8290f52a85856ba1fd0874d-2881x1620.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event Virtual Webinars & Workshops Manufacturing Visual AI in Manufacturing: How Multimodal Data Powers Adaptive Process Control Oct 8, 2025 9:00 - 10:30 Pacific Time Online. Register for the Zoom! About this event Join this webinar to learn how multimodal visual AI drives defect detection and adaptive control in manufacturing. Host ![](https://cdn.sanity.io/images/h6toihm1/production/33c8a7b110bd086417d17f82c8a160fe2bc38d00-256x256.png?auto=format&dpr=2&fit=max&q=75&w=96) Nick Lotz Technical Marketing Engineer Bio Join us for a fast-paced session that reveals how a data-centric visual AI workflow turns disparate sensor streams into a single, actionable source of truth on the factory floor. Through live demos and real-world case studies, you’ll see how fusing images, video, thermal, and PLC data accelerates defect detection, drives adaptive process control, and boosts product yield. **What you’ll learn:** - **Building a multimodal dataset**: ingesting images, video, and industrial sensor logs into a unified visual dataset - **Curating for quality**: rapid exploration, filtering, and labeling techniques that surface the edge-cases traditional QA misses - **Training & evaluating models**: fuse visual and non-visual features into one training and evaluation pipeline, track every experiment, and mine scenarios that contribute to model failures - **Closing the loop**: stream model outputs back to the line to trigger real-time corrective actions - **Measuring model impact**: align model performance with key metrics like scrap rate, OEE, and MTTR **Who should attend:** Manufacturing engineers, computer-vision practitioners, data scientists, and operations leaders looking to use visual AI to optimize their production lines. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-92-lllmstxt|> ## RIF Robotics Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/f5bfb518a5b173e357f7c318576cda65d29e7f59-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) RIF Robotics RIF Robotics uses FiftyOne to enable robot-assisted surgery Apr 22, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [RIF Robotics](https://www.rifrobotics.com/) is combining bleeding edge computer vision with modern robotics and building a flexible solution that can identify, inspect, and manipulate surgical instruments. Their goal: to boost sterile processing productivity and eliminate preparation errors that cause costly operating room delays and patient infections. > "RIF Robotics uses FiftyOne daily to translate between different dataset formats and visualize our computer vision datasets." – Kevin DeMarco, Ph.D, CEO & Co-Founder at RIF Robotics ![](https://cdn.sanity.io/images/h6toihm1/production/114550d2246a686d396d53cf828559fef3408399-955x538.png?auto=format&dpr=2&fit=max&q=75&w=955) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-93-lllmstxt|> ## AI, ML, and Computer Vision Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/81c77dd40160a833db0cf8aeebc0596d56fc4fcb-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/81c77dd40160a833db0cf8aeebc0596d56fc4fcb-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event Virtual Americas Meetups AI, ML and Computer Vision Meetup en Español - October 23, 2025 Oct 23, 2025 9 AM Pacific Virtually over Zoom. Sign up! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/257a66cef95943b20ef466b3c96237dfd96a0a0b-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Jose Blasco Instituto Valenciano de Investigaciones Agrarias Bio ![](https://cdn.sanity.io/images/h6toihm1/production/10cf7db945b2b78145505b2a512f7c909ad1670b-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Carlos Bustillo Platzi Bio ![](https://cdn.sanity.io/images/h6toihm1/production/52cd94981200404f81cf9d5aa131b26d33cd3008-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Adonai Vera Voxel51 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/f8273d705275b6740c75bd5d7df84254a62d3a34-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Brayan Monroy Universidad Industrial de Santander Bio About this event Join the Meetup to hear talks in Spanish from experts on cutting-edge topics across AI, ML, and computer vision. Schedule Del campo al dato: oportunidades, obstáculos y adopción de la IA en agricultura ![](https://cdn.sanity.io/images/h6toihm1/production/257a66cef95943b20ef466b3c96237dfd96a0a0b-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Jose Blasco Instituto Valenciano de Investigaciones Agrarias Bio La inteligencia artificial (IA) se ha consolidado como una herramienta clave para afrontar algunos de los mayores desafíos en la agricultura actual, desde la optimización del uso de insumos hasta el monitoreo preciso de cultivos y la predicción de rendimientos. En este seminario se presentarán varios casos prácticos reales en los que la IA ha demostrado su potencial para mejorar la eficiencia, sostenibilidad y rentabilidad en distintas etapas de la producción agrícola. Sin embargo, a pesar de sus beneficios, la adopción de estas tecnologías en el sector sigue siendo limitada. Factores como la falta de relevo generacional, el envejecimiento de la población agraria, la baja digitalización en el medio rural y la percepción de complejidad o desconfianza hacia las herramientas digitales suponen importantes barreras. Esta charla, se aborda éxitos y obstáculos actuales, destacando la importancia de diseñar soluciones accesibles, acompañadas de formación y apoyo técnico, que se ajusten a la realidad de un sector tradicional en proceso de transformación. Crea tu primer Visual Search desde cero ![](https://cdn.sanity.io/images/h6toihm1/production/10cf7db945b2b78145505b2a512f7c909ad1670b-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Carlos Bustillo Platzi Bio ¿Te imaginas poder buscar un objeto dentro de una imagen de la misma forma que lo harías en Google Images o Bing? En esta charla veremos paso a paso cómo diseñar e implementar un sistema de búsqueda visual desde cero utilizando redes neuronales y Python. Multimodalidad con sesgos: Entiende y evalúa VLMs para conducción autónoma con FiftyOne ![](https://cdn.sanity.io/images/h6toihm1/production/52cd94981200404f81cf9d5aa131b26d33cd3008-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Adonai Vera Voxel51 Bio ¿Tus VLMs realmente ven el peligro? Con FiftyOne te muestro cómo entender y evaluar modelos visión-lenguaje para conducción autónoma, haciendo visible el riesgo y el sesgo en segundos. Compararemos modelos en las mismas escenas, revelaremos fallos y edge cases, y verás un dashboard simple para decidir qué datos curar y qué ajustar. Te llevas un método claro, práctico y replicable para subir el listón de seguridad. Deep Learning Techniques for HDR modulo imaging ![](https://cdn.sanity.io/images/h6toihm1/production/f8273d705275b6740c75bd5d7df84254a62d3a34-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Brayan Monroy Universidad Industrial de Santander Bio This talk explores the transformative potential of modulo imaging for achieving unlimited dynamic range capture, fundamentally reimagining how we approach high dynamic range photography beyond traditional sensor limitations. By introducing cyclical intensity wrapping through the modulo operator, we unlock new opportunities for computational imaging that transcends conventional well-capacity constraints. The modulo imaging paradigm presents fascinating new challenges in distinguishing authentic scene structure from artificial wrap discontinuities, a problem that pushes the boundaries of classical phase unwrapping into unexplored territory. Deep learning has emerged as a natural solution, providing advanced pattern-recognition capabilities to resolve ambiguities that traditional optimization methods cannot effectively address. We present complementary approaches leveraging unrolled optimization networks and feature lifting strategies that teach neural architectures to handle wrapped measurements effectiveness. The introduction of scaling equivariance principles enables robust adaptation across varying exposure conditions, while physics-informed input representations guide networks toward meaningful reconstructions. This work addresses fundamental questions about unlimited sampling theory in practical imaging systems, revealing how modulo measurements can codify arbitrarily bright scenes within finite bit depths. The implications extend beyond photography into autonomous systems, scientific imaging, and any application demanding extreme dynamic range. These advances establish computational modulo imaging as a viable pathway toward truly unlimited dynamic range capture, opening new frontiers in computational photography. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-94-lllmstxt|> ## Visual AI for Manufacturing [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Whitepaper # Visual AI for Defect Detection in Manufacturing Can AI-based visual inspection really deliver factory-grade reliability? In this whitepaper, we analyze what breaks most manufacturing defect detection systems—and how top teams overcome the hardest challenges with data-centric strategies, synthetic defects, and scalable QA workflows. Get actionable insights and proven tools for: - **Reducing defect escapes:** Why 65% of manufacturers still rely on error-prone manual inspection - **Preventing model collapse:** How data scarcity and drift silently degrade production AI - **Building robust datasets:** Tools for generating rare defect types with GANs and diffusion - **Scaling quality control:** How Tesla, TSMC, and Samsung use AI to boost yield and cut costs - **Fixing failure modes:** Labeling inconsistencies, generalization gaps, and explainability bottlenecks ![](https://cdn.sanity.io/images/h6toihm1/production/2b9cf0dc8ac97d8cfb20b412b7eec7435249ea1f-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=1600) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-95-lllmstxt|> ## ADT Commercial Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/675fd154dadbf64742785ce336ce4ba3be8722a9-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) ADT FiftyOne helps ADT leverage computer vision for security systems Apr 26, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) ADT provides safe, smart, and sustainable solutions for people, homes, and businesses. Through innovative products, partnerships, and the largest network of smart home, security, and rooftop solar professionals in the United States, ADT empowers its customers to protect and connect what matters most. > " [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) has helped us manage our huge datasets, collaborate on model evaluation, tighten our production schedule, and ultimately deliver solutions that help our customers better manage their risk. FiftyOne Teams has added tremendous value to our computer vision processes." – Philippe Sawaya, Director of Artificial Intelligence at ADT Commercial ![](https://cdn.sanity.io/images/h6toihm1/production/cc3471489867ed253209f00d794f896848866b93-1024x801.webp?auto=format&dpr=2&fit=max&q=75&w=1024) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-96-lllmstxt|> ## Enhancing YOLOv8 Segmentation [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) Enhancing YOLOv8 Segmentation: Precision, Efficiency, and Robustness Apr 3, 2025 • 8 min read Article content In this article [Balancing Speed and Segmentation Accuracy](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness#366c9d16aac3) [Common Pitfalls in YOLOv8 Segmentation](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness#564c471d817a) [Comprehensive Evaluation Beyond Single Metrics](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness#42740d283636) [YOLOv8 Under Real-World Conditions](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness#c91f82e8ae85) [Explore the Jupyter Notebook](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness#3938126b15a4) In this article [Balancing Speed and Segmentation Accuracy](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness#366c9d16aac3) [Common Pitfalls in YOLOv8 Segmentation](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness#564c471d817a) [Comprehensive Evaluation Beyond Single Metrics](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness#42740d283636) [YOLOv8 Under Real-World Conditions](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness#c91f82e8ae85) [Explore the Jupyter Notebook](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness#3938126b15a4) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/878ad4ffe165325bdff6478982b74c1ce9f9aa16-560x155.png?auto=format&dpr=2&fit=max&q=75&w=560) YOLO’s reputation rests on speed: one pass through the network, bounding boxes drawn, and you’re done. However, the moment instance segmentation is required, things can become more complicated. While [YOLOv8](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1/) supports segmentation masks, these masks can be imperfect: parts of objects may disappear, boundaries can bleed together, and performance may degrade from ultra-fast bounding boxes to questionable outlines. Yet, real-time instance segmentation remains compelling. We created [FiftyOne](https://voxel51.com/fiftyone/) to ease the challenges of working with large datasets and segmentation tasks, particularly when iterating quickly with YOLO models. The goal of this article is to show how you can refine YOLO’s masks while preserving the speed that makes YOLO so appealing. We’ve also prepared a companion [Jupyter notebook](https://colab.research.google.com/drive/1a7FoEQdVzdXJbVuZzqAxH2qd8rI7iT9L?usp=sharing) that walks through the data-centric tips outlined below. - Inference with YOLOv8 instance segmentation - Visualizing predictions in FiftyOne - Handling class imbalance with weighted sampling - Synthetic occlusion for more robust masks - Grad-CAM to understand how YOLO “sees” each object ## **Balancing Speed and Segmentation Accuracy** One-stage detectors like YOLO are optimized for bounding boxes, which works well for tasks such as drawing a box around a cat. [Instance segmentation](https://voxel51.com/resources/learn/a-guide-to-ai-image-segmentation/), however, demands identifying each pixel of that dog or cat. For overlapping objects, YOLO’s single-shot approach can begin to stretch thin. Two-stage approaches (Mask R-CNN and others) often produce sharper boundaries but run more slowly. If you work in robotics, real-time safety monitoring, or any domain where heavy compute overhead is not an option, YOLO’s efficiency remains compelling. The key is finding ways to refine YOLO’s segmentation rather than resorting to more resource-intensive solutions. Real-time segmentation is already appearing in autonomous vehicles, advanced medical imaging, and streaming analytics for security cameras. As soon as you ask a single model to generate both bounding boxes and segmentation masks at high speed, you recognize that you cannot simply press “train” and expect perfect results. Although YOLO’s speed is advantageous, data curation, labeling consistency, and a purposeful training strategy have a major role to play. ## **Common Pitfalls in YOLOv8 Segmentation** ### **Inconsistent Data and Bad Labels** [Data quality](https://voxel51.com/blog/data-quality-the-hidden-driver-of-ai-success/) is paramount, and pixel-level labeling invites a wide margin for error compared to bounding boxes. A bounding box can be off by a few pixels without severely harming training, but instance segmentation requires precision around edges. If your annotation tool mislabels portions of the background as part of an object or neglects certain occlusions, YOLO will internalize those inconsistencies. The result could be masks that truncate arms or merge adjacent objects. Before concluding that the model is at fault, inspect your ground truth carefully. FiftyOne proves highly useful for diagnosing such issues by overlaying predicted versus ground-truth masks. By comparing them side by side, you can pinpoint where a label may have been erroneous or an object boundary was drawn incorrectly. ![](https://cdn.sanity.io/images/h6toihm1/production/878ad4ffe165325bdff6478982b74c1ce9f9aa16-560x155.png?auto=format&dpr=2&fit=max&q=75&w=560) This kind of visualization is critical in fields like medical imaging, where a single mislabeled pixel may mark the difference between healthy and diseased tissue. ### **Class Imbalance: A Persistent Challenge** Consider a dataset containing 10,000 dog images and only 200 bird images. YOLO will excel at detecting dogs but struggle with birds. That might be acceptable if detecting dogs is your main objective, but not if both classes matter equally. Weighted sampling, oversampling, or carefully adding more data are all potential remedies. By examining per-class performance in FiftyOne, you can identify if “bird” IoU is inadequate, indicating you may need additional data or class-specific weighting. Weighted loss functions and synthetic data strategies are frequent go-tos for restoring balance so that minority classes do not become an afterthought. **Key Data IssuesData Issue** **Impact on Segmentation** **Mitigation Strategies** _Class Imbalance_ Leads to biased models that perform well on frequent classes but poorly on underrepresented ones; rare classes may be ignored. Use oversampling or undersampling, apply weighted loss functions, or add synthetic examples to balance the class distribution. _Occlusions & Overlaps_ Causes segmentation masks to merge or blur together when objects partially block one another, reducing the accuracy of boundaries. Implement multi-scale training, use synthetic occlusion augmentations (e.g., cutouts or blending), and apply post-processing steps to refine masks. _Labeling Errors_ Inaccurate or inconsistent annotations can cause the model to learn incorrect boundaries, leading to truncated or merged object masks. Rigorously review and clean data annotations, use tools like FiftyOne to catch mislabels, and perform quality assurance on the dataset before training. ## **Comprehensive Evaluation Beyond Single Metrics** It’s easy to focus exclusively on a metric like mean Intersection over Union (mIoU) or mean Average Precision (mAP), as these are convenient to quote. However, when you need refined YOLO models for mission-critical or high-stakes applications, you must [look more deeply](https://docs.voxel51.com/tutorials/evaluate_detections.html) into your metrics. ### **Uncertainty Analysis** If your model is overly confident in every prediction, it’s either exceptionally robust or inflating its certainty. By examining predictions with low confidence, you can locate where YOLO’s coverage may be insufficient. FiftyOne can highlight these low-confidence regions so you can diagnose whether the issue arises from sparse data, class confusion, or unusual occlusions. By fixing those subsets, retraining, and iterating, you can progressively refine your model. ### **Edge Cases** Investigate the relatively few images that produce poor IoU. These may share unusual factors, such as very dark lighting, extreme angles, or challenging objects placed near image boundaries. Focusing your [data augmentation](https://voxel51.com/blog/data-augmentation-is-still-data-curation/) on these edge cases can yield considerable improvements. If YOLO repeatedly struggles with small objects near the frame’s corner, multi-scale training or curated examples of small-corner objects may be valuable. The more lightweight YOLO variants, like YOLOv8n, facilitate fast experimentation so that you can adapt quickly to new findings. **Optimizing YOLOv8 for Superior Segmentation** ### **Speeding Up Iterations** No one wants to spend a week waiting for a single training run to finish before determining whether a particular augmentation strategy worked. Smaller YOLOv8 “nano” models are a practical option for fast experimentation. When a method shows promise, you can scale up to a larger model for those final increments in accuracy. This strategy saves both time and GPU resources. ### **Clean Up Labels with FiftyOne** Whenever you notice strange predictions, investigate whether labels might be inaccurate. It’s not unusual to discover entire objects labeled as background or partially trimmed. With FiftyOne’s visual interface, you can directly compare predicted masks to your ground-truth annotations and resolve labeling inconsistencies. Even a well-planned training workflow will replicate errors if many segmentation masks in your dataset are off by a few pixels. ### **Addressing Class Imbalance** We cannot overemphasize how challenging class imbalance can be. If your dataset features wide disparities between classes, the underrepresented classes risk being overlooked. Weighted losses or oversampling are straightforward techniques in PyTorch (with WeightedRandomSampler, for instance) that ensure each class receives due attention. While it may sound basic, it can greatly reduce the “this class never shows up” problem. ### **Improving Robustness with Synthetic Occlusions** Real-world scenes are usually messy, with objects partially hiding behind one another. Introducing random black or blurred patches, overlays, or other forms of partial occlusion in your training images helps YOLO learn that an object can remain identifiable even when partially obscured. Although initial results may dip, the model generally adapts, leading to fewer catastrophic merges when real-world occlusions occur. ![](https://cdn.sanity.io/images/h6toihm1/production/878ad4ffe165325bdff6478982b74c1ce9f9aa16-560x155.png?auto=format&dpr=2&fit=max&q=75&w=560) ### **Check YOLO’s Brain with Grad-CAM** It can be insightful to discover which parts of an image YOLO relies on when generating segmentation masks. [Grad-CAM](https://voxel51.com/blog/exploring-gradcam-and-more-with-fiftyone/) (or Eigen-CAM) can highlight the critical regions that determine YOLO’s outputs. Sometimes, the model focuses exactly where you’d expect; other times, it may attend to irrelevant background details. These insights can help you decide whether to refine your data or detect any unintended correlations YOLO is learning. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop **Putting This into Action** At Voxel51, we developed FiftyOne to [streamline](https://docs.voxel51.com/teams/query_performance.html) these processes: - Overlay predicted masks and ground-truth masks - Filter by class or confidence level - Tag anomalous edge cases or labeling errors - Evaluate performance through detailed per-class and per-sample metrics ![](https://cdn.sanity.io/images/h6toihm1/production/fec7b23838b23cb94088f2ba8618011646ca544c-1203x528.png?auto=format&dpr=2&fit=max&q=75&w=1203) ## **YOLOv8 Under Real-World Conditions** Segmentation models don’t always encounter ideal conditions. Real-world scenarios with dim lighting, blurry cameras, and noisy environments can quickly challenge even highly accurate models. The four images below demonstrate YOLOv8’s segmentation performance under common image degradations: original, dimmed, noisy, and blurred. Interestingly, notice how YOLOv8 predictions shift with each distortion. For example, under dimmed conditions, detection confidence dips across objects, reflecting the model’s uncertainty without strong visual cues. Noise further erodes accuracy, reducing confidence sharply, while blur slightly improves predictions for some of the objects compared to dimming, even if it still mislabels clasped hands as a remote. This visual exploration underscores the need to expose YOLOv8 models to realistic, degraded conditions during training. Incorporating augmentation strategies like slight blurring, brightness variations, and mild noise can greatly boost YOLO’s robustness in real-world applications, maintaining the fast performance users expect. ![](https://cdn.sanity.io/images/h6toihm1/production/e5bc47b5671632458a8df63e1a8099c35e12689a-1831x910.png?auto=format&dpr=2&fit=max&q=75&w=1600) **Streamlining Segmentation Workflows with FiftyOne** FiftyOne removes the guesswork from your YOLO segmentation model runs, allowing you to perform instance segmentation without manually sorting through large volumes of predictions. Quickly identify issues with your custom dataset, refine your labels, and streamline your workflow when you [train YOLOv8](https://docs.voxel51.com/tutorials/yolov8.html) instance segmentation models, even lightweight versions like the nano model. Continual refinement and iterative retraining help you identify individual objects and generate accurate segmentation masks, transforming a rough dataset into a structured resource. Even if your use case is robotics, medical imaging, or everyday object detection, FiftyOne enables a detailed understanding of your data and improves real-time performance. By addressing common challenges in instance segmentation tasks, you can confidently deploy your fully trained model to reliably segment objects in real-world conditions. ## **Explore the Jupyter Notebook** Check out the [Jupyter notebook](https://drive.google.com/open?id=10s7Wx7lvmpxajgmw1M1-TNEorm9aU9iu) that walks through the data-centric tips discussed above. By following the notebook, you can apply these techniques to your dataset, experiment with different augmentations, and inspect how YOLOv8 behaves under challenging conditions. **Image Citations** - Carine06. _Ed Corrie & Kevin Anderson Shaking Hands After a Match._ Photograph. Wikimedia Commons. CC BY-SA 2.0. [https://commons.wikimedia.org/wiki/File:Handshake\_(27106664813).jpg](https://commons.wikimedia.org/wiki/File:Handshake_(27106664813).jpg). Sedlecký, David. _Václav Lebeda (Voxel), Czech Musician_. Photograph. June 9, 2015. Wikimedia Commons. CC BY-SA 4.0. [https://commons.wikimedia.org/wiki/File:V\_Lebeda\_Voxel\_2015.JPG](https://commons.wikimedia.org/wiki/File:V_Lebeda_Voxel_2015.JPG). [YOLOv8](https://voxel51.com/blog/tag/yolov8) [segmentations](https://voxel51.com/blog/tag/segmentations) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/9252e8bb5db5c4805f4a6f315b51527ee5c22072-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Guide to AI Image Segmentation\\ \\ Learn\\ \\ • \\ \\ Dec 19, 2024](https://voxel51.com/blog/a-guide-to-ai-image-segmentation) [![](https://cdn.sanity.io/images/h6toihm1/production/c845648a6e64e00a3cac3d875280cb462c63633d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Complete Guide to Auto Labeling\\ \\ Learn\\ \\ • \\ \\ Jun 16, 2025](https://voxel51.com/blog/the-complete-guide-to-auto-labeling) [![](https://cdn.sanity.io/images/h6toihm1/production/ca4f84addae3f8e97daad00cf856754301cbb5e0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Are Image Segmentation Maps Superior to Bounding Boxes?\\ \\ Learn\\ \\ • \\ \\ Feb 26, 2025](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-97-lllmstxt|> ## Understanding Good Data [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) What Makes ‘Good’ Data? A View from the Front Lines of AI Jul 17, 2025 • 4 min read Article content In this article [From open source code to open source data](https://voxel51.com/blog/what-makes-good-data-a-view-from-the-front-lines-of-ai#1daf9d20c80e) [Why we built FiftyOne for data understanding](https://voxel51.com/blog/what-makes-good-data-a-view-from-the-front-lines-of-ai#e35ecb083a9c) [What makes data “good”?](https://voxel51.com/blog/what-makes-good-data-a-view-from-the-front-lines-of-ai#925055bd54e7) [Ready to see your data clearly?](https://voxel51.com/blog/what-makes-good-data-a-view-from-the-front-lines-of-ai#0bfa107f59a3) In this article [From open source code to open source data](https://voxel51.com/blog/what-makes-good-data-a-view-from-the-front-lines-of-ai#1daf9d20c80e) [Why we built FiftyOne for data understanding](https://voxel51.com/blog/what-makes-good-data-a-view-from-the-front-lines-of-ai#e35ecb083a9c) [What makes data “good”?](https://voxel51.com/blog/what-makes-good-data-a-view-from-the-front-lines-of-ai#925055bd54e7) [Ready to see your data clearly?](https://voxel51.com/blog/what-makes-good-data-a-view-from-the-front-lines-of-ai#0bfa107f59a3) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) For much of the last decade, the prevailing narrative in AI has been that scale wins. More data, more compute, larger models—that has been the formula. And to be fair, we’ve gone pretty far with it. But in my experience—both in academia and industry—this emphasis on quantity has come at the cost of a more nuanced truth: it’s not just how much data you have. It’s how well you understand and curate that data. As a computer vision and machine learning researcher, I’ve spent years working on models to interpret the visual world. But over time, I kept running into the same friction point: the data itself. Was it representative? Was it biased? Was it even usable? And perhaps most importantly—how would I even know? ## **From open source code to open source data** The open source mindset has always been part of my work—long before it became mainstream in machine learning. As a graduate student, most of us spent countless hours re-implementing algorithms from papers, line by line, because authors rarely shared their code. That made replication slow and sometimes frustrating. But it also reinforced how important transparency and reproducibility was if we wanted the field to move forward. In the early 2000s, that started to change. More researchers began publishing MATLAB code alongside their papers, and suddenly it became easier to reproduce, test, critique, and build on each other’s work. It wasn’t just more efficient—it helped advance the broader field faster and more dynamically. That shift convinced me to make open source code a requirement once I became a professor with my own research lab. If we published a paper, we released the code and instructions to reproduce the results. It was a simple rule with a big impact. But after years of doing that, something still felt incomplete. We were sharing our code. But the data behind the results—the decisions we made in collecting, filtering, labeling, or cleaning it—were rarely as visible. Around 2003, more open computer vision datasets, such as Caltech 101 and the KTH action dataset, were created, setting the precedent that data sharing was also viable. As datasets became increasingly available and larger, it was a huge challenge to understand and analyze them. The tools to inspect or understand them just didn’t exist. In a world where models were only getting more complex and data more central, that seemed like a blind spot we couldn’t afford to ignore. ## Why we built FiftyOne for data understanding Around the time that large-scale deep learning started reshaping the field, something else changed: the role of data shifted from supporting cast to co-star. It became clear that the performance of a model wasn’t just about architecture or training tricks—it was fundamentally tied to the quality and characteristics of the data itself. That shift was both exciting and disorienting. We were training increasingly powerful models, but often without understanding _why_ they worked—or didn’t. I saw it again and again in my own research and in conversations with other engineers and scientists. The data might be noisy, imbalanced, redundant, or just a poor fit for the task, but we didn’t have tools to diagnose that. It felt like flying blind. This lack of visibility sparked the idea that maybe we needed to rethink the interface between humans and machine learning systems—not at the level of models, but at the level of data. What if you could ask questions of your dataset the way you’d inspect model logs or weight distributions? What if developers had the same observability for their training data that they have for their model architectures? That’s what we set out to build with the founding of Voxel51: tools like [FiftyOne](https://voxel51.com/fiftyone) that allow machine learning engineers to inspect, slice, visualize, and experiment with image and video datasets in meaningful ways. Not because it’s trendy, but because understanding your data is essential if you want your model to generalize, behave ethically, or even just to work at all. ## **What makes data “good”?** There’s no universal checklist for “good” data—it depends entirely on context. But in practice, we’ve found that a few patterns come up again and again: redundancy, class imbalance, mislabeled examples, and poor edge-case coverage, to name a few. These aren’t just academic concerns—they’re the reason models fail in production. What makes data “good” is whether it’s right for the problem you’re trying to solve. That might mean reducing duplicates, balancing class distributions, or specifically _not_ balancing them if your use case demands it. Sometimes it means discovering the rare, hard-to-label examples that matter most. The point is, you can’t fix what you can’t see. Data observability is the prerequisite for actionable improvement. In machine learning, we’ve made incredible progress on modeling techniques, architectures, and tooling. But models are only as good as the data they learn from. And understanding that data—truly analyzing it, questioning it, improving it—is still one of the most underdeveloped parts of the workflow. We don’t need more data. We need better ways to work with the data we already have, and that starts with data-centric AI practices. If we want to build models that are not just accurate, but robust, fair, and reliable, then investing in data understanding isn’t optional—it’s foundational. ## **Ready to see your data clearly?** [Try FiftyOne today](https://voxel51.com/sales) and discover how faster dataset visualization and curation can unlock the next level of performance for your visual-AI models. [dataset improvement](https://voxel51.com/blog/tag/dataset-improvement) ![](https://cdn.sanity.io/images/h6toihm1/production/ad9fb967c5455e0f763411fb81956767d7f26482-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Jason Corso Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/6af33def6d297e2382d387e224e16451c95876af-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Why the “Annotate Everything” Era in Automotive AI Is Over\\ \\ Computer Vision\\ \\ • \\ \\ Jul 24, 2025](https://voxel51.com/blog/smarter-automotive-datasets-selection) [![](https://cdn.sanity.io/images/h6toihm1/production/65dd49021e60c3c8d3f0982f888781871ba4ed26-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ From Prototype to Production: What it Really Takes to Deploy Computer Vision in Manufacturing\\ \\ Computer Vision\\ \\ • \\ \\ Aug 6, 2025](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing) [![](https://cdn.sanity.io/images/h6toihm1/production/1798d34efc8956a6696377fb7776886ec0e61092-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ UnCommon Objects in 3D\\ \\ Datasets, Event Recaps\\ \\ • \\ \\ Jun 18, 2025](https://voxel51.com/blog/uncommon-objects-in-3d) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-98-lllmstxt|> ## Building GUI Agents Workshop [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/2bc63f0cc384ef11ad9a6fbf2b11217505725ec2-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/2bc63f0cc384ef11ad9a6fbf2b11217505725ec2-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event Virtual Americas Webinars & Workshops From Research to Reality: Building GUI Agents That Actually Work - August 29, 2025 Aug 29, 2025 9 AM Pacific Online. Register for the Zoom! About this event Welcome to the Visual Agents Workshop Series, your virtual pass to learn about visual agents - how they work, how to develop them and how to fine-tune them. Host ![](https://cdn.sanity.io/images/h6toihm1/production/c96544cfe7c8fc1e23601a34ca6a5fc11ccd6aa5-320x320.png?auto=format&dpr=2&fit=max&q=75&w=96) Harpreet Sahota Voxel51 Bio ### Part 3: Teaching Machines to See and Click - Model Finetuning From Foundation Models to GUI Specialists Foundation models, such as Qwen2.5-VL, demonstrate impressive visual understanding, but they require specialized training to master GUI interactions. In this final session, you'll transform a general-purpose vision-language model into a GUI specialist that can navigate interfaces with human-like precision. We'll explore modern fine-tuning strategies specifically designed for GUI tasks, from selecting the right architecture to handling the unique challenges of coordinate prediction and multi-step reasoning. You'll implement training pipelines that can handle the diverse formats and platforms in your dataset, evaluate models on metrics that actually matter for GUI automation, and deploy your trained model in a real-world testing environment. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-99-lllmstxt|> ## Motion Control for Video [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Motion Prompting: Generalized Motion Control for Video Generation Jun 26, 2025 • 2 min read Article content In this article [Method Overview](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#20ae18134590) [Key Applications](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#c3434eb4c3ef) [Limitations](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#f2910b72caae) [Relevance to FiftyOne](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#0a70e7a6175d) [Conclusion](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#e8b30c960a6b) [What is next?](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#d402ba7ff4ef) In this article [Method Overview](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#20ae18134590) [Key Applications](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#c3434eb4c3ef) [Limitations](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#f2910b72caae) [Relevance to FiftyOne](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#0a70e7a6175d) [Conclusion](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#e8b30c960a6b) [What is next?](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation#d402ba7ff4ef) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _**CVPR 2025 Insights \#3:** Learning from Movement. [Paper](https://openaccess.thecvf.com/content/CVPR2025/papers/Geng_Motion_Prompting_Controlling_Video_Generation_with_Motion_Trajectories_CVPR_2025_paper.pdf) Presented at CVPR 2025 Oral Session \| Poster \#173_ Recent advances in video generation have introduced methods for text-to-video and region-specific control. But what if we could condition a video model on any motion using a unified, intuitive representation? Motion Prompting does precisely that. This work introduces a method to control AI-generated videos using point trajectories, user-defined paths over space and time, enabling general and flexible motion conditioning. ## Method Overview - The team fine-tunes Lumiere, a base video diffusion model, with a ControlNet that accepts a rasterized space-time volume of point tracks. - Tracks are extracted using video tracking algorithms and embedded with 64D vectors acting as unique identifiers. - It supports arbitrary density, duration, and location of motion signals, which are far more general than bounding boxes or sparse keypoints. ![](https://cdn.sanity.io/images/h6toihm1/production/20d4eb734323f94c9bc0fdc030974d9125462bde-1342x1328.webp?auto=format&dpr=2&fit=max&q=75&w=1342) ## Key Applications - **Interactive Video Editing:** Click-and-drag input turns still images into dynamic videos with localized, consistent motion. - **Camera & Object Control:** Depth-based point clouds allow synthetic camera movement (e.g., dolly zooms). - **Motion Transfer:** Animate new images with motion from a reference video. - **Motion Magnification:** Subtle motions like breathing are amplified by scaling point trajectories. ## Limitations - Bidirectional generation leads to non-causal effects (e.g., motion anticipation). - Ambiguities in overlapping motion regions may produce unintended results. - Requires ~10 minutes per video; real-time interaction is still under exploration. ## Relevance to FiftyOne While **Motion Prompting** does not directly intersect with FiftyOne, there are clear synergy points: - Point trajectory data (input/output) could be analyzed, visualized, or labeled using FiftyOne’s spatial-temporal tools. - Motion transfer outputs might benefit from frame-by-frame evaluation, error diagnosis, or comparative visualization. - Future work could integrate trajectory-based interaction logs into FiftyOne for dataset curation or model debugging. ## Conclusion “Motion Prompting” introduces a scalable, general-purpose way to control video synthesis via point tracks, unlocking a wide range of editing and interactive generation capabilities. It’s a valuable contribution for researchers in **video synthesis, human-computer interaction, and creative AI**. ## What is next? If you’re interested in following along as I dive deeper into the world of AI and continue to grow professionally, feel free to connect or follow me on [LinkedIn](https://www.linkedin.com/in/paula-ramos-phd/). Let’s inspire each other to embrace change and reach new heights! You can find me at some Voxel51 events ( [https://voxel51.com/computer-vision-events/](https://voxel51.com/events)), or if you want to join this fantastic team, it’s worth taking a look at this page: [https://voxel51.com/jobs/](https://voxel51.com/careers) [video](https://voxel51.com/blog/tag/video) [generative AI](https://voxel51.com/blog/tag/generative-ai) [CVPR](https://voxel51.com/blog/tag/cvpr) ![](https://cdn.sanity.io/images/h6toihm1/production/e926c07c7d1426c0fde8fdefa637c528d47b16f4-512x512.webp?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a558b86370f2f17212fb2f2c894d590101458a85-5760x3241.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Multimodal Frontier in Computer Vision, Medicine, and Agriculture— CVPR 2025 Reflections\\ \\ Event Recaps, Industry Solutions\\ \\ • \\ \\ Jun 24, 2025](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections) [![](https://cdn.sanity.io/images/h6toihm1/production/97cf3d887735ab9574b2b3e2d3825146016d875e-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Embodied Computer Vision at CVPR 2025: The Next AI Frontier\\ \\ Event Recaps\\ \\ • \\ \\ Jun 30, 2025](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier) [![](https://cdn.sanity.io/images/h6toihm1/production/8bf3c50a5edd9f83e1013ad5df86ae159d519dae-1200x626.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Best of CVPR 2025: Conversations at the Cutting Edge of AI\\ \\ Event Recaps\\ \\ • \\ \\ Jul 3, 2025](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-100-lllmstxt|> ## Auto Labeling Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) The Complete Guide to Auto Labeling Jun 16, 2025 • 10 min read Article content In this article [The Annotation Challenge, and the AI-Powered Solution](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#8722180bb8c7) [Selecting the Right Foundation Model for Your Task](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#a8f6b801d4f7) [Setting Confidence Thresholds: Balancing Precision and Recall](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#acd9e6e97fa3) [Quality Assurance of Auto-Labels](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#d2fd983a693e) [Training Downstream Inference Models on Auto-Labels](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#e29394eef17b) [Conclusion: Evaluating Annotation Effort Trade-offs](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#99a5eabd36bb) In this article [The Annotation Challenge, and the AI-Powered Solution](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#8722180bb8c7) [Selecting the Right Foundation Model for Your Task](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#a8f6b801d4f7) [Setting Confidence Thresholds: Balancing Precision and Recall](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#acd9e6e97fa3) [Quality Assurance of Auto-Labels](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#d2fd983a693e) [Training Downstream Inference Models on Auto-Labels](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#e29394eef17b) [Conclusion: Evaluating Annotation Effort Trade-offs](https://voxel51.com/blog/the-complete-guide-to-auto-labeling#99a5eabd36bb) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Manually labeling images is one of the biggest bottlenecks in developing computer-vision models. Automated data labelling, using AI to annotate data, promises to drastically reduce this cost and delay. Recent advances in foundation models have made automatic data labeling surprisingly effective, and [new approaches](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance) can rival human annotation at a fraction of the effort and cost. In this guide to automated annotation, we’ll first understand the labeling challenge. We’ll then walk through how to automate data labelling with foundation models, learn how to tune and validate an auto-labeling pipeline, and discuss best practices to get the most from an automated annotation tool. We’ll cover model selection (classification, detection, segmentation), confidence-threshold tuning, quality-assurance (QA) workflows (using FiftyOne’s embeddings and model-evaluation panels), and how to balance precision, recall, and annotation effort. By the end, you’ll have a conceptual how-to for leveraging annotation automation in your ML workflow. ## The Annotation Challenge, and the AI-Powered Solution In modern machine learning, great labels are necessary for building great models. Yet traditional hand-labeling is costly and time consuming. For years, the prevailing wisdom was that more _human_-labeled data yields better models, fueling a huge human labeling industry. However, this approach doesn’t scale well: adding human labels for every new scenario or edge case is expensive, particularly in domains requiring specialized expertise. At the same time, the rise of vision-language foundation models (VLMs) offers a new opportunity. Models like YOLO-World and Grounding DINO are pretrained on enormous datasets and already“know” how to identify many visual objects. The question is can we then leverage pretrained models to label new data automatically, and also ensure these auto-generated labels are accurate and useful for training downstream models? The latest tools aim to have a model _entirely_ annotate a dataset in a zero-shot manner and make the results good enough to train production models. In auto labeling, we can use foundation models to [generate “pseudo-ground-truth” labels](https://arxiv.org/abs/2506.02359) for unlabeled data. Using a tool like FiftyOne’s [Verified Auto Labeling](https://voxel51.com/annotation), a foundation model predicts labels (with confidence scores) for each image, and after verifying their usefulness, those predictions become your training labels. ## Selecting the Right Foundation Model for Your Task Choosing the right model for auto-labeling depends on the task type, i.e., classification or detection or segmentation, as well as your requirements for accuracy, speed, and coverage. Here are some general guidelines: ### Image Classification For image classification, consider zero-shot classification models. A standard choice is [OpenAI’s CLIP](https://voxel51.com/blog/a-history-of-clip-model-training-data-advances), which predicts if an image contains a certain concept by comparing embeddings of the image and text labels. Other transformer-based classifiers (e.g. BiT, ViT, etc.) fine-tuned on large datasets can also serve as auto-labelers for classification. Ultimately you need to choose a model whose vocabulary covers your domain. For instance, if labeling everyday objects, a CLIP-based approach might work out-of-the-box. If labeling medical images, a specialty model pretrained on medical data might be necessary. ### Object Detection For detecting and labeling multiple objects per image with bounding boxes, vision-language models that support open-vocabulary detection are ideal. [YOLO-World](https://docs.ultralytics.com/models/yolo-world/) is a real-time open-vocabulary detector built on YOLO with a CLIP-like text encoder that can efficiently find objects based on arbitrary text prompts. It was shown to label simpler datasets like PASCAL VOC in mere minutes with impressive accuracy. [Grounding DINO](https://huggingface.co/docs/transformers/en/model_doc/grounding-dino) is also a powerful open-vocabulary detector known for high accuracy, but it’s heavier. It can take orders of magnitude longer on large datasets and requires more GPU memory. If speed and scale are priorities (e.g. annotating millions of images), a faster model like YOLO variants might be preferable. If you need higher precision on a broader or more complex set of classes (especially with descriptive labels), a model like Grounding DINO or Google’s [OWL-ViT](https://huggingface.co/docs/transformers/en/model_doc/owlvit) could be better despite the slower speed. Also consider class granularity and domain. If your project involves very fine-grained classes or a long-tail distribution (like identifying specific animal species or products), check if the foundation model can handle it. In such cases, a hybrid approach might be needed – e.g. use auto-labeling for the common classes and supplement with human labels for the rare ones. ### Segmentation For segmentation masks, the leading foundation model is [Meta’s Segment Anything Model (SAM)](https://segment-anything.com/). SAM can generate segmentation masks for any object in an image given minimal prompts (points or boxes). An automated annotation strategy for segmentation could be to first use a detector (like Grounding DINO) to find object regions and then apply SAM to get exact masks for those regions. This two-step approach can automatically produce segmentation labels: the detector proposes _what_ and _where_, and SAM delineates the exact shape. If class labels are needed for each mask, you might still need a classification step. Either the detector provides the class from text prompt, or use an image classifier on the masked region. There are also emerging one-shot segmentation models and fully open-vocabulary segmentation models, but those are less mature. In practice, a combination of detection + segmentation model works. While more involved, this is still programmatic data labeling done in minutes, not requiring a human to trace polygons. If you are doing segmentation auto-labeling, verify the model you choose can capture the level of detail you need. SAM is very general, but on highly domain-specific structures (like medical imagery), a domain-specific model might be necessary for best results. ### Model-Selection Trade-offs Accuracy, speed, scalability, and compatibility all need to be balanced. If your dataset is huge (millions of images or more), a slightly less-accurate but orders of magnitude faster model could actually yield better results overall because you can label _all_ your data instead of timing out on half of it. Also consider memory and deployment: some open-vocabulary models have quirks (e.g. Grounding DINO had memory issues when given very long class lists like LVIS). A practical tip is to start with a fast model to get an initial set of labels, then perhaps re-label a subset of data with a more powerful model for classes that were missed or low-quality. FiftyOne makes it easy to swap in different models thanks to its integration with libraries like [Ultralytics](https://docs.voxel51.com/integrations/ultralytics.html) and [Hugging Face Transformers](https://docs.voxel51.com/integrations/huggingface.html). ## Setting Confidence Thresholds: Balancing Precision and Recall Choosing a confidence threshold is one of the most important configuration choices in auto labeling. This threshold determines how sure a model’s prediction must be to accept it as a label. In object-detection tasks a model might output 50 candidate boxes with confidence scores ranging from 0.1–0.99. If the threshold might is set at 0.5, any prediction below 50% confidence is discarded and only predictions ≥ 0.5 become auto labels. Intuitively, you might think “the higher the threshold, the cleaner the labels.” Indeed, higher thresholds give higher precision (fewer false-positive labels). However, there is a catch: too high a threshold _dramatically_ lowers recall. The model may not label many true objects that it was less confident about, which can harm the final trained model’s performance. [Voxel51’s research](https://arxiv.org/abs/2506.02359) suggests that ultra-high-confidence auto-labels (e.g. 0.8 – 0.9+) actually led to worse downstream model performance than using moderately confident labels. The sweet spot observed was thresholds in the 0.2 – 0.5 range, which provided a good balance of precision and recall and yielded the highest mAP when training models on the auto-labeled data. In practice, a recommended strategy is to start with a moderately low threshold (≈ 0.3 **)** for initial labeling. This prioritizes high recall. Then use QA processes to clean up the false positives (we’ll discuss how in the next section). This aligns with the idea that _“clean labels aren’t always better”_ if achieving them means sacrificing too much recall. FiftyOne also provides tools such as an [“Optimal Confidence Threshold” plugin](https://voxel51.com/blog/finding-the-optimal-confidence-threshold) that can scan a range of thresholds and find which yields the best F1 against ground truth for a given model. ## Quality Assurance of Auto-Labels Even after choosing an appropriate model and threshold, the auto labels will not be perfect. Remember that a moderate threshold results at higher recall at the expense of more false positives. This is deliberate, as it’s easier to QA incorrect labels than hunt for false negatives, or objects that _should_ have been labeled among the data that were not. FiftyOne provides some powerful visualization tools to make this QA process efficient. There’s no singular prescription. Organizations will need to choose the workflow that works best for their team, data description, and quality of auto labels in their data set. But below are a few best practices. - **Focus on Low-Confidence Predictions:** Low confidence predictions (for example, ɑ < 0.3) are prime candidates for review. Verified Auto Labeling lets you create a filter view of them automatically with the confidence slider. - **One-Click Accept/Reject:** During initial label generation, FiftyOne lets you batch labels for review, and then reviewers can batch-approve or discard labels. - **Use the Embeddings Panel to Spot Anomalies:** FiftyOne computes object embeddings and lets you lasso outliers that can often indicate mislabels. - **Leverage Similarity Search:** After finding one mistake, [search for visually similar samples](https://docs.voxel51.com/brain.html#similarity) to identify trends and bulk-fix. - **Double-Check Edge Cases:** Use FiftyOne’s Data Quality workflow to sort/filter by metadata (scene type, brightness, etc.) and find model blind-spots. ## Training Downstream Inference Models on Auto-Labels After QA, train your downstream ML model on the verified auto-labels. Voxel51’s research showed models trained _solely on auto-labeled data_ achieved 90–95 % of the accuracy of models trained on human-labeled data in many cases. On challenging datasets like LVIS, the gap between was larger. In those cases, a hybrid approach of using manual labels for the hardest 5-10% of samples is recommended.The steps might look like the following. 1. Auto-label the entire dataset to maximize recall, understanding that there may be some mistakes in long-tail classes. 2. Manually label 5-10%of samples. These are typically rare classes, edge-case conditions, or high-risk scenarios. Selecting those samples is easy in FiftyOne: 1. Filter by low auto-label F1 or low sample-confidence. 2. Use the embeddings panel to surface outliers or sparsely represented clusters. 3. FiftyOne integrates with common annotation tools like [CVAT](https://docs.voxel51.com/integrations/cvat.html). You can also [create labels directly](https://docs.voxel51.com/user_guide/annotation.html) with the FiftyOne SDK. 3. [Union the two label sets](https://docs.voxel51.com/user_guide/using_views.html#editing-fields) (auto + human labels) and train your inference model. Even a few hundred high-quality human labels can close most of the long-tail gap, while still minimizing overall annotation costs. You can then evaluate the model with FiftyOne’s [Model-Evaluation Panel](https://voxel51.com/blog/unified-model-insights-with-fiftyone-model-evaluation-workflows) to examine precision/recall, F1 score, and confusion matrices. The Scenario Analysis tab also provides per-class metrics and sample-level errors to show necessary context to fix labels or adjust thresholds for retraining. Here’s how a model eval workflow might look. 1. [Load model predictions into FiftyOne](https://docs.voxel51.com/user_guide/dataset_creation/#model-predictions) and open the Model-Evaluation Panel to compute metrics like precision, recall, F1, and mAP. 2. Inspect confusion matrices to identify systematic mix-ups. Clicking any cell to drill into the exact images affected, where you can bulk-tag them for relabeling if needed. 3. Plot precision-recall curves [for the overall model](https://docs.voxel51.com/api/fiftyone.core.plots.matplotlib.html#fiftyone.core.plots.matplotlib.plot_pr_curves) and [per class](https://docs.voxel51.com/api/fiftyone.core.plots.matplotlib.html#fiftyone.core.plots.matplotlib.plot_pr_curves). These curves inform the confidence threshold you’ll deploy for future auto-labeling iterations. For examples, ɑ = 0.35 for the inference model may maximize F1 score even if the model used to apply auto labels was set to ɑ = 0.25. 4. Drill into sample-level errors. That is, sort by number of false negatives to identify corner cases like occlusions or unusual lighting. Feed those images back into the auto-label loop. 5. Continue iterating with small patches of relabeling or confidence threshold tuning. ## Conclusion: Evaluating Annotation Effort Trade-offs Automated data labeling helps you annotate and build vision models at a scale and speed that was previously impossible. By carefully choosing foundation models, tuning confidence thresholds, and using intelligent QA workflows, you can obtain training data that is nearly as good as human-labeled in a fraction of the time and cost. Tools like FiftyOne’s Verified Auto Labeling feature and visualization panels are key to making this approach practical, allowing you to integrate model predictions with human insight efficiently. The result is a complete tool that augments your annotation process with focused human effort where it matters most. [auto-labeling](https://voxel51.com/blog/tag/auto-labeling) [annotation](https://voxel51.com/blog/tag/annotation) [segmentations](https://voxel51.com/blog/tag/segmentations) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/9252e8bb5db5c4805f4a6f315b51527ee5c22072-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Guide to AI Image Segmentation\\ \\ Learn\\ \\ • \\ \\ Dec 19, 2024](https://voxel51.com/blog/a-guide-to-ai-image-segmentation) [![](https://cdn.sanity.io/images/h6toihm1/production/e54b1e5afaaede40db246a7681a611856ebd8782-2260x1268.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Enhancing YOLOv8 Segmentation: Precision, Efficiency, and Robustness\\ \\ Learn\\ \\ • \\ \\ Apr 3, 2025](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness) [![](https://cdn.sanity.io/images/h6toihm1/production/ca4f84addae3f8e97daad00cf856754301cbb5e0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Are Image Segmentation Maps Superior to Bounding Boxes?\\ \\ Learn\\ \\ • \\ \\ Feb 26, 2025](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-101-lllmstxt|> ## Ai.Fish Fisheries Management [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/d2cffb30499c8bfdf723500dbde020b5b834199d-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Ai.Fish Ai.Fish revolutionizes fishery management with FiftyOne Apr 11, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Ai.Fish](https://www.ai.fish/) is creating fisheries of the future by using computer vision and AI to help the commercial fishing industry reduce costs and improve efficiency of their electronic monitoring operations—with the goal of driving sustainable fishing practices and marine conservation. > "Ai.Fish applies automated computer vision annotation and analysis to video footage to deliver the fishery management of the future, today. FiftyOne is a key technology that helps Ai.Fish curate higher volumes of quality data faster, leading to more accurate models." – Jimmy Freese, CEO at AI.Fish ![](https://cdn.sanity.io/images/h6toihm1/production/021ed701e35df3e49fb87c1a8b798e223019d1a5-1161x720.jpg?auto=format&dpr=2&fit=max&q=75&w=1161) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-102-lllmstxt|> ## AI Assistant for Data [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # VoxelGPTAI Assistant for Computer Vision VoxelGPT is an open source plugin for FiftyOne that translates your natural language prompts into actions that organize and explore your data. [Try in browser](https://try.fiftyone.ai/datasets) [Install it locally](https://github.com/voxel51/voxelgpt) Use Cases ## Use chat to interact with your data and models ### Explore your data using natural language, not code Ask VoxelGPT to search your image, video, and 3D point cloud data to find what you’re looking for, no matter how complex the request. _“Show me the most unique images with a false positive prediction.”_ _“Retrieve just the samples with at least two people holding suitcases.”_ _“Give me 10 random samples taken at night and in the snow.”_ ### Search documentation, tutorials and API​ VoxelGPT has access to the entire collection of FiftyOne documentation, including tutorials, user guides, and API docs, making it easy to get instant answers to your questions as you work. _“How do I load my custom dataset into FiftyOne?”_ _“How can I export my dataset in COCO format?”_ _“How can I generate a 2D image for a point cloud?”_ ### Get answers to complex machine learning problems VoxelGPT can answer questions about computer vision, machine learning, and data science while you’re working in FiftyOne. _“What is the difference between precision and recall?”_ _“How can I detect smiling faces in my images?”_ _"How can I reduce redundancy in my dataset?”_ ## Onboard your AI assistant today VoxelGPT is open source and easy to use in your browser or install locally. Get started in seconds and experience the future of dataset management. [Try in browser](https://try.fiftyone.ai/datasets) 0M+ Installs of FiftyOne 0K+ GitHub stars 0K+ Meetup members [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-103-lllmstxt|> ## Car Damage Detection Workshop [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/b510b86e5cba7fb38fcc4ed274f8f0d9e8776979-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/b510b86e5cba7fb38fcc4ed274f8f0d9e8776979-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) In-person EMEA Webinars & Workshops Advanced Car Damage Detection with FiftyOne and the CarDD Dataset - July 12 Jul 12, 2025 10:00 AM - 16:00 PM Saarland University Saarbrücken Campus 66123 Saarbrücken Building E1.3, Foyer Germany About this event This hands-on workshop explores computer vision techniques for automotive damage detection using the CarDD dataset - the largest public dataset specifically designed for vehicle damage analysis. Participants will learn how to leverage FiftyOne, a powerful computer vision experimentation platform, to explore, visualize, and build models for detecting six common types of vehicle damage: dents, scratches, cracks, glass shatter, tire flats, and broken lamps. The workshop consists of five comprehensive modules that guide participants from initial setup to advanced model deployment: 1. **Environment Setup:** Configure a Python environment with essential libraries including FiftyOne, PyTorch, and other dependencies needed for computer vision tasks. 2. **Dataset Exploration:** Load and analyze the CarDD dataset, which contains 4,000 high-resolution images with over 9,000 expertly annotated instances of vehicle damage. Explore dataset statistics, annotation formats (COCO), and understand the dataset's split into training, validation and test sets. 3. **Visual Embeddings Analysis:** Implement state-of-the-art vision models (CLIP, SigLIP) to gain deeper understanding of the dataset. Visualize relationships between damage types using dimensionality reduction techniques and explore similarities through natural language search capabilities. 4. **Model Evaluation:** Learn techniques for evaluating model performance on instance segmentation and detection tasks. Export data for training and evaluate models using FiftyOne's built-in evaluation tools. 5. **Extended Functionality:** Enhance your workflow using the FiftyOne plugin ecosystem for specialized visualizations, search capabilities, and integration with other tools. Deploy pre-trained "zoo" models for real-world car damage detection applications. This workshop is designed for computer vision practitioners, automotive industry professionals, and data scientists interested in applying modern deep learning techniques to vehicle damage detection. Participants will gain practical experience working with a production-grade dataset while learning best practices for model development, evaluation, and deployment. Host ![](https://cdn.sanity.io/images/h6toihm1/production/d65929dbf208894122fd87995a590c7491dbb937-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Harpreet Sahota Voxel51 Bio [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-104-lllmstxt|> ## Valencia AI Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/5d8257912f602205a536d40e7c3890b68fe8c187-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/5d8257912f602205a536d40e7c3890b68fe8c187-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event In-person EMEA Meetups Valencia AI, ML and Computer Vision Meetup - September 25, 2025 Sep 25, 2025 5:00 - 8:30 PM Universidad de Valencia Salon de Grados Avinguda de l'Universitat, s/n 46100 Burjassot, Valencia Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/e6444be43c86d447bb0700e01d15ce03a90c12e9-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Luis San Martin Exceltic Bio ![](https://cdn.sanity.io/images/h6toihm1/production/b3973bbcce3bcb66a5ed3ad6f343a6dd33bf4e8a-481x481.png?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Voxel51 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/22309094488de28707201bd68c28c1de44e7370a-481x481.png?auto=format&dpr=2&fit=max&q=75&w=42) Sandra Lancheros Bio ![](https://cdn.sanity.io/images/h6toihm1/production/ce842c59548a2d7969baa1e093e436b21e3d7cba-481x481.png?auto=format&dpr=2&fit=max&q=75&w=42) Jose Blasco IVIA Bio ![](https://cdn.sanity.io/images/h6toihm1/production/d14925380c377fe4c0a22bb397f108c8db3bf37d-513x512.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Emilio Soria-Olivas Universidad de Valencia Bio About this event Acompáñanos para escuchar charlas de expertos en IA, ML y Visión por Computadora. Abrimos puertas a las 4:30 PM , comenzamos a las 5 PM hasta las 8:30 PM, picoteo y networking. **Evento patrocinado por:** Voxel51, [Generalitat Valenciana](https://www.gva.es/en/inicio/presentacion), [ACI ARA](https://aciara.gva.es/en/), [Institut Valencia d' Investigacions Agraries (IVIA)](https://ivia.gva.es/va/), [Universitat de Valencia](https://www.uv.es/uvweb/college/en/university-valencia-1285845048380.html), [CitCom](https://citcomtef.eu/) Schedule IA Generativa Responsable y Gobernada: MLOps/GenAIOps para Entornos On-Premise Éticos y Seguros ![](https://cdn.sanity.io/images/h6toihm1/production/e6444be43c86d447bb0700e01d15ce03a90c12e9-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Luis San Martin Exceltic Bio La rápida adopción de modelos generativos exige entornos controlados que garanticen trazabilidad, privacidad y gobernanza de extremo a extremo. En esta charla exploraremos cómo aplicar el paradigma MLOps/GenAIOps para desplegar soluciones de IA generativa ética en infraestructuras on-premise, cumpliendo regulaciones como el AI Act y ENS. Revisaremos patrones de arquitectura, herramientas open-source, y estrategias para garantizar observabilidad, reproducibilidad y control de datos en cada etapa del ciclo de vida del modelo. Además, compartiremos casos reales de colaboración internacional en el co-desarrollo de soluciones IA responsables con startups, sector público y comunidades técnicas. IA que Huele a Café: Explorando Datos Agrícolas con FiftyOne ![](https://cdn.sanity.io/images/h6toihm1/production/b3973bbcce3bcb66a5ed3ad6f343a6dd33bf4e8a-481x481.png?auto=format&dpr=2&fit=max&q=75&w=96) Paula Ramos Voxel51 Bio La inteligencia artificial en la agricultura solo es tan buena como los datos que la respaldan, pero los conjuntos de datos desordenados, las anotaciones deficientes y los sesgos ocultos frenan el progreso. Acompaña a Paula en una sesión dinámica sobre segmentación semántica, donde mostrará cómo FiftyOne puede transformar la curación de datos, el análisis de anotaciones y la evaluación de modelos en proyectos de IA agrícola. Usando conjuntos de datos reales de café provenientes de Colombia, exploraremos la segmentación de frutos de café en diferentes etapas de maduración, aprovechando las potentes herramientas de FiftyOne: desde la detección de datos únicos hasta la búsqueda por similitud y la visualización de embeddings. Ya sea que trabajes en robótica agrícola, teledetección o fenotipado de plantas, esta charla te brindará técnicas prácticas para refinar tus conjuntos de datos y potenciar tus flujos de trabajo de IA. IA Agéntica y RAG Visual: Sistemas Inteligentes que Ven, Recuperan y Actúan ![](https://cdn.sanity.io/images/h6toihm1/production/22309094488de28707201bd68c28c1de44e7370a-481x481.png?auto=format&dpr=2&fit=max&q=75&w=96) Sandra Lancheros Bio El auge de los LLMs ha transformado la interacción con el lenguaje, pero sus capacidades son limitadas sin percepción visual. Por otro lado, los sistemas de visión computacional carecen de razonamiento contextual. Esta charla explora cómo combinar visión por computador, recuperación aumentada (RAG) y agentes autónomos para construir sistemas que no solo ven, sino que entienden y actúan. Presentaremos una arquitectura práctica basada en herramientas open source como CLIP, LangChain y FiftyOne, que permite crear agentes visuales capaces de etiquetar imágenes, responder preguntas y tomar decisiones informadas. Una charla para quienes quieren ir más allá del análisis visual y construir flujos inteligentes basados en percepción, razonamiento y acción. Del campo al dato: oportunidades, obstáculos y el camino por Recorrer de la IA en agricultura ![](https://cdn.sanity.io/images/h6toihm1/production/ce842c59548a2d7969baa1e093e436b21e3d7cba-481x481.png?auto=format&dpr=2&fit=max&q=75&w=96) Jose Blasco IVIA Bio La inteligencia artificial (IA) se ha consolidado como una herramienta clave para afrontar algunos de los mayores desafíos en la agricultura actual, desde la optimización del uso de insumos hasta el monitoreo preciso de cultivos y la predicción de rendimientos. En este seminario se presentarán varios casos prácticos reales en los que la IA ha demostrado su potencial para mejorar la eficiencia, sostenibilidad y rentabilidad en distintas etapas de la producción agrícola. Sin embargo, a pesar de sus beneficios, la adopción de estas tecnologías en el sector sigue siendo limitada. Factores como la falta de relevo generacional, el envejecimiento de la población agraria, la baja digitalización en el medio rural y la percepción de complejidad o desconfianza hacia las herramientas digitales suponen importantes barreras. Esta charla, se aborda éxitos y obstáculos actuales, destacando la importancia de diseñar soluciones accesibles, acompañadas de formación y apoyo técnico, que se ajusten a la realidad de un sector tradicional en proceso de transformación. Pasado (reciente), Presente (variable) y Futuro (incierto) de la IA ![](https://cdn.sanity.io/images/h6toihm1/production/d14925380c377fe4c0a22bb397f108c8db3bf37d-513x512.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Emilio Soria-Olivas Universidad de Valencia Bio En este charla se abordará el pasado reciente de la IA (desde 2010); el presente que varía diariamente por la gran cantidad de grandes empresas que están en este campo y el futuro que, cualquier innovación disruptiva (como los transformers en el 2017) puede cambiar. ## Sponsors [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-105-lllmstxt|> ## Verified Auto Labeling Workshop [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![Auto Labeling Workshop: Smarter annotation at scale](https://cdn.sanity.io/images/h6toihm1/production/5212603b698f079790dcca848109776fbc871d0a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![Auto Labeling Workshop: Smarter annotation at scale](https://cdn.sanity.io/images/h6toihm1/production/5212603b698f079790dcca848109776fbc871d0a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Webinars & Workshops Verified Auto Labeling: Smarter Annotation at Scale – June 24, 2025 This event has ended, but you can still catch up! Watch the on-demand recordings and register for our [future events.](https://voxel51.com/events) Jun 24, 2025 9:00 AM to 10:00 AM Pacific Time Virtually over Zoom! About this event Want to take your computer vision workflows to the next level? Then this workshop is for you! Join us for 90 minutes as we demonstrate the full power of the enterprise FiftyOne platform so your teams can deploy visual AI applications faster, with visibility and control. Host ![](https://cdn.sanity.io/images/h6toihm1/production/33c8a7b110bd086417d17f82c8a160fe2bc38d00-256x256.png?auto=format&dpr=2&fit=max&q=75&w=96) Nick Lotz Technical Marketing Engineer Bio ## About the Workshop Manual annotation is one of the biggest bottlenecks in computer vision. Join us to learn how Verified Auto Labeling delivers near-human labeling performance at a fraction of the time and cost. In this webinar, you’ll get a first look at Verified Auto Labeling before it releases to the public. You will learn how to: - Set up auto-labeling quickly: load samples, select a model, define classes, and configure confidence thresholds - Generate labels at scale using built-in recommendation engines and customizable workflows - Review results with side-by-side comparisons, confidence filters, and embedding-based tools to spot poor-quality labels - Streamline QA by auto-approving high-confidence labels and tagging edge cases for human review - Evaluate performance with precision, recall, and F1 scores, and compare outcomes between human and auto-labeled datasets - Identify the most valuable data to collect or relabel to improve model performance Check out our blog post on [Verified Auto Labeling](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance) with our ML research that shows how you can save up to 100,000x on annotation costs. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-106-lllmstxt|> ## Neural 3D Vision Approach [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) VGGT is a Pure Neural Approach to 3D Vision Jun 25, 2025 • 4 min read Article content In this article [One Model, Multiple Tasks](https://voxel51.com/blog/vggt-is-a-pure-neural-approach-to-3d-vision#5820c8e10590) [Rethinking Multi-Task Learning](https://voxel51.com/blog/vggt-is-a-pure-neural-approach-to-3d-vision#b73c7a622ff8) In this article [One Model, Multiple Tasks](https://voxel51.com/blog/vggt-is-a-pure-neural-approach-to-3d-vision#5820c8e10590) [Rethinking Multi-Task Learning](https://voxel51.com/blog/vggt-is-a-pure-neural-approach-to-3d-vision#b73c7a622ff8) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Does This Mark the End of Geometric Post-Processing? VGGT is not like traditional 3D vision pipelines that rely on geometric optimization. I came across this paper at CVPR, where it won the Best Paper Award at the conference. The newly introduced Visual Geometry Grounded Transformer processes multiple images of a scene and directly outputs camera parameters, depth maps, point maps, and 3D tracks in a single forward pass. This feed-forward approach eliminates the need for expensive post-processing steps, such as Bundle Adjustment, which have been considered essential for decades. Most remarkably, this purely neural approach achieves superior results to optimization-based methods while completing reconstructions in under a second, compared to 7–20+ seconds for previous approaches. This performance leap suggests we’ve reached an inflection point where data-driven neural approaches can finally outperform traditional geometric methods. ### Simplicity in Architecture Beats Specialized Design ![](https://cdn.sanity.io/images/h6toihm1/production/8c32da08d29f8c60d3da3197a700900cb3dd9ec5-1400x629.webp?auto=format&dpr=2&fit=max&q=75&w=1400) VGGT’s architecture is straightforward, avoiding complex 3D-specific components. The model employs a standard transformer backbone with a novel alternating-attention mechanism that switches between frame-wise and global self-attention layers. This design allows the network to balance local image understanding with cross-image geometric reasoning without specialized 3D inductive biases. Their ablation studies demonstrate that this approach significantly outperforms both global-only attention and cross-attention alternatives, highlighting that architectural simplicity combined with sufficient training data can be more effective than hand-engineered geometric constraints. The lack of geometry-specific components marks a philosophical shift toward letting the data define the solution. ## One Model, Multiple Tasks VGGT handles an impressive range of input scenarios that typically require specialized solutions. The model processes anywhere from a single image to hundreds of views in a single forward pass, eliminating the need for separate models or processing pipelines for different view counts. This stands in stark contrast to previous state-of-the-art approaches, such as [DUSt3R](https://github.com/naver/dust3r) and [MASt3R](https://github.com/naver/mast3r), which can only process image pairs and require complex post-processing to handle additional views. VGGT also demonstrates strong generalization to challenging scenarios, such as paintings, non-overlapping frames, and scenes with repeating textures. This versatility significantly simplifies the 3D reconstruction workflow for practitioners. ### Using VGGT with FiftyOne [VGGT is now available as a FiftyOne Zoo Model](https://github.com/harpreetsahota204/vggt), making it accessible for immediate use in computer vision workflows. The implementation provides a seamless way to generate depth maps, camera parameters, and 3D point clouds from images with just a few lines of code. FiftyOne’s visualization capabilities allow practitioners to immediately inspect results in an interactive 3D environment, exploring reconstructed scenes from different viewpoints. The model can be configured with different preprocessing modes and confidence thresholds to handle various image types and quality requirements. First, download an example dataset and register the model source: ```python 1#pip install fiftyone 2# 3import fiftyone as fo 4from fiftyone.utils.huggingface import load_from_hub 5import fiftyone.zoo as foz 6 7dataset = load_from_hub("Voxel51/Total-Text-Dataset") 8 9foz.register_zoo_model_source( 10 "https://github.com/harpreetsahota204/vggt", 11 overwrite=True 12) ``` Next, load the VGGT model with your preferred configuration and apply it to your dataset: ```python 1model = foz.load_zoo_model( 2 "facebook/VGGT-1B", 3 install_requirements=True, 4 mode="crop", # you can also pass "pad", 5 confidence_threshold=0.7 6 ) 7 8# Apply to your dataset 9dataset.apply_model(model, "depth_map_path") ``` Finally, create a grouped dataset to visualize RGB, depth, and 3D together: ```python 1import os 2from pathlib import Path 3import fiftyone as fo 4 5grouped_dataset = fo.Dataset("vggt_results", overwrite=True) 6grouped_dataset.add_group_field("group", default="rgb") 7samples = [] 8for filepath in dataset.values("filepath"): 9 path = Path(filepath) 10 base_dir = path.parent 11 base_name = path.stem 12 13 # Create paths for each modality 14 rgb_path = filepath 15 depth_path = os.path.join(base_dir, f"{base_name}_depth.png") 16 threed_path = os.path.join(base_dir, f"{base_name}.fo3d") 17 18 group = fo.Group() 19 samples.extend([\ 20 fo.Sample(filepath=rgb_path, group=group.element("rgb")),\ 21 fo.Sample(filepath=depth_path, group=group.element("depth")),\ 22 fo.Sample(filepath=threed_path, group=group.element("threed"))\ 23 ]) 24 25grouped_dataset.add_samples(samples) 26fo.launch_app(grouped_dataset) # View results interactively ``` This practical implementation demonstrates how quickly cutting-edge research can be integrated into production workflows. ## Rethinking Multi-Task Learning VGGT’s multi-task approach reveals a counterintuitive advantage in predicting “redundant” outputs. The model simultaneously predicts camera parameters, depth maps, and point maps, despite these quantities being mathematically related (depth + cameras can produce point maps). Their ablation studies show that this joint prediction actually improves overall accuracy compared to predicting each quantity individually. This suggests that multi-task learning creates beneficial inductive biases that help the model learn more robust representations, even when the tasks have theoretical overlap. These findings challenge conventional wisdom about task separation in neural network design. ### Not Yet the Complete Solution I tested VGGT on a dataset of Marvel trading cards VGGT still faces challenges with certain specialized imaging scenarios. The current implementation [doesn’t support fisheye](https://huggingface.co/datasets/Voxel51/fisheye8k) or panoramic images and shows reduced performance with extreme input rotations. It also struggles with substantial non-rigid deformations, limiting its application in dynamic scene reconstruction. The authors note that addressing these limitations should be straightforward through fine-tuning on targeted datasets, but these remain areas for improvement. No single approach has yet solved all 3D reconstruction challenges. ### A New Foundation for 3D Vision VGGT represents a potential cornerstone for future 3D vision research and applications. The authors demonstrate that VGGT’s pre-trained features significantly enhance downstream tasks, such as point tracking in dynamic videos and novel view synthesis. This positions VGGT as not just a reconstruction tool but a foundation model for 3D understanding, similar to how CLIP and DINO serve as backbones for 2D vision tasks. With the model now accessible through tools like FiftyOne, the barrier to entry for working with state-of-the-art 3D reconstruction has been substantially lowered. We may be witnessing the beginning of a pure neural approach to 3D vision that finally renders traditional geometric methods obsolete. [3D](https://voxel51.com/blog/tag/3d) [generative AI](https://voxel51.com/blog/tag/generative-ai) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [deep learning](https://voxel51.com/blog/tag/deep-learning) ![](https://cdn.sanity.io/images/h6toihm1/production/a41a0477c7a98264f600772e9568607d070eea59-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Sahota Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/272733186f35c5c6aca68b421960cd41c570e24b-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Composed Image Retrieval at CVPR 2025\\ \\ Event Recaps\\ \\ • \\ \\ Jun 2, 2025](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/7ec61c80f387b16f24b4b2fe33f804264a38b486-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Rethinking How We Evaluate Multimodal AI\\ \\ Event Recaps\\ \\ • \\ \\ Jun 12, 2025](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/e647f4490ed6ad75d63c9f28b67105aebc819e37-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Motion Prompting: Generalized Motion Control for Video Generation\\ \\ Event Recaps\\ \\ • \\ \\ Jun 26, 2025](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-107-lllmstxt|> ## Allstate Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/112f66cd1604b5cee5d9cb2e4039a25260112f0b-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Allstate FiftyOne plays a vital role in processing claims at Allstate India Apr 26, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Allstate India](https://www.allstateindia.com/) provides software development, testing, business process management, technology support, analytics and other IT-enabled services to Allstate and its subsidiaries. > "At Allstate, my team works on auto vehicle damage inspection. Verifying the damage to a vehicle can take an insurance claim agent hours to verify, but using computer vision and FiftyOne, we can segment the parts of vehicles first, then detect the damages, and finally match the damage to repair costs and generate reports for the adjusters." – Pavan Nanjundappa, Data Science Manager at Allstate India ![](https://cdn.sanity.io/images/h6toihm1/production/9803eabd2741b77183a593c53e06c919833e683e-1024x683.jpg?auto=format&dpr=2&fit=max&q=75&w=1024) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-108-lllmstxt|> ## Video Insights and Updates [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Video [![](https://cdn.sanity.io/images/h6toihm1/production/e647f4490ed6ad75d63c9f28b67105aebc819e37-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Motion Prompting: Generalized Motion Control for Video Generation\\ \\ Event Recaps\\ \\ • \\ \\ Jun 26, 2025](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation) [![](https://cdn.sanity.io/images/h6toihm1/production/663dd6a3e6f3e57a03425f932440b5d242133451-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Search and curate video data with FiftyOne, Twelve Labs, and Databricks Vector Search\\ \\ Product & News, Integrations\\ \\ • \\ \\ Jun 5, 2025](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs) ## Enough data wrangling.
 Request a demo. [Get started](https://voxel51.com/link-catcher) [Explore the Demo](https://voxel51.com/link-catcher) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-109-lllmstxt|> ## AI Predictive Maintenance Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) AI for Predictive Maintenance Using Computer Vision Apr 17, 2025 • 9 min read Article content In this article [Adoption of Predictive Maintenance](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#8887a6aa03af) [Limitations of Traditional Methods](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#f8cfcc40c9ae) [Computer Vision's Transformative Potential](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#2a1723d28294) [Advanced Computer Vision Techniques for Predictive Maintenance](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#f07cd547f95f) [Addressing the Challenges of Implementing AI for Predictive Maintenance](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#52d8f697a1f4) [The Role of FiftyOne in Accelerating AI for Predictive Maintenance](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#ab29d327cb7d) [Make Downtime a Thing of the Past](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#543e96eb816b) In this article [Adoption of Predictive Maintenance](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#8887a6aa03af) [Limitations of Traditional Methods](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#f8cfcc40c9ae) [Computer Vision's Transformative Potential](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#2a1723d28294) [Advanced Computer Vision Techniques for Predictive Maintenance](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#f07cd547f95f) [Addressing the Challenges of Implementing AI for Predictive Maintenance](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#52d8f697a1f4) [The Role of FiftyOne in Accelerating AI for Predictive Maintenance](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#ab29d327cb7d) [Make Downtime a Thing of the Past](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision#543e96eb816b) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/41a2af0d84d341158132e10054699704adc756c5-1536x1024.png?auto=format&dpr=2&fit=max&q=75&w=1536) [Predictive maintenance](https://en.wikipedia.org/wiki/Predictive_maintenance) is quickly gaining traction across industries like manufacturing, energy, and transportation due to its ability to optimize maintenance schedules and prevent unexpected equipment failures. Unlike traditional methods that depend on fixed intervals or reactive responses, AI-powered predictive maintenance uses machine learning and computer vision to identify issues before they disrupt operations. This article explores how advanced computer vision techniques and strong data practices in combination with [FiftyOne](https://voxel51.com/fiftyone/) are reshaping maintenance strategies. ## **Adoption of Predictive Maintenance** ![](https://cdn.sanity.io/images/h6toihm1/production/7f66d6c5db1e32a1d4566d472d0d7d9be50389c6-960x1289.jpg?auto=format&dpr=2&fit=max&q=75&w=960) Organizations are increasingly adopting predictive maintenance to streamline schedules and reduce overall maintenance costs. By analyzing real-time operational data and visual indicators from machinery, teams can proactively detect early signs of failure. In heavy industries like oil & gas or automotive, unplanned equipment failures can cause significant downtime, revenue loss, and inflate repair expenses. As a result, AI-driven predictive maintenance continues to grow, promising lower operational expenditures, higher customer satisfaction, and improved operational efficiency. ## **Limitations of Traditional Methods** Historically, industries relied on traditional preventative maintenance strategies like routine inspections, vibration analysis, and scheduled maintenance intervals. While these methods prevented some major breakdowns, they often missed real-time anomalies and subtle mechanical changes preceding failures. Additionally, traditional approaches could be resource-intensive or require specialized equipment. As systems grow increasingly complex and operate at greater scales, relying solely on scheduled or reactive maintenance leaves organizations vulnerable to overlooked failure signals and delayed interventions. ## **Computer Vision's Transformative Potential** \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Computer vision is becoming a powerful tool in AI-driven predictive maintenance. By applying AI algorithms to images or video, organizations can detect wear, misalignment, or overheating in real time. High-resolution cameras, thermal imaging, and advanced analytics reveal subtle issues like surface cracks or abnormal movements that traditional sensors might miss. This post explores advanced computer vision techniques and how they integrate with other data sources, such as vibration or temperature readings, to build more accurate predictive models. ## **Advanced Computer Vision Techniques for Predictive Maintenance** ![](https://cdn.sanity.io/images/h6toihm1/production/c6b4a3c329e0b739dc67dcd13abd5deedcaa0495-2560x1920.jpg?auto=format&dpr=2&fit=max&q=75&w=1600) ### **Beyond Object Detection** Traditional object detection, which uses bounding boxes to identify components, offers limited insight. It can locate parts like valves or bearings but often misses subtle changes or gradual shifts. Predictive maintenance instead requires methods that can detect abnormal patterns, minor surface defects, and early behaviors signaling potential failures. ### **Anomaly Detection** - **One-Class Classification**: In many industrial scenarios, there's ample data on "normal" equipment conditions but limited examples of failures. One-class classifiers learn the expected characteristics of normal machine operations and flag anything that deviates significantly. - **Autoencoders**: A popular machine learning approach for unsupervised anomaly detection. By reconstructing the input image (or sequence of images), the autoencoder indicates potential faults through high reconstruction errors. For instance, a slight fracture in a gear might not reconstruct cleanly, pointing to a possible issue. ### **Semantic Segmentation** Object detection is useful, but semantic segmentation offers pixel-level precision, beneficial for monitoring detailed regions such as cracks in conveyor belts or rust on gears. By categorizing images into areas like "healthy steel surface," "corroded region," or "lubricant residue," maintenance teams can accurately measure degradation severity and scope. Early detection allows teams to optimize maintenance schedules or order replacements before minor defects become major failures. ### **Action Recognition** Some types of predictive maintenance hinge on dynamic behavior rather than static imagery. For instance, a robotic arm that moves erratically or a conveyor belt that stutters intermittently can signal underlying mechanical issues. Action recognition models analyze motion patterns across video frames, detecting deviations from normal operating states (like "smooth rotation"). Sudden jerks or abnormal frequency changes in movement can trigger alerts for further inspection. ### **Leveraging Temporal Information** Predicting failures effectively often requires looking beyond single snapshots: - **Importance of Video Sequences**: Static images don't capture evolving changes, such as heat buildup or subtle vibrations that grow over time. Video data provides context: not just what is happening, but how it changes over seconds or minutes. - **Temporal Convolutional Networks (TCNs)**: TCNs and related architectures excel at modeling time dependencies, offering insights into cyclical equipment behaviors. For instance, if a pump displays increasing vibration amplitude in each cycle, TCN-based anomaly detection can warn of an impending breakdown. ### **Integrating Sensor Data** Combining computer vision with additional sensor data, such as temperature, vibration, or pressure readings, significantly enhances predictive models by creating multi-modal datasets that reveal hidden discrepancies. For instance, a furnace showing normal temperature readings but visually indicating soot accumulation could harbor undiscovered faults. Multi-modal deep learning models that integrate visual and sensor inputs capture a broader range of failure indicators, helping detect subtle faults, like unusual motor currents in a slow-moving conveyor, more effectively than vision-only methods. However, effectively deploying these multi-modal predictive models involves overcoming practical challenges. Organizations must navigate issues like data quality, interpretability, and scalability to ensure successful real-world implementation. ## **Addressing the Challenges of Implementing AI for Predictive Maintenance** \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ### **Data Quality and Quantity** Despite the allure of AI in predictive maintenance, one practical hurdle is curating large, [well-labeled image/video datasets](https://voxel51.com/blog/data-quality-the-hidden-driver-of-ai-success/) covering both normal and fault states. **Challenge** **Description** **Mitigation Strategies** _Data Curation_ Need for large, high-quality, well-labeled image/video datasets is a core requirement, representing both normal and fault states. Establish standardized data collection, thorough annotation processes, and rigorous validation procedures to ensure data accuracy and consistency. _Impact of Poor Data_ Inadequate labeling or biased data sampling significantly compromises the effectiveness of AI models. Rigorous labeling processes (e.g., crowd-sourcing), careful data sampling strategies. _Data Collection Issues_ Gathering consistent imagery is difficult due to varying equipment, diverse operating conditions, and proprietary data constraints across different sites. Standardized data collection protocols where possible, targeted data capture using cameras/drones. _Scarcity of Fault Data_ Examples of equipment failures, especially rare ones, are often hard to obtain in sufficient quantities. Generating synthetic images of faults, using other sensor data to help identify or approximate fault conditions, actively documenting observed failures with cameras/drones. ### **Model Interpretability and Explainability** Maintenance teams and plant operators often hesitate to trust opaque, black-box AI solutions, demanding explainability to confidently adopt predictive maintenance. Without clarity on why a system predicts failures, operators may revert to less accurate traditional methods. Techniques such as Grad-CAM visualization, which highlights critical image regions indicating faults, and SHAP, which quantifies individual feature contributions, provide the necessary transparency to build trust in AI-driven processes. ### **Deployment and Scalability** Even when a predictive model performs well in the lab, deployment can involve real-world integration challenges. - **Real-World Challenges**: Legacy infrastructure in factories and oil rigs complicates software integration; data pipelines must manage noise, missing frames, or sensor failures. - **Edge Computing vs. Cloud Solutions**: Edge computing offers real-time analysis but requires device-level deployment; cloud solutions centralize data but involve latency, bandwidth, and maintenance trade-offs. ## **The Role of FiftyOne in Accelerating AI for Predictive Maintenance** FiftyOne provides powerful tools for data exploration, augmentation, and model evaluation in computer vision, making it ideal for predictive maintenance applications. Here's how: ### **Data Exploration and Visualization** FiftyOne's user-friendly interface simplifies: - **Dataset Inspection**: Large volumes of industrial images and videos can be quickly sorted, filtered, and previewed. This helps teams spot anomalies or identify camera alignment issues. - **Interactive Analysis**: Maintenance engineers can drill down into clusters of images representing different faults, rapidly verifying whether they're correctly labeled or if new labels are needed. ### **Data Curation and Augmentation** Predictive models require robust datasets, involving efficient labeling of thousands of images, such as consistently annotating cracked turbine blades, and applying augmentation techniques like rotation, flipping, or noise addition. These processes help models handle variations in lighting, angle, or occlusion, enhancing their resilience and accuracy. ### **Model Development and Evaluation** The platform also streamlines the model development cycle: - **Data Loading and Preprocessing**: Efficiently import large volumes of factory-floor imagery, organize them by timestamp or device ID, and create training splits. - **Model Predictions Visualization**: Visual overlays of predicted anomalies, like bounding boxes, enable quick iteration. Maintenance staff can rapidly identify misclassifications and provide immediate feedback, streamlining data annotation and model improvement. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## **Make Downtime a Thing of the Past** Organizations looking to improve predictive maintenance should explore advanced computer vision techniques, whether developing solutions in-house or leveraging specialized predictive maintenance services. Tools like FiftyOne support rapid iteration, streamlined dataset refinement, and improved analysis of model outputs, enabling more effective real-time maintenance regardless of the implementation path. For data scientists, engineers, and maintenance managers, adopting AI-driven predictive maintenance is a strategic investment in reliability and efficiency, helping machinery stay ahead of unexpected failures. **Image Citations** - Snelson, Brian. _Final assembly 3_. Photograph. September 14, 2008. Wikimedia Commons. CC BY 2.0. [https://commons.wikimedia.org/wiki/File:Final\_assembly\_3.jpg](https://commons.wikimedia.org/wiki/File:Final_assembly_3.jpg) - Elliot, Raul (Sgt., U.S. Army). _Factory truck parks at the loading dock, Iraq in 2003_. Photograph. September 13, 2003. Wikimedia Commons. Public domain. [https://commons.wikimedia.org/wiki/File:Factory\_truck\_parks\_at\_the\_loading\_dock,\_Iraq\_in\_2003.jpeg](https://commons.wikimedia.org/wiki/File:Factory_truck_parks_at_the_loading_dock,_Iraq_in_2003.jpeg) - 분당선M. 20131123 _TTC repair bus_. Photograph. November 23, 2013. Wikimedia Commons. CC BY-SA 3.0. [https://commons.wikimedia.org/wiki/File:20131123\_TTC\_repair\_bus.jpg](https://commons.wikimedia.org/wiki/File:20131123_TTC_repair_bus.jpg) - Gobierno de Castilla-La Mancha. _2021-12-10 - Visita a las instalaciones de la bodega 'Jesús del Perdón'_. Photograph. December 10, 2021. Wikimedia Commons. CC BY-SA 2.0. [https://commons.wikimedia.org/wiki/File:2021-12-10\_-Visita\_a\_las\_instalaciones\_de\_la\_bodega%E2%80%98Jes%C3%BAs\_del\_Perd%C3%B3n%E2%80%99\_-\_51737235201.jpg](https://commons.wikimedia.org/wiki/File:2021-12-10_-Visita_a_las_instalaciones_de_la_bodega%E2%80%98Jes%C3%BAs_del_Perd%C3%B3n%E2%80%99_-_51737235201.jpg) - Newman, Rob. _Train under repair, Tyseley Depot_. Photograph. June 28, 2008. Wikimedia Commons. CC BY-SA 2.0. [https://commons.wikimedia.org/wiki/File:Train\_under\_repair,Tyseley\_Depot-geograph.org.uk-\_2347616.jpg](https://commons.wikimedia.org/wiki/File:Train_under_repair,Tyseley_Depot-geograph.org.uk-_2347616.jpg) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [predictive maintenance](https://voxel51.com/blog/tag/predictive-maintenance) [AI](https://voxel51.com/blog/tag/ai) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/d2e24d0a14de508f8ccd36eaffd3b909c9f193b9-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Image Embeddings Transform Computer Vision Capabilities\\ \\ Learn\\ \\ • \\ \\ Nov 25, 2024](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities) [![](https://cdn.sanity.io/images/h6toihm1/production/9252e8bb5db5c4805f4a6f315b51527ee5c22072-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Guide to AI Image Segmentation\\ \\ Learn\\ \\ • \\ \\ Dec 19, 2024](https://voxel51.com/blog/a-guide-to-ai-image-segmentation) [![](https://cdn.sanity.io/images/h6toihm1/production/b4fd054bf7574e4060ffc7a9a8f201c31a66c327-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Using Computer Vision to Enhance Customer Experience in Retail\\ \\ Learn\\ \\ • \\ \\ Jan 23, 2025](https://voxel51.com/blog/using-computer-vision-to-enhance-customer-experience-in-retail) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-110-lllmstxt|> ## Visual AI Event 2025 [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/f738f742abf67718f860efe462f9042b6639ec99-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/f738f742abf67718f860efe462f9042b6639ec99-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event Virtual Americas Meetups Manufacturing Visual AI in Manufacturing and Robotics - September 11, 2025 Sep 11, 2025 9 AM Pacific Online. Register for the Zoom! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/a4fae3a08d3954497752168dee8ff10db1dd4c8e-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Clinton J Smith, PhD RIOS Intelligent Machines Bio ![](https://cdn.sanity.io/images/h6toihm1/production/ad4c3057e20dcf55d853bb65e25e4b91e2ffe152-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Steve Xie, PhD Lightwheel Bio ![](https://cdn.sanity.io/images/h6toihm1/production/1973cc398f537cccdd1e9ee41c497661c2fe59c0-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Samet Akcay Intel Bio ![](https://cdn.sanity.io/images/h6toihm1/production/a4b60eb5164caab096007f2c07502eebab7e33f4-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Allen Lee Voxel51 Bio About this event Join us for dat two in a series of virtual events to hear talks from experts on the latest developments at the intersection of Visual AI, Manufacturing and Robotics. Schedule Bringing Specialist Agents to the Physical World to Improve Manufacturing Output ![](https://cdn.sanity.io/images/h6toihm1/production/a4fae3a08d3954497752168dee8ff10db1dd4c8e-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Clinton J Smith, PhD RIOS Intelligent Machines Bio U.S. manufacturing productivity (output per labor hour) has been stagnant since 2008, driven by a stall in technology integration as well as available workers. RIOS Agents are collaborative AI perception and control systems that act as plant managers' eyes on the ground. Our Agents become specialists in a process, observing process steps, reporting on them, and ultimately controlling them by integrating into new or existing equipment. This enables factory production to be optimized in a way that was previously not possible. Accelerating Robotics with Simulation ![](https://cdn.sanity.io/images/h6toihm1/production/ad4c3057e20dcf55d853bb65e25e4b91e2ffe152-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Steve Xie, PhD Lightwheel Bio In this session, Steve Xie, CEO of Lightwheel, shares how simulation-first workflows and high-quality SimReady assets are transforming the development of visual AI in manufacturing. From warehouse anomaly detection to worker safety and object identification, Steve will explore how physics-accurate simulation and synthetic datasets can drive scalable AI training with minimal real-world data. Drawing from Lightwheel’s deployment of robot models like GR00T N1 in factory environments, the talk highlights how unifying vision, language, and action in simulation accelerates real-world deployment while improving safety, generalization, and efficiency. Anomalib 2.0: Edge Inference and Model Deployment ![](https://cdn.sanity.io/images/h6toihm1/production/1973cc398f537cccdd1e9ee41c497661c2fe59c0-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Samet Akcay Intel Bio When deploying models for inference, just exporting the models and calling them via the inferencers do not work. There are challenges related to pre-processing and post-processing. Any deviation in these steps during inference impacts performance. This talk is about how we re-designed components of Anomalib to integrate pre and post-processing steps in the model graph. Exploring Robotic Manipulation Datasets using FiftyOne: DROID and Amazon Armbench ![](https://cdn.sanity.io/images/h6toihm1/production/a4b60eb5164caab096007f2c07502eebab7e33f4-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Allen Lee Voxel51 Bio [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-111-lllmstxt|> ## Powering Physical AI [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) How Voxel51 is Powering Physical AI with Databricks Aug 20, 2025 • 5 min read Article content In this article [Why Physical AI starts with better data](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks#857f71ed6a8d) [Powering the AV data pipeline](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks#eec79860cd40) [The Physical AI stack: Databricks + FiftyOne](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks#ca7dc96a0cb9) [Use case: Surfacing rare events in AV datasets](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks#13ee9e9835a8) In this article [Why Physical AI starts with better data](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks#857f71ed6a8d) [Powering the AV data pipeline](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks#eec79860cd40) [The Physical AI stack: Databricks + FiftyOne](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks#ca7dc96a0cb9) [Use case: Surfacing rare events in AV datasets](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks#13ee9e9835a8) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## **Why Physical AI starts with better data** In autonomous vehicle (AV) and advanced driver-assistance system (ADAS) development, the hardest problems often hide in the long-tail rare events, such as a pedestrian crossing in the rain at night or a cyclist partially obscured in a crosswalk. These edge cases are exactly where models struggle most, and yet they’re buried across petabytes of unstructured sensor data. Too often, teams spend months writing custom queries, trawling through metadata, or hand-labeling samples just to uncover a handful of the moments that matter. Even then, they’re rarely sure they’ve found all the right examples. Voxel51 changes that. By combining the scalable data infrastructure of Databricks with the powerful discovery and curation tools offered by [FiftyOne](https://voxel51.com/fiftyone), the data engine for visual and multimodal AI that unlocks the full potential of your model performance, teams can now search, slice, and surface critical AV/ADAS scenarios in hours instead of weeks. This joint stack brings structure to unstructured data and makes it possible to iteratively build the datasets that truly move model performance forward. ## **Powering the AV data pipeline** This Databricks + FiftyOne integration lays the groundwork for something even more powerful: the AV/ADAS data engine. Instead of treating data discovery and annotation as one-off tasks, teams can now build continuous pipelines that identify edge cases, validate quality, and trigger labeling workflows, all with human-in-the-loop oversight. This isn’t just about finding better data once. It’s about creating a feedback loop where model insights guide data curation, the curated data fuels better models, and the cycle repeats. With the scalable compute of Databricks and the visual intelligence of FiftyOne, we're enabling a new kind of infrastructure, one that automates the search for signal in the noise and accelerates the development of safer, smarter vehicles. ## **The Physical AI stack: Databricks + FiftyOne** The Databricks + FiftyOne integration forms a seamless pipeline that connects a scalable data infrastructure with visual-first discovery tools - Databricks provides the scalable backbone to: - Store massive AV datasets in [Unity Catalog-managed Volumes](https://www.databricks.com/product/unity-catalog) - Capture structured metadata in Delta tables - Index data for rapid retrieval with [Databricks Vector Search](https://www.databricks.com/product/machine-learning/vector-search) - FiftyOne builds on top of that foundation and provides the data engine for visual and multimodal AI to: - [Explore, curate, and analyze](https://voxel51.com/curation) visual datasets and models - Leverage [Data Lens](https://docs.voxel51.com/enterprise/data_lens.html) to visually query and filter events at scale - Discover unstructured sensor data with embeddings, similarity search, and scenario filters - Identify rare conditions and curate high-value subsets for model development Together, they enable AV/ADAS teams to work fluidly across structured and unstructured worlds. Databricks powers the storage, governance, and indexing, while FiftyOne brings those indexed volumes to life in a rich, interactive interface. Let’s take a look at how. ## **Use case: Surfacing rare events in AV datasets** Autonomous vehicle (AV) datasets such as [nuScenes](https://www.nuscenes.org/nuscenes), [BDD100K](https://bair.berkeley.edu/blog/2018/05/30/bdd/), or internal fleet collections are massive, complex, and filled with edge cases that directly impact model performance. The challenge isn’t collecting data; it’s finding the moments that matter most buried within millions of frames. ### Starting with Databricks Volumes The first step is staging AV datasets in Unity Catalog-managed Volumes, which centralizes all your sensor data (camera images, LiDAR sweeps, radar, and labels) in a governed storage layer. We then structure key metadata (e.g., weather, time of day, object counts per frame) into a Delta table. This structured index provides the foundation for efficient queries without combing through raw files. ### **Searching with [Data Lens](https://docs.voxel51.com/enterprise/data_lens.html) \+ [Databricks Vector Search](https://www.databricks.com/product/machine-learning/vector-search)** With the dataset indexed, connect it to [FiftyOne Data Lens](https://voxel51.com/blog/streamline-visual-data-discovery-with-fiftyone-data-lens). Powered by Databricks Vector Search, Data Lens allowed us to visually query for combinations of attributes and embeddings that define key events, such as: - _Adverse weather conditions_ (e.g., rain, fog, snow) - _Pedestrian or cyclist presence in crosswalks_ - _Nighttime scenarios with multiple overlapping objects_ These queries return focused subsets of frames in seconds, even across millions of samples. ### **Drilling down with scenario analysis and embedding view** It's not enough to just find a scenario; you need to know exactly what you are missing. Voxel51’s Model [Evaluation Panel](https://docs.voxel51.com/user_guide/evaluation.html), which includes [analysis of different scenarios](https://docs.voxel51.com/user_guide/evaluation.html#analyzing-scenarios-sub-new), lets you understand where your model is failing and where you need to bolster your dataset. Then you can search both your current dataset and your large data lake using similarity search powered by Databricks to find that data. ### **Results: finding what matters** In just a few minutes, this workflow can isolate AV scenarios crucial to training. The curated datasets become the foundation for improved model training and continuous evaluation. No more wasted hours searching across petabytes of data for what you are looking for, get what you need in minutes. The right data creates the best model. ### **Finding rare events and edge cases** In AV/ADAS, the most important data is often the hardest to find. A single missed pedestrian in a crosswalk or a rare combination of bad weather and traffic conditions can have an outsized impact on model performance and on safety. This is exactly what Voxel51 enables with [FiftyOne on Databricks](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai). FiftyOne and Databricks bridge the gap between managing and making sense of visual data. By bringing together Databricks’ scalable storage and indexing with FiftyOne’s visual-first discovery and curation, teams can now surface these critical events in ways that simply weren’t possible before: - **Pinpoint long-tail scenarios instantly:** Search millions of frames for combinations of conditions, objects, and behaviors that define your edge cases. - **See your data in new ways**: Embeddings and similarity search let you uncover patterns and outliers you might not even know to look for. - **Continuously close the loop:** As new data flows into Databricks, FiftyOne makes it easy to expand and refine curated datasets without starting over. This isn’t just incremental improvement, it’s a step change. For the first time, AV/ADAS teams can see their entire dataset, find what truly matters, and act on it in real time. That’s the promise of Physical AI, and it’s being enabled by Voxel51. Join our upcoming webinar on **Sept 4, 2025, @9am PT** to learn how Porsche is advancing its autonomous vehicle (AV) development by leveraging the power of Voxel51 and Databricks. [Register here](https://events.databricks.com/FY260904-WB-EngineeringRD/registration?scid=701Vp00000U6EaCIAV&utm_medium=Partner&utm_source=n/a) → [integrations](https://voxel51.com/blog/tag/integrations) [dataset curation](https://voxel51.com/blog/tag/dataset-curation) ![](https://cdn.sanity.io/images/h6toihm1/production/3b39056326e925c10b46da1324bc3c5840a1629c-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Dan Gural Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/6aafb2b5fa699824c252fabfe2607eaeb820616a-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Enabling the AV Datasets of the Future with NVIDIA NuRec and FiftyOne\\ \\ Product & News\\ \\ • \\ \\ Aug 11, 2025](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/663dd6a3e6f3e57a03425f932440b5d242133451-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Search and curate video data with FiftyOne, Twelve Labs, and Databricks Vector Search\\ \\ Product & News, Integrations\\ \\ • \\ \\ Jun 5, 2025](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs) [![](https://cdn.sanity.io/images/h6toihm1/production/14713df0d4dec67cd3bb5e9c292c607820df061d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Databricks and Voxel51: Scaling Data-Centric Visual AI on the Data Intelligence Platform\\ \\ Product & News\\ \\ • \\ \\ Jul 22, 2025](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-112-lllmstxt|> ## NVIDIA C-RADIOv3 Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps), [Integrations](https://voxel51.com/blog/category/integrations) NVIDIA’s C-RADIOv3 is the Vision Encoder You Should Be Using Jun 23, 2025 • 8 min read Article content In this article [CVPR 2024 Might Spark Interest in Agglomerative Vision Models](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#80a868507fa2) [The Rise of Agglomerative Vision Models](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#e182632daf70) [What Exactly Is “Agglomerative” Modeling?](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#ca2559cc6157) [What Makes RADIOv2.5 Different](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#954479c50b0c) [The Token Compression Breakthrough](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#c1fc7cb0f37b) [Practical Performance Advantages](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#9eb9a211eb9f) [Real-World Applications](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#1dd24bd38a74) [Getting Started with RADIO in FiftyOne](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#e841c1291a7c) [Advanced FiftyOne Workflows](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#95718fd80086) [Lessons from Implementing RADIO in FiftyOne](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#780eb48bc225) [Embracing the Agglomerative Future](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#1d2a97562117) In this article [CVPR 2024 Might Spark Interest in Agglomerative Vision Models](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#80a868507fa2) [The Rise of Agglomerative Vision Models](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#e182632daf70) [What Exactly Is “Agglomerative” Modeling?](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#ca2559cc6157) [What Makes RADIOv2.5 Different](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#954479c50b0c) [The Token Compression Breakthrough](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#c1fc7cb0f37b) [Practical Performance Advantages](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#9eb9a211eb9f) [Real-World Applications](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#1dd24bd38a74) [Getting Started with RADIO in FiftyOne](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#e841c1291a7c) [Advanced FiftyOne Workflows](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#95718fd80086) [Lessons from Implementing RADIO in FiftyOne](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#780eb48bc225) [Embracing the Agglomerative Future](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using#1d2a97562117) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## CVPR 2024 Might Spark Interest in Agglomerative Vision Models I stumbled upon the [RADIOv2.5 poster at CVPR this year](https://arxiv.org/abs/2412.07679), and it immediately caught my attention amid the sea of incremental improvements. The conference floor was packed with the usual suspects — minor tweaks to transformers, yet another CLIP variant, and countless papers promising fractional gains on ImageNet. But this poster was different. The visualizations, which show consistent feature maps across resolutions, made me stop in my tracks, especially when contrasted with the bizarre mode-switching behaviour of previous approaches. After a brief chat with the authors about their multi-resolution training strategy, I knew this was something I needed to investigate further once I got home. Sometimes the most important advances don’t come with the flashiest presentations, but they fundamentally change how we approach problems. ## The Rise of Agglomerative Vision Models Foundation models have transformed computer vision, but the real power lies in combining their strengths. The vision landscape is cluttered with specialized models: - CLIP excels at connecting images to text - DINO captures semantic structure - SAM delivers precise segmentation masks. Each model is impressive on its own, but each also has significant limitations when applied outside its comfort zone. **Agglomerative models like RADIOv2.5 solve this problem by distilling knowledge from multiple teacher models into a single, versatile student.** > “For the student model to be consistently accurate across resolutions, it is sufficient to match all teachers at all resolutions, and to train at two resolutions simultaneously in the final training stage.” **The era of single-purpose vision models is ending.** ## What Exactly Is “Agglomerative” Modeling? Agglomerative vision models represent a fundamental shift from how we’ve traditionally built computer vision systems. Most practitioners are familiar with two approaches: using a single pre-trained model (like CLIP or ResNet) fine-tuned for a specific task, or creating ensembles that average predictions from multiple models at inference time. Agglomerative models take a radically different approach — they use knowledge distillation to transfer the learned representations from multiple “teacher” models into a single “student” model during training. This isn’t simple ensembling; the student learns to produce features that simultaneously match those of CLIP, DINO, and SAM for the same input image, effectively compressing multiple specialized models into one. It’s knowledge distillation on steroids, creating a single backbone that inherits the superpowers of all its teachers. ## What Makes RADIOv2.5 Different [RADIOv2.5](https://cvpr.thecvf.com/media/PosterPDFs/CVPR%202025/34144.png?t=1748786629.8344061) is the first agglomerative model that truly works across all input resolutions. Previous agglomerative models like AM-RADIO suffered from “mode switching” — they’d produce DINO-like features at low resolutions but SAM-like features at high resolutions. This inconsistency made them unreliable for production use. RADIOv2.5 solves this through multi-resolution training, carefully balancing teacher influence, and applying PHI Standardization to normalize feature distributions across models. Resolution robustness is the killer feature you didn’t know you needed. ## The Token Compression Breakthrough > “Token Merging is very effective at retaining the most diverse information under high compression ratios.” Token compression is the secret to efficiently integrating high-resolution vision with language models. Traditional approaches like pixel unshuffling blindly compress visual tokens without considering information content, treating every region of an image as equally important. RADIOv2.5 introduces a bipartite matching approach that intelligently merges similar tokens, preserving detail in information-rich regions while aggressively compressing homogeneous areas. This selective compression means OCR, document understanding, and fine-grained tasks all perform dramatically better. Your LLM doesn’t need 4,096 tokens to understand a simple image. ## Practical Performance Advantages [The numbers don’t lie](https://arxiv.org/pdf/2412.07679): RADIOv2.5 consistently outperforms specialized models across diverse benchmarks. On semantic segmentation, RADIOv2.5-B scores 48.94 mIoU on ADE20k, surpassing DINOv2-g’s 48.79 despite being one-tenth the size. For vision-language tasks, RADIOv2.5-H achieves 68.9% on TextVQA and 53.9% on DocVQA, handily beating SigLIP’s 67.6% and 57.1% respectively. Most impressively, RADIOv2.5 maintains consistent performance as resolution increases, while competitors degrade. Better features translate directly to better downstream performance. ## Real-World Applications RADIOv2.5 shines in complex real-world scenarios that demand multi-faceted understanding. Document AI applications benefit enormously from RADIOv2.5’s ability to understand both high-level document structure and fine text details simultaneously. Robotics applications leverage its combination of SAM-like spatial awareness and DINO’s semantic understanding for more robust perception. In medical imaging, the model’s resolution flexibility means it can process everything from whole-slide pathology scans to focused ROIs without quality degradation. The next generation of vision applications will be built on agglomerative models. ## Getting Started with RADIO in FiftyOne The fastest way to experiment with RADIO models is through the FiftyOne computer vision platform’s integration. FiftyOne makes it dead simple to leverage RADIO’s powerful embeddings for similarity search, clustering, and data curation. The integration provides access to all model variants (B/L/H/g) with both summary and spatial feature outputs. Setting up is straightforward — just register the model source, load your preferred variant, and start extracting features. The real power comes from combining RADIO’s rich embeddings with FiftyOne’s Brain workflows for duplicate detection, outlier discovery, and representative sample selection. Five minutes of setup saves days of custom implementation work. ## Advanced FiftyOne Workflows RADIO’s dual-output capability opens up powerful workflows that most models can’t support. The spatial features option is particularly valuable for understanding what your model is focusing on — something CLIP simply can’t provide. By loading a spatial model variant with `output_type="spatial"`, you can generate attention heatmaps that reveal which image regions are driving your model's decisions. This is invaluable for debugging, bias detection, and understanding failure cases. Combine this with the global embeddings for a complete understanding of your dataset's structure and individual sample characteristics. The ability to extract both global and spatial features from a single model is a workflow accelerator. `` ## Lessons from Implementing RADIO in FiftyOne I spent an entire day implement NVIDIA’s RADIO model as a FiftyOne zoo model. What started as a “simple” model wrapper turned into a fascinating journey through the complexities of modern vision transformers. Here are my key takeaways: ### Work WITH the Model, Not Against It **The biggest lesson**: Don’t fight the model’s native output format. I spent hours trying to reshape patch tokens back to spatial grids, making assumptions about patch ordering that were wrong. The breakthrough came when I realized RADIO already provides spatial structure in NCHW format — I just needed to use it correctly. - **Wrong approach**: `sqrt(num_patches)` → guess spatial dimensions → reshape and hope - **Right approach**: Trust RADIO's spatial layout → apply PCA across channels → preserve native structure ### The Devil is in the Data Types Storage efficiency matters more than you think. High-resolution spatial features can easily exceed MongoDB’s 16MB document limit. Converting from `float32` to `uint8` gave us a 75% size reduction with negligible quality loss. Sometimes the "obvious" solution (store raw floats) isn't practical at scale. ### Dual Output Strategies Are Powerful Rather than forcing everything into one output type, I implemented separate pathways: - **Summary embeddings** → 1D vectors for similarity/search - **Spatial features** → 2D heatmaps for interpretability This gives users the best of both worlds and avoided awkward compromises. ### Modern Models Need Modern Optimization Mixed precision ( `bfloat16`) and GPU compatibility checking aren't just nice-to-haves anymore. However, graceful fallbacks are crucial - not every user has an RTX 4090. Always provide a path that works on older hardware. ### Iteration Speed Matters I went through probably 15+ different approaches to the spatial heatmap problem. Having fast iteration cycles — clear error messages, good debugging output, easy rollbacks — was essential. The solution that worked was actually quite simple, but finding it required lots of experimentation. ### Domain Knowledge is Irreplaceable Understanding the difference between RADIO’s NCHW and NCL output format was the key insight that unlocked everything. No amount of clever engineering could substitute for understanding what the model actually outputs. Read the papers, understand the architecture, don't just treat models as black boxes. ### The Final Touch Makes All the Difference The solution that finally worked included one small detail that transformed everything: `gaussian_filter(attention_1d, sigma=1.0)`. This tiny addition turned noisy, hard-to-interpret attention maps into smooth, professional-looking visualizations. Never underestimate the importance of that final 10% of polish. > "The best technical solutions often look obvious in retrospect, but the path to get there rarely is. Today was a perfect reminder that building good developer tools is as much about understanding the problem domain as it is about writing code." ## Embracing the Agglomerative Future The vision field stands at an inflection point where single-purpose models are giving way to agglomerative approaches that combine the best of multiple worlds. RADIOv2.5 exemplifies this shift with its multi-resolution training that eliminates mode switching, token compression that preserves critical information, and teacher balancing that creates truly unified representations. The practical benefits are clear across benchmarks: outperforming DINOv2-g on segmentation with one-tenth the parameters, beating SigLIP on document understanding tasks, and maintaining consistency across resolutions where other models falter. Implementing this model in production environments like FiftyOne unlocks these capabilities while teaching us valuable lessons about modern vision architectures — from respecting native tensor formats to understanding the importance of dual output pathways and optimization techniques. What makes these models truly special isn’t just their theoretical elegance but their practical utility across a stunning range of applications: document understanding, medical imaging, robotics, and more. As specialized teachers continue to emerge in the vision landscape, the agglomerative approach will only become more powerful, creating a virtuous cycle where progress in specialized models accelerates the capabilities of generalist backbones. The future belongs to models that don’t just excel at one thing but combine multiple specialized strengths into a cohesive, adaptable whole — and RADIOv2.5 is showing us the way. [NVIDIA](https://voxel51.com/blog/tag/nvidia) [CVPR](https://voxel51.com/blog/tag/cvpr) [ML research](https://voxel51.com/blog/tag/ml-research) ![](https://cdn.sanity.io/images/h6toihm1/production/a41a0477c7a98264f600772e9568607d070eea59-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Sahota Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/7ec61c80f387b16f24b4b2fe33f804264a38b486-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Rethinking How We Evaluate Multimodal AI\\ \\ Event Recaps\\ \\ • \\ \\ Jun 12, 2025](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/1798d34efc8956a6696377fb7776886ec0e61092-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ UnCommon Objects in 3D\\ \\ Datasets, Event Recaps\\ \\ • \\ \\ Jun 18, 2025](https://voxel51.com/blog/uncommon-objects-in-3d) [![](https://cdn.sanity.io/images/h6toihm1/production/3f54c19e72060ff2fa99673840a472f4d6e85c8e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Van der Maaten’s Three-System Roadmap to AGI Is Brilliantly Pragmatic\\ \\ Event Recaps\\ \\ • \\ \\ Jun 17, 2025](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-113-lllmstxt|> ## Boston AI/ML Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/79482d42c785829703e25f224beeb33fce11d9b3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/79482d42c785829703e25f224beeb33fce11d9b3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) In-person Americas Meetups Boston AI, ML and Computer Vision Meetup - June 26, 2025 Jun 26, 2025 5:00 - 8:00 PM Microsoft Research Lab – New England (NERD) at MIT Deborah Sampson Conference Room One Memorial Drive, Cambridge, MA, 02142 Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/53d98fe6e2a1967d473888a0d0666a00b3d0f12a-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Jingnan Shi MIT Bio ![](https://cdn.sanity.io/images/h6toihm1/production/8d52eaf6fba96b9fb1ae94a9355541a49dd449d0-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Suprateem Banerjee InterSystems Bio ![](https://cdn.sanity.io/images/h6toihm1/production/b435c2b911cd69bd4cf4d2ad7699d4c8e5500915-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Pooja Mistry Postman Bio ![](https://cdn.sanity.io/images/h6toihm1/production/789c4d2e94a74f14ce09ab95663668161039216f-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Dan Gural Voxel51 Bio About this event Hear talks from experts on cutting-edge topics in AI, ML, and computer vision on June 26 at Microsoft NERD Schedule Resilient Object Perception for Robotics ![](https://cdn.sanity.io/images/h6toihm1/production/53d98fe6e2a1967d473888a0d0666a00b3d0f12a-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Jingnan Shi MIT Bio A broad array of applications, ranging from search and rescue to self-driving vehicles, require robots to perceive and understand the geometry of objects in the environment. Object perception needs to reliably work in a variety of scenarios and preserve a desired level of performance in the face of outliers and shifts from the training domain. Obtaining such a level of performance requires robust estimation algorithms that are able to identify and reject outliers, as well as techniques to continually improve performance of learning-based perception modules during test-time. In this talk, I discuss my three projects on this topic: (1) solvers and a graph-theoretic framework that together help achieve state-of-the-art pose estimation performance even under high outlier rates, (2) self-supervised object pose estimators that can improve performance during test-time with accuracy comparable to state-of-the-art supervised methods and (3) a test-time adaptation method for both object shape reconstruction and pose estimation without the need for CAD models. Pixie: Building a Local ChatGPT Alternative using Ollama ![](https://cdn.sanity.io/images/h6toihm1/production/8d52eaf6fba96b9fb1ae94a9355541a49dd449d0-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Suprateem Banerjee InterSystems Bio I built Pixie out of a desire to replace my ChatGPT workflows with a local alternative. This project runs parallel to projects like LLMStudio, caters more towards Ollama models, and thus allows us to optimize for the user experience for Ollama workflows. I will go through the different philosophies at play here, design choices and how to create a system that can substitute for the "ChatGPT experience" while remaining local and open source. You Can’t Do AI Without Quality APIs ![](https://cdn.sanity.io/images/h6toihm1/production/b435c2b911cd69bd4cf4d2ad7699d4c8e5500915-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Pooja Mistry Postman Bio The Agentic Era is here — and it runs on APIs. In today’s AI revolution, success isn’t about who has the biggest model, but who builds the highest-quality, AI-ready APIs. From powering intelligent agents to enabling seamless orchestration, APIs are the backbone of modern AI systems. At Postman, we see how collaboration, testing, and documentation are essential to delivering APIs that truly support AI innovation. This talk explores why robust APIs are the foundation of the AI future—because in this new era, you simply can’t do AI without APIs. Using NVIDIA Omniverse + FiftyOne to Build the AI Datasets of the Future ![](https://cdn.sanity.io/images/h6toihm1/production/789c4d2e94a74f14ce09ab95663668161039216f-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Dan Gural Voxel51 Bio Nothing is driving physical AI forward faster than autonomous vehicles. In this talk, I'll walk through how to build cutting-edge AV datasets using synthetic data, NeRFs, smarter curation, and vector search. It's all powered by NVIDIA Omniverse and FiftyOne, and it's a glimpse at what the future of data looks like. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-114-lllmstxt|> ## Porsche AV Development Webinar [![Databricks Logo](https://assets.swoogo.com/uploads/medium/3680870-65f41e39a16c8.png)](https://databricks.com/) Virtual How Porsche Uses Auto-Labeling to Supercharge AV Development ![](https://assets.swoogo.com/uploads/medium/4747099-67781052f3f33.png) - [Registration](https://events.databricks.com/fy260904-wb-engineeringrd/register?utm_medium=Partner&utm_source=n%2Fa) - Meetings - [1Event Registration](https://events.databricks.com/fy260904-wb-engineeringrd/registration?utm_medium=Partner&utm_source=n%2Fa) - 2Confirmation ## Register Now Submit ## **Thursday, September 4, 2025 \| 9:00-10:00 AM PT** Join us for an exclusive webinar showcasing how Porsche is advancing its autonomous vehicle (AV) development by leveraging the power of Visual AI from Voxel51 and the scalability of Databricks. Discover how cutting-edge video-language models (VLMs), automated data pipelines, and human-in-the-loop feedback loops are transforming the way AV systems are trained and refined—dramatically reducing reliance on manual annotation while accelerating model iteration. **Revolutionizing the AV Data Engine** Learn how Porsche is using Voxel51 and Databricks to: - Reduce manual annotation workloads and labeling costs through AI-powered video auto-labeling - Automate ingestion and model retraining loops, enabling faster iteration at lower cost - Use scalable compute and Vector Search at enterprise scale with Databricks’ Data Intelligence Platform - Deliver smarter, safer AV systems by optimizing their data engine pipeline **What You’ll Learn:** - A behind-the-scenes look at Porsche’s data architecture for AV development - How Voxel51’s Visual AI toolkit integrates with Databricks to power scalable, automated data workflows - A live walkthrough of the data engine cycle, from ingest to annotation to retraining - Insights and lessons from Porsche’s Tin Stribor Sohn, PhD, on building data-centric AV systems **Don’t Miss Out** This is your chance to see real-world impact from enterprise AI integration—featuring one of the world’s most innovative car manufacturers. ➡️ Reserve your spot now and stay ahead in the race for fully autonomous, intelligent vehicles. ## **Co-sponsors** [![Voxel51](https://assets.swoogo.com/uploads/medium/5570221-687e77f5f10af.png)](https://events.databricks.com/FY260904-WB-EngineeringRD/sponsor/870198/voxel51 "Sponsor Details") * * * ## **Agenda** | | | | --- | --- | | 9:00 AM - 9:10 AM | Welcome | | 9:10 AM - 9:20 AM | The Data Challenge in AV Development Workflows | | 9:20 AM - 9:40 AM | How Porsche Uses Voxel51 + Databricks for Visual AI | | 9:40 AM - 9:50 AM | Voxel51 + Databricks Solution Demo | | 9:50 AM - 10:00 AM | Event Wrap up and Q&A | ## **Speakers** ![](https://assets.swoogo.com/uploads/medium/5670692-68a6329e45f87.png) **Dan Gural** Voxel51 Machine Learning Evangelist ![](https://assets.swoogo.com/uploads/medium/5586633-6883ce4e8fa59.png) **Tin Stribor Sohn** Porsche Autonomous Driving Data Technical Lead ![](https://assets.swoogo.com/uploads/medium/3946640-664735c85a157.jpeg) **Shiv Trisal** Databricks Global Manufacturing & Energy Industry Leader ![](https://databricks.com/wp-content/uploads/2021/10/databricks_logo_sm.svg) © Databricks 2025. All rights reserved. Apache, Apache Spark, Spark and the Spark logo are trademarks of the [Apache Software Foundation](https://www.apache.org/). [Privacy Notice](https://www.databricks.com/privacynotice) \| [Event Terms](https://www.databricks.com/legal/terms/event-terms-and-conditions) \| [Terms of Use](https://www.databricks.com/terms-of-use) \| [California Privacy](https://www.databricks.com/legal/supplemental-privacy-notice-california-residents) \| Your Privacy Choices![](https://www.databricks.com/sites/default/files/2022-12/gpcicon_small.png) <|firecrawl-page-115-lllmstxt|> ## CVPR 2025 AI Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Best of CVPR 2025: Conversations at the Cutting Edge of AI Jul 3, 2025 • 5 min read Article content In this article [CVPR 2025 Insights #5: Dialogues in AI Research](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#36508bae658d) [OpticalNet: Seeing Beyond the Diffraction Limit](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#953c513f86da) [Nonisotropic Gaussian Diffusion for Human Motion Prediction](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#87a9e8474852) [Interpretable Medical AI with Human-in-the-Loop Models](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#0375c5f72e46) [From Occlusion to Imagination: OFER’s Face Reconstruction Breakthrough](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#85ba9673455a) [Driving with Foundation Models: A Vision for Autonomous Cars](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#35ce7dabb5e5) [RANGE: Encoding the Earth’s Secrets](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#3ca2a8893a03) [Few-Shot Agriculture Vision with Grounding DINO](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#a735e0956e94) [SmartHome-Bench: Multimodal AI for Real-World Anomalies](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#185937fb36fa) [Multi-View Anomaly Detection with Normalizing Flows](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#7859292225bb) [CLIP++: Scaling Down to Go Further](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#5eef51757623) [Maxing Out Medical Benchmarks](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#290ff5558613) [Why This Series Matters](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#297d3b1c3998) [Join Us for the Best of CVPR Series Meetups](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#3c7d42ccaccc) [What is next?](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#67c83e89a725) In this article [CVPR 2025 Insights #5: Dialogues in AI Research](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#36508bae658d) [OpticalNet: Seeing Beyond the Diffraction Limit](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#953c513f86da) [Nonisotropic Gaussian Diffusion for Human Motion Prediction](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#87a9e8474852) [Interpretable Medical AI with Human-in-the-Loop Models](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#0375c5f72e46) [From Occlusion to Imagination: OFER’s Face Reconstruction Breakthrough](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#85ba9673455a) [Driving with Foundation Models: A Vision for Autonomous Cars](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#35ce7dabb5e5) [RANGE: Encoding the Earth’s Secrets](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#3ca2a8893a03) [Few-Shot Agriculture Vision with Grounding DINO](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#a735e0956e94) [SmartHome-Bench: Multimodal AI for Real-World Anomalies](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#185937fb36fa) [Multi-View Anomaly Detection with Normalizing Flows](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#7859292225bb) [CLIP++: Scaling Down to Go Further](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#5eef51757623) [Maxing Out Medical Benchmarks](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#290ff5558613) [Why This Series Matters](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#297d3b1c3998) [Join Us for the Best of CVPR Series Meetups](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#3c7d42ccaccc) [What is next?](https://voxel51.com/blog/best-of-cvpr-2025-conversations-at-the-cutting-edge-of-ai#67c83e89a725) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## **CVPR 2025 Insights \#5**: Dialogues in AI Research As CVPR 2025 wrapped up in Nashville, a spark of inspiration stayed with us — not just from the keynotes, papers, and posters, but from the passionate individuals behind the science. That’s why we created the **Best of CVPR Series** — a series of meetups and short interviews spotlighting the authors behind some of the most intriguing research presented this year. These weren’t just technical Q&As. We discussed career journeys, real-world impact, dreams, and occasionally gave a shoutout to a mother who inspired it all. Join us on July 9th, 10th, and 11th to meet these authors in our online meetup series and dive deeper into their groundbreaking work. Now let’s meet the Researchers and their stories. ## OpticalNet: Seeing Beyond the Diffraction Limit **Ruyi and Benquan** from NTU Singapore and UT Austin are using AI to challenge physical limits in microscopy. Their fully software-based solution could disrupt semiconductor inspection with subwavelength imaging, eliminating the need for specialized optics. What excites them is that AI is starting to rival fundamental physics. They’re calling on the CVPR community to collaborate using their rich experimental datasets. > _“AI at the edge of Big Data really has the ability to challenge some physical laws.”_ _—_ **Benquan** ## Nonisotropic Gaussian Diffusion for Human Motion Prediction **Cecilia Corelli** from TUM brings a deeply human angle to technical innovation. Her work enables predictive motion modeling in robots, which is essential for elderly care and autonomous vehicles. With a solid grasp on mathematics and empathy, she reminds us that understanding AI means understanding your tools, not treating them like infallible gods. > _“Know your data. Know your tools. And stay grounded.”_ _—_ **Cecilia** ## Interpretable Medical AI with Human-in-the-Loop Models **Huy** from the University of Adelaide is building medical imaging tools that talk back with doctors. He’s bridging trust and performance in clinical AI by creating models that improve with expert feedback. Next up? Enabling doctors to chat with AI for collaborative diagnostics literally. > _“The power of AI and human is complementary, not competitive.”_ _—_ **Huy** ## From Occlusion to Imagination: OFER’s Face Reconstruction Breakthrough **Pratheba** of the Max Planck Institute is helping AI imagine faces behind occlusions using diffusion models. Initially sparked during her Microsoft internship in a COVID-era project, her work shows how inspiration meets technical depth. Her next goal? Tackling temporal consistency and audio-visual integration in videos > _“You fail at so many attempts, but you push through. Then suddenly… something clicks.” —_ **Pratheba** ## Driving with Foundation Models: A Vision for Autonomous Cars **Tin**, PhD at Porsche and KIT, is envisioning a “Knight Rider” world, where foundational models can comprehend traffic scenes with context and physics. His advice is deceptively simple: follow your gut and keep it fun. > _“Think simple. Don’t let people talk you into a niche. Just do what makes sense for you.” —_ **Tin** ## RANGE: Encoding the Earth’s Secrets **Ayush** from Washington University is pioneering representation learning for geospatial data. His RANGE model achieves strong performance on downstream tasks by leveraging the unique characteristics of GPS-based data. He’s now moving into generative territory, asking: What if we could generate our planet’s patterns? > _“Find exciting problems. Read widely. Look for the gaps. Then start filling them.”_ _—_ **Ayush** ## Few-Shot Agriculture Vision with Grounding DINO **Dr. Sudhir Sornapudi** from Corteva Agriscience is the leading innovator in AgTech. His team fine-tuned vision-language models for crop detection with only a handful of labels. Next? Optimizing for edge devices and applying LoRA techniques. He invites researchers to explore AgTech, where real-world impact meets scientific challenge. > _“Ag is complex — and that’s what makes it beautiful. Come work with biologists and change the world.”_ _—_ **Sudhir** ## SmartHome-Bench: Multimodal AI for Real-World Anomalies **Xinyi & Congjing** from the University of Washington created a benchmark for video anomaly detection using MLLMs. Their work, rooted in user-centric design, achieved a whopping 11% performance boost. Their next step? Personalizing anomaly detection based on household-specific preferences. > _“Don’t just study what happened. Study why it happened.”_ _—_ **Xinyi & Congjing** ## Multi-View Anomaly Detection with Normalizing Flows **Mathis** from Leibniz University Hannover is adding a new dimension — literally — to anomaly detection. His model achieves higher accuracy and better object understanding by leveraging multiple views. The next stop is flow matching and even 3D modeling. > _“Pick a topic you’re passionate about. Then it won’t feel like work.”_ _—_ **Mathis** ## CLIP++: Scaling Down to Go Further **Rui** from TUM trained CLIP-like models with synthetic captions and just 30 million images, outperforming the original at scale. His work unlocks new fine-grained representations for large language models and multimodal tasks. > _“You don’t need a billion images. You need a clever setup.”_ _—_ **Rui** ## Maxing Out Medical Benchmarks **Max Gutbrod** from OTH Regensburg introduced a new benchmark for out-of-distribution detection in medical imaging. His work reveals that what works in natural images might not translate to healthcare. Max is now returning to the core challenge — creating more trustworthy detection methods in medicine. > _“Medical imaging needs fairness, trust, and robust benchmarks. That’s where the real impact lies.”_ _—_ **Max** ## Why This Series Matters This series is more than just research summaries. It’s about people. Their stories, setbacks, motivations, and advice for the next generation. Whether you’re a seasoned researcher, an industry practitioner, or a student exploring AI, these voices offer inspiration and insight. ## Join Us for the Best of CVPR Series Meetups [RSVP Here](https://medium.com/@paularamos_phd/0b9cb3328f77#), [July 9th,](https://voxel51.com/events/best-of-cvpr-july-9-2025) [July 10th](https://voxel51.com/events/best-of-cvpr-july-10-2025), [July 11th](https://voxel51.com/events/best-of-cvpr-july-11-2025). [Explore the Blogs](https://medium.com/@paularamos_phd/0b9cb3328f77#), [Day 1 (July 9th)](https://medium.com/@paularamos_phd/the-best-of-cvpr-2025-series-day-1-7b1a7925da39), [Day 2 (July 10th)](https://medium.com/@paularamos_phd/the-best-of-cvpr-2025-series-day-2-b914d331e178), [Day 3 (July 11th)](https://medium.com/@paularamos_phd/the-best-of-cvpr-2025-series-day-3-85de08633c97) ## What is next? If you’re interested in following along as I dive deeper into the world of AI and continue to grow professionally, feel free to connect or follow me on [LinkedIn](https://www.linkedin.com/in/paula-ramos-phd/). Let’s inspire each other to embrace change and reach new heights! You can find me at some Voxel51 events ( [https://voxel51.com/computer-vision-events/](https://voxel51.com/computer-vision-events/)), or if you want to join this fantastic team, it’s worth taking a look at this page: [https://voxel51.com/jobs/](https://voxel51.com/jobs/) [CVPR](https://voxel51.com/blog/tag/cvpr) [robotics](https://voxel51.com/blog/tag/robotics) [human perception](https://voxel51.com/blog/tag/human-perception) [agriculture](https://voxel51.com/blog/tag/agriculture) [medical imaging](https://voxel51.com/blog/tag/medical-imaging) ![](https://cdn.sanity.io/images/h6toihm1/production/e926c07c7d1426c0fde8fdefa637c528d47b16f4-512x512.webp?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a558b86370f2f17212fb2f2c894d590101458a85-5760x3241.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Multimodal Frontier in Computer Vision, Medicine, and Agriculture— CVPR 2025 Reflections\\ \\ Event Recaps, Industry Solutions\\ \\ • \\ \\ Jun 24, 2025](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections) [![](https://cdn.sanity.io/images/h6toihm1/production/97cf3d887735ab9574b2b3e2d3825146016d875e-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Embodied Computer Vision at CVPR 2025: The Next AI Frontier\\ \\ Event Recaps\\ \\ • \\ \\ Jun 30, 2025](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier) [![](https://cdn.sanity.io/images/h6toihm1/production/e60eea36edcff16d65c62e2c3fea99a66d36a8d5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 @CVPR 2025: Smarter, Faster Visual AI\\ \\ Event Recaps, Product & News\\ \\ • \\ \\ Jun 3, 2025](https://voxel51.com/blog/cvpr-2025) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) Best of CVPR 2025: Conversations at the Cutting Edge of AI <|firecrawl-page-116-lllmstxt|> ## AI, ML, and Computer Vision Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/e9dd901fc82b4b27b0342db4fe50e1380153c507-960x541.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/e9dd901fc82b4b27b0342db4fe50e1380153c507-960x541.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event Virtual Americas Meetups AI, ML and Computer Vision Meetup - Aug 28, 2025 Aug 28, 2025 10 AM - 12 PM Pacific Online. Fill in the form to register! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/4a3ddde96c24e503d8d0554f301e27f8757a9a81-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Md Mostafijur Rahman UT Austin Bio ![](https://cdn.sanity.io/images/h6toihm1/production/fbae23c5265f673f7c16a24c5ae15e46739f72ec-240x241.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Elisa Chen Meta Bio ![](https://cdn.sanity.io/images/h6toihm1/production/d42432c84e43daad328ba19a006b2a8a3f59237e-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Dan Gural Voxel51 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/7eabf0b2f5620c60bdd119cc0c0e4ac5905f31c4-481x480.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Constantin Seibold University Hospital Heidelberg Bio About this event Hear talks from experts on the latest topics in AI, ML and Computer Vision! Schedule EffiDec3D: An Optimized Decoder for High-Performance and Efficient 3D Medical Image Segmentation ![](https://cdn.sanity.io/images/h6toihm1/production/4a3ddde96c24e503d8d0554f301e27f8757a9a81-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Md Mostafijur Rahman UT Austin Bio Recent 3D deep networks such as SwinUNETR, SwinUNETRv2, and 3D UX-Net have shown promising performance by leveraging self-attention and large-kernel convolutions to capture the volumetric context. However, their substantial computational requirements limit their use in real-time and resource-constrained environments. In this paper, we propose EffiDec3D, an optimized 3D decoder that employs a channel reduction strategy across all decoder stages and removes the high-resolution layers when their contribution to segmentation quality is minimal. Our optimized EffiDec3D decoder achieves a 96.4% reduction in #Params and a 93.0% reduction in #FLOPs compared to the decoder of original 3D UX-Net. Our extensive experiments on 12 different medical imaging tasks confirm that EffiDec3D not only significantly reduces the computational demands, but also maintains a performance level comparable to original models, thus establishing a new standard for efficient 3D medical image segmentation. Exploiting Vulnerabilities In CV Models Through Adversarial Attacks ![](https://cdn.sanity.io/images/h6toihm1/production/fbae23c5265f673f7c16a24c5ae15e46739f72ec-240x241.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Elisa Chen Meta Bio As AI and computer vision models are leveraged more broadly in society, we should be better prepared for adversarial attacks by bad actors. In this talk, we'll cover some of the common methods for performing adversarial attacks on CV models. Adversarial attacks are deliberate attempts to deceive neural networks into generating incorrect predictions by making subtle alterations to the input data. What Makes a Good AV Dataset? Lessons from the Front Lines of Sensor Calibration and Projection ![](https://cdn.sanity.io/images/h6toihm1/production/d42432c84e43daad328ba19a006b2a8a3f59237e-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Dan Gural Voxel51 Bio Getting autonomous vehicle data ready for real use, whether for training, simulation, or evaluation, isn’t just about collecting LIDAR and camera frames. It’s about making sure every point lands where it should, in the right frame, at the right time. In this talk, we’ll break down what it actually takes to go from raw logs to a clean, usable AV dataset. We’ll walk through the practical process of validating transformations, aligning coordinate systems, checking intrinsics and extrinsics, and making sure your projected points actually show up on camera images. Along the way, we’ll share a checklist of common failure points and hard-won debugging tips. Finally, we’ll show how doing this right unlocks downstream tools like Omniverse Nurec and Cosmos—enabling powerful workflows like digital reconstruction, simulation, and large-scale synthetic data generation Clustering in Computer Vision: From Theory to Applications ![](https://cdn.sanity.io/images/h6toihm1/production/7eabf0b2f5620c60bdd119cc0c0e4ac5905f31c4-481x480.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Constantin Seibold University Hospital Heidelberg Bio In today’s AI landscape, these techniques are crucial. Clustering methods help organize unstructured data into meaningful groups, aiding knowledge discovery, feature analysis, and retrieval-augmented generation. From k-means to DBSCAN and hierarchical approaches like FINCH, selecting the right method is key: including balancing scalability, managing noise sensitivity, and fitting computational demands. This presentation provides an in-depth exploration of the current state-of-the-art of clustering techniques with a strong focus on their applications within computer vision. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-117-lllmstxt|> ## Visual AI Event [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/63d5925d12404c349aa01929b8e5802b4a794f47-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/63d5925d12404c349aa01929b8e5802b4a794f47-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event Virtual Americas Meetups Manufacturing Visual AI in Manufacturing and Robotics - September 12, 2025 Sep 12, 2025 9 AM Pacific Online. Register for the Zoom! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/31a70fb4eaf83b0900c43e67cd46648ee8bb7364-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Jiafei Duan University of Washington / Ai2 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/671016056c10952d5725d483e97d1059e610d6f7-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Aimira Baitieva Valeo Bio ![](https://cdn.sanity.io/images/h6toihm1/production/861e3ef09449da22eb403a7926d93bcd14d31b3c-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Vlad Larichev Accenture Industry X Bio ![](https://cdn.sanity.io/images/h6toihm1/production/d99d54682a356a82c7e46b68f273fd8cbdd67a96-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Michael Hart D-Robotics Bio About this event Join us for a series of virtual events to hear talks from experts on the latest developments at the intersection of Visual AI, Manufacturing, and Robotics. Schedule Towards Robotics Foundation Models that Can Reason ![](https://cdn.sanity.io/images/h6toihm1/production/31a70fb4eaf83b0900c43e67cd46648ee8bb7364-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Jiafei Duan University of Washington / Ai2 Bio In recent years, we have witnessed remarkable progress in generative AI, particularly in language and visual understanding and generation. This leap has been fueled by unprecedentedly large image–text datasets and the scaling of large language and vision models trained on them. Increasingly, these advances are being leveraged to equip and empower robots with open-world visual understanding and reasoning capabilities. Yet, despite these advances, scaling such models for robotics remains challenging due to the scarcity of large-scale, high-quality robot interaction data, limiting their ability to generalize and truly reason about actions in the real world. Nonetheless, promising results are emerging from using multimodal large language models (MLLMs) as the backbone of robotic systems, especially in enabling the acquisition of low-level skills required for robust deployment in everyday household settings. In this talk, I will present three recent works that aim to bridge the gap between rich semantic world knowledge in MLLMs and actionable robot control. I will begin with AHA, a vision-language model that reasons about failures in robotic manipulation and improves the robustness of existing systems. Building on this, I will introduce SAM2Act, a 3D generalist robotic model with a memory-centric architecture capable of performing high-precision manipulation tasks while retaining and reasoning over past observations. Finally, I will present [MolmoAct](https://allenai.org/blog/molmoact), AI2’s flagship robotic foundation model for action reasoning, designed as a generalist system that can be post-trained for a wide range of downstream manipulation tasks. Beyond Academic Benchmarks: Critical Analysis and Best Practices for Visual Industrial Anomaly Detection ![](https://cdn.sanity.io/images/h6toihm1/production/671016056c10952d5725d483e97d1059e610d6f7-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Aimira Baitieva Valeo Bio In this talk, I will share our recent research efforts in visual industrial anomaly detection. It will present a comprehensive empirical analysis with a focus on real-world applications, demonstrating that recent SOTA methods perform worse than methods from 2021 when evaluated on a variety of datasets. We will also investigate how different practical aspects, such as input size, distribution shift, data contamination, and having a validation set, affect the results. The Digital Reasoning Thread in Manufacturing: Orchestrating Vision, Simulation, and Robotics ![](https://cdn.sanity.io/images/h6toihm1/production/861e3ef09449da22eb403a7926d93bcd14d31b3c-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Vlad Larichev Accenture Industry X Bio Manufacturing is entering a new phase where AI is no longer confined to isolated tasks like defect detection or predictive maintenance. Advances in reasoning AI, simulation, and robotics are converging to create end-to-end systems that can perceive, decide, and act – in both digital and physical environments. This talk introduces the Digital Reasoning Thread – a consistent layer of AI reasoning that runs through every stage of manufacturing, connecting visual intelligence, digital twins, simulation environments, and robotic execution. By linking perception with advanced reasoning and action, this approach enables faster, higher-quality decisions across the entire value chain. We will explore real-world examples of applying reasoning AI in industrial settings, combining simulation-driven analysis, orchestration frameworks, and the foundations needed for robotic execution in the physical world. Along the way, we will examine the key technical building blocks – from data pipelines and interoperability standards to agentic AI architectures – that make this level of integration possible. Attendees will gain a clear understanding of how to bridge AI-driven perception with simulation and robotics, and what it takes to move from isolated pilots to orchestrated, autonomous manufacturing systems. The Road to Useful Robots ![](https://cdn.sanity.io/images/h6toihm1/production/d99d54682a356a82c7e46b68f273fd8cbdd67a96-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Michael Hart D-Robotics Bio This talk explores the current state of AI-enabled robots and the issues with deploying more advanced models on constrained hardware, including limited compute and power budgets. It then moves on to what's next for developing useful, intelligent robots. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-118-lllmstxt|> ## Defect Detection Webinar August 21st at 12:00pm Preventing Critical Misses in Defect Detection: A Data-Centric Approach with MongoDB and Voxel51 Virtual [RSVP](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#rsvp) Text goes here [X](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#) ![](https://d24wuq6o951i2g.cloudfront.net/img/events/splash/shapes-highcontrast.png?stp) Mrs Robinson Wishes you a wonderful day Peter Lloyd is a host of exceptional ability. Studies show that a vast majority of guests attending events by Peter have been known to leave more elated than visitors to Santa's Workshop, The Lost of Continent of Atlantis, and the Fountain of Youth. Mrs Robinson Wishes you a wonderful day Peter Lloyd is a host of exceptional ability. Studies show that a vast majority of guests attending events by Peter have been known to leave more elated than visitors to Santa's Workshop, The Lost of Continent of Atlantis, and the Fountain of Youth. ![Logo](https://d3m889aznlr23d.cloudfront.net/img/events/id/459/459228401/assets/2e6f2f76ed5e6e15fd52325b9e5a8359.Group-1-1-.png) [Join The Webinar](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#rsvp) Text goes here [X](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#) Preventing Critical Misses in Defect Detection: A Data-Centric Approach with MongoDB and Voxel51 [RSVP](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#rsvp) Text goes here [X](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#) Thursday, August 21 12:00pm - 1:00pm CDT Live Webcast Preventing Critical Misses in Defect Detection: A Data-Centric Approach with MongoDB and Voxel51 [Join The Webinar](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#rsvp) Text goes here [X](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#) August 21st 12:00pm – 1:00pm Virtual ![](https://cached-services.splashthat.com/service/image-transform?src=//d3m889aznlr23d.cloudfront.net/img/events/id/458/458659429/assets/2d9dfc81af97e7c350cbab2330998fdd.what-to-expect.svg) [Save my Spot](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#rsvp) Text goes here [X](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#) Transform Your Inspection Process from Model-Centric to Data-Driven In critical industries, what you can't see can hurt you. Missed defects in industrial applications have far-reaching and expensive consequences. From micro-cracks in manufacturing to leaks in oil pipelines and stress fractures in wind turbines, undetected flaws lead to catastrophic failures, safety risks, and costly regulatory violations. While new computer vision models are improving how manufacturers and energy operators detect these issues, success doesn’t come from algorithms alone.  The real challenge lies in building workflows that connect powerful models with high-quality data, at scale, in a way your teams can actually use. Join this session where experts Dan Gural from Voxel51 and Robert Jones from MongoDB will show you how data-centric practices transform defect detection. Moving from general manufacturing to the specific needs of energy and industrial inspection, they will demonstrate how to master your visual data. You'll discover proven strategies to ensure data quality and build systems that enable accurate, scalable defect detection. Leave the session ready to move beyond model-centric thinking and start building robust, data-driven pipelines. What You'll Learn: - Grasp the critical impact of defect detection challenges across manufacturing, energy, and industrial sectors. - See how Visual AI pinpoints hard-to-find defects like cracks, corrosion, leaks, and surface anomalies with precision. - Design modern data architectures to efficiently manage massive visual inspection datasets. - Build robust data pipelines for consistently accurate models and effective production monitoring. - Learn practical strategies to integrate AI into existing inspection workflows without disrupting operations. Audience: - Digital Transformation Leaders - AI Engineers and Data Scientists working on industrial applications - Maintenance and Reliability Managers - Inspection and Quality Engineers - Operations Managers in manufacturing, energy, or infrastructure sectors [Register Now](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#rsvp) Text goes here [X](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#) Webinar Speakers This session will be led by technical experts from MongoDB and Voxel51 who will guide you through data-centric best practices. ![](https://d3m889aznlr23d.cloudfront.net/img/events/id/459/459228401/assets/317276a4343d07256ea2c02e898ab796.Daniel-Gural.jpg) Dan Gural Machine Learning Evangelist Voxel51 ![](https://d3m889aznlr23d.cloudfront.net/img/events/id/459/459228401/assets/1651a6ca1a68883739610787e2345f8b.Robert-Jones.jpg) Robert Jones Advisory Solutions Architect MongoDB Lorem ipsum dolor sit amet, consectetur adipiscing elit. Morbi in dapibus mi. Proin eget mi ut dolor tincidunt fermentum sit amet consectetur lacus. Read the blogs Check out related recaps, highlights, and latest announcements. [Read blogs >](https://www.mongodb.com/blog) Stay connected Follow us on LinkedIn to stay up-to-date on all things MongoDB. Follow us > Attend events Connect with the MongoDB team at an event near you! View events > The Final Countdown! Time left for the event 0 days 0 hours 0 minutes 00 seconds The countdown doesn't work if the event start date is set to TBD ![](https://d3m889aznlr23d.cloudfront.net/img/events/id/458/458658274/assets/b85cb09361f3422aac541cb343145e62.Headshot-1.png) Minnie Redding Minnie Redding will be speaking about their experience as an art director and organizer of design events. \[confirmation\_headline\] \[confirmation\_messaging\] [Add to Calendar](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#cal) Text goes here [X](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#) Agenda Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. ![](https://cached-services.splashthat.com/service/image-transform?src=//d3m889aznlr23d.cloudfront.net/img/events/id/458/458661103/assets/fb16e543e4972be509afd4f7c5479f62.Layer-1.svg) [Skip to content](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#to-content) [Share with Friends](https://events.mongodb.com/voxel51-defect-detection?utm_source=VXL&utm_medium=EML&utm_id=701Ue00000LhBFSIA3#) Facebook Twitter LinkedIn Link ### RSVP [![Google Icon](https://d24wuq6o951i2g.cloudfront.net/img/site-assets/google-icon.svg)\\ Google](https://events.mongodb.com/voxel51-defect-detection?calendar-Gmail&u=22271a3ad4a7eddb) [![Outlook Icon](https://d24wuq6o951i2g.cloudfront.net/img/site-assets/outlook-icon.svg)\\ Outlook](https://events.mongodb.com/voxel51-defect-detection?calendar-Outlook&u=22271a3ad4a7eddb) [![Apple Icon](https://d24wuq6o951i2g.cloudfront.net/img/site-assets/apple-icon.svg)\\ Apple](https://events.mongodb.com/voxel51-defect-detection?calendar-iCal&u=22271a3ad4a7eddb) [![Yahoo Icon](https://d24wuq6o951i2g.cloudfront.net/img/site-assets/yahoo-icon.svg)\\ Yahoo](https://events.mongodb.com/voxel51-defect-detection?calendar-Yahoo&u=22271a3ad4a7eddb) <|firecrawl-page-119-lllmstxt|> ## Cost Estimation Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Cost-Estimation [![](https://cdn.sanity.io/images/h6toihm1/production/990a31a85d850d41fb482ae29960a8fa10b18ecd-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Behind the math: how we built the annotation savings calculator\\ \\ Product & News\\ \\ • \\ \\ Jun 4, 2025](https://voxel51.com/blog/how-we-built-annotation-savings-estimator) ## Enough data wrangling.
 Request a demo. [Get started](https://voxel51.com/link-catcher) [Explore the Demo](https://voxel51.com/link-catcher) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-120-lllmstxt|> ## Visual AI for Security [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Visual AI in Security Identify safety and security threats instantly, accurately, and efficiently with solutions built using FiftyOne. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/7357d4b89ca05e262a4b01cdf9a5b0f0bf0978d0-628x395.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/ecd68d5ae6e4d889ccbba649f5def5ad7afd81f4-628x813.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/4f0bb4ba87b48dcd92b7f75c727f3ee88caa9ca7-628x817.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/b4739edcd078da2fcb351de2fb3e8f6919c09b7a-628x393.png?auto=format&dpr=2&fit=max&q=75&w=314) Benefits ## Streamline security & safety AI systems with FiftyOne Building computer vision solutions that keep people and property safe and secure requires being able to efficiently sift through large streams of data. FiftyOne helps you bring data insights that enable top model performance so no threat is left behind. ### Increase productivity Save ½ FTE or more of valuable engineering time so you can build more features your customers love ### Get to production faster Shave off 5+ months of development time while delivering better models with the high performance you need. ### Save money Save significant amounts of money on annotation by identifying just the samples you need to send for labeling. Use cases ## Visual AI use cases in security powered by FiftyOne Computer vision enables solutions to key use cases in security. That’s why leaders and innovators building security solutions rely on FiftyOne. ![](https://cdn.sanity.io/images/h6toihm1/production/5f9269b1629c1ed7a39f49208d018ddc6ae0b297-542x318.jpg?auto=format&dpr=2&fit=max&q=75&w=271) Real-time video analysis Use computer vision algorithms to analyze live video streams from security cameras to recognize potential threats and risks. ![](https://cdn.sanity.io/images/h6toihm1/production/69e3c278087cfb99853ba6fb91a62d4d0c09ce20-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Detection of people, vehicles, PPE, and more Apply detection and classification models to identify objects, people, and other items in visual data. ![](https://cdn.sanity.io/images/h6toihm1/production/9bb258ecad1d3819ebc32300adff8fb6ae535073-768x549.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Smoke and fire detection Identify and alert when fire is detected to increase safety and minimize damage. ![](https://cdn.sanity.io/images/h6toihm1/production/e4558dacfd1649267066090c793a829dfa02a424-768x467.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Motion detection and tracking Track people and objects across cameras and locations to better identify safety and security risks. ![](https://cdn.sanity.io/images/h6toihm1/production/5b7b1b76c984abd399ac40b2df6a2519e423f916-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Slip and fall detection Immediately recognize hazard detection and accidents across sites to enhance safety. ![](https://cdn.sanity.io/images/h6toihm1/production/0419c10b97ad4707762af86b836b9ce34fbde2c0-768x513.jpg?auto=format&dpr=2&fit=max&q=75&w=384) License plate detection and recognition Automatically detect identifying information for vehicles and equipment. Features ## How visual AI can help you ### Easily organize millions of incoming samples Your app relies on data, metadata, and labels that matter most to you. FiftyOne makes it easy to manage your samples across dozens of formats. Multiple data types: images, videos, clips, frames, geolocation, and 3D point clouds Any metadata you need: time of day, camera or device ID, location information, weather conditions, and anything else you need in your AI workflows Any model you’re working with: person tracking, object detection, action localization, and many more ![](https://cdn.sanity.io/images/h6toihm1/production/6f824d37fd55c2af92eb6cc07e4f4c6fce0f1923-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Quickly find the subsets of data you want Sifting through massive amounts of data is like searching for a needle in a haystack. With FiftyOne you can pinpoint samples of interest in seconds. ![](https://cdn.sanity.io/images/h6toihm1/production/45ba7cf572e77494727098d5909bb006a450f0dc-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Only annotate what you need Generating annotations can be complex, cumbersome, and costly. FiftyOne integrates with your favorite annotation tools to become your mission control for annotation workflows. Stop passing data around: collaborate with any human in the loop on a single source of truth Stop overpaying for annotations: identify your most valuable samples to annotate, then automatically send them to your annotation vendor Mitigate annotation mistakes: assess the quality of your annotations to improve both your datasets and models ![](https://cdn.sanity.io/images/h6toihm1/production/a8c9cf3580fb1bd8ff0489e06e7f4d682b4d96ca-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Continuously build models that perform Models don’t always perform on new, unseen data. FiftyOne gives you the ability to visualize and compare model performance so you can deploy into production with peace of mind. Understand your model’s failure modes: browse model performance at the sample level so you can take the right steps to fix them Embrace continuous evaluation: integrate FiftyOne into your training pipeline to evaluate and improve model performance and datasets with every model update Withstand the test of time: detect and manage drift in your datasets that naturally occur over time ![](https://cdn.sanity.io/images/h6toihm1/production/a736b3be1dec803e2ffc6edca02656aee07363ec-1536x1084.png?auto=format&dpr=2&fit=max&q=75&w=600) features ## Designed for AI builders FiftyOne natively supports and enables the computer vision building blocks needed to develop robust automotive AI solutions. ![](https://cdn.sanity.io/images/h6toihm1/production/b65da79f5913eb1e20235f7bc09363e94b9dca14-2560x960.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,2560,960&w=1280) - Classification - Detection - Segmentation - Polygons and polylines - Keypoints - Pointclouds - Heatmaps - Geolocation - Embeddings - Multiview datasets - Images, videos, and 3D data > “FiftyOne has helped us manage our huge datasets, collaborate on model evaluation, tighten our production schedule, and ultimately deliver solutions that help our customers better manage their risk. FiftyOne has added tremendous value to our computer vision processes.” > > **Philippe Sawaya** > > Director of Artificial Intellignce, ADT Commercial ![](https://cdn.sanity.io/images/h6toihm1/production/edd47f2994e95d3cf5cb91da52f5ac9a8bc738d4-720x720.png?auto=format&dpr=2&fit=max&q=75&w=100) > “It’s a great thing to have centralized dataset management in the form of FiftyOne, which really makes our lives easier when curating datasets and training new models. All the team members have the same view on the datasets, which ensures that everyone understands the data used to train the models. This really saves a few hours of back and forth between team members.” > > **Ivan Ralašić** > > CTO & Co-Founder, Forsight ![](https://cdn.sanity.io/images/h6toihm1/production/e4504939c1f6ff2d9c79b2e0367181ec187c88b4-1024x219.png?auto=format&dpr=2&fit=max&q=75&rect=0,9,1024,206&w=100) > “FiftyOne has helped us manage our huge datasets, collaborate on model evaluation, tighten our production schedule, and ultimately deliver solutions that help our customers better manage their risk. FiftyOne has added tremendous value to our computer vision processes.” > > **Lanny Lin** > > Sr. Director of AI & Data Science, Vivint ![](https://cdn.sanity.io/images/h6toihm1/production/abc37a33007a1a8026be33e42bdae5f8b981e3a9-1024x238.png?auto=format&dpr=2&fit=max&q=75&w=100) > “It’s a great thing to have centralized dataset management in the form of FiftyOne Teams, which really makes our lives easier when curating datasets and training new models. All the team members have the same view on the datasets, which ensures that everyone understands the data used to train the models. This really saves a few hours of back and forth between team members.” > > **Patrick Rowsome** > > Lead Computer Vision Engineer, Protex AI ![](https://cdn.sanity.io/images/h6toihm1/production/4bfcfa58759a470844fb87537ca042e819274a06-1024x229.png?auto=format&dpr=2&fit=max&q=75&w=100) Resources ## Learn more about visual AI in security Learn how leading companies are using visual AI to build solutions that ensure safety and security. ### How computer vision is changing safety & security Computer Vision and AI are driving innovation in the security industry, augmenting CCTV cameras, security checkpoints, search and rescue, and more. [Read more](https://voxel51.com/blog/how-computer-vision-is-changing-security) ![](https://cdn.sanity.io/images/h6toihm1/production/8c9980c3e5700224c5907404b6ee0af6f6e60cd1-1920x1200.png?auto=format&dpr=2&fit=max&q=75&w=960) ## Data eats models for lunch Talk to our computer vision experts to start building better datasets and models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-121-lllmstxt|> ## Deploying Computer Vision [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision) From Prototype to Production: What it Really Takes to Deploy Computer Vision in Manufacturing Aug 6, 2025 • 4 min read Article content In this article [The data imbalance no one talks about](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing#f575590ea8d4) [Real-world use cases: from micro to macro](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing#799e5c22dd68) [False positives aren’t the enemy](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing#c46a9a6440d2) [From prototypes to production](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing#0882964d745e) [The role of human-in-the-loop](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing#4b2191df9688) In this article [The data imbalance no one talks about](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing#f575590ea8d4) [Real-world use cases: from micro to macro](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing#799e5c22dd68) [False positives aren’t the enemy](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing#c46a9a6440d2) [From prototypes to production](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing#0882964d745e) [The role of human-in-the-loop](https://voxel51.com/blog/deploy-computer-vision-in-manufacturing#4b2191df9688) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Over the past few years, I’ve had the privilege of working closely with manufacturing teams that are deploying visual AI to solve real problems on the factory floor. From detecting subtle defects in circuit boards to identifying structural anomalies in large-scale assemblies, the use of computer vision in manufacturing is no longer at the experimental stage—it’s operational, as highlighted in our [2025 AI in Manufacturing Landscape](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape). For all the excitement around foundation models and automated visual inspection, the reality of deploying these systems at scale is far from plug-and-play. Most of the cost, complexity, and risk doesn’t lie in GPUs or code—it lies in the data. And in manufacturing, where defects are rare by design, that data is especially hard to come by. Here’s what we’ve learned at Voxel51 about building reliable AI defect detection systems in the real world, and how manufacturers are rethinking their approach to AI. ## The data imbalance no one talks about When people think of AI in manufacturing, they often picture high-speed vision systems spotting flaws with superhuman precision. But behind that glossy outcome is a deeply asymmetric data problem. On a typical production line, more than 99.99% of parts are perfectly fine. That’s exactly what you want from a manufacturing system. But from a machine learning standpoint, this is a nightmare. We’re trying to train a model to detect rare and often subtle anomalies—hairline cracks, misalignments, poor solder joints—when there might be only one true defect in every 200,000 examples. That extreme imbalance is a bottleneck not just for modeling, but for data collection and labeling. It’s also why so many efforts to use off-the-shelf models fail. ## Real-world use cases: from micro to macro At Voxel51, we’ve worked with customers detecting defects at multiple scales. On the macro end, we’ve seen computer vision used to catch panel misalignments on vehicle assembly lines and identify damage to large mechanical components. On the micro scale, manufacturers are inspecting PCB boards for imperfect solder joints, missing components, or subtle wear that could lead to failure. In all of these cases, the process starts with a visual media pipeline—RGB, depth, sometimes infrared. That data flows into a visual dataset platform like [FiftyOne](https://voxel51.com/fiftyone), where QA teams or machine learning engineers [inspect, sort, and flag anomalies](https://voxel51.com/blog/anomaly-detection-with-fiftyone-and-anomalib). These aren’t giant labeling farms. They’re small, focused teams identifying the few meaningful edge cases that make all the difference in training. And those hard examples? They become the gold nuggets—the foundation for model improvement. ## False positives aren’t the enemy One of the most surprising insights from our [research on auto-labeling](https://voxel51.com/whitepapers/auto-labeling-data-for-object-detection) is that false positives are often less harmful than we think, especially in early-stage model training. With our [Verified Auto Labeling](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance) work, we’ve studied the performance of models when trained on noisy labels produced by foundation models. What we found is that it’s better to have a wide range of data with some incorrect labels than a small, perfectly labeled dataset. Bigger is often better than better, so long as you have mechanisms to verify and improve iteratively. That’s particularly relevant in manufacturing, where getting more of the right data (i.e., real defect cases) is often infeasible. Instead, teams are increasingly augmenting rare examples or simulating variability in appearance, lighting, or viewpoint. The goal isn’t perfection—it’s coverage. ## From prototypes to production Many teams see early success with their first AI prototypes. A model achieves 80% accuracy after just a few weeks of training, and the proof-of-concept demo looks great. But 80% isn’t good enough when you’re shipping products at scale. If your system flags 1 in every 5 defects incorrectly—or worse, misses them altogether—you still need humans reviewing every frame. So how do you move beyond 80%? One key strategy we’ve seen is scenario and failure mode analysis. Instead of just evaluating a model’s top-line accuracy, teams inspect how it performs across different edge cases. Where is it weak? Which conditions consistently trip it up? Should we collect more data, refine the labels, or change the model architecture? This isn’t glamorous work. It’s data-centric, engineering-heavy, and often iterative. But it’s the work that separates toy models from real production systems. ## The role of human-in-the-loop In manufacturing, we often hear the concern: “Will AI replace human inspectors?” But in practice, the most successful systems are built around human-in-the-loop design. Models can be used to triage incoming data—automatically passing along the confident “green” cases, flagging the uncertain “yellow” ones for review, and discarding or quarantining the obvious errors. This workflow saves time, reduces QA fatigue, and ensures that humans focus their attention where it matters most. Over time, those reviewed edge cases improve the model, which leads to fewer yellow flags, and more automation. But the human role doesn’t disappear—it becomes more strategic. Defect detection is one of the most compelling applications of computer vision in manufacturing, but it’s also one of the hardest to get right. The challenges aren’t algorithmic—they’re data-driven. You’re not just building a model; you’re building a system that can learn from the rare, the subtle, and the complex. If there’s one takeaway from the work we’ve done at Voxel51, it’s this: the teams that succeed treat data not as a byproduct, but as a product. They invest in the right tools, build feedback loops, and design with humans in mind. Because at the end of the day, good AI doesn’t just see the world—it learns from it. I recently joined Manufacturing Tomorrow, an Ohio State University podcast, to discuss this in detail. [Listen to the full episode](https://podcast.osu.edu/mfgtmw/2025/07/14/jason-corso-voxel51/) if you’re bringing visual AI applications from concept to the production floor. [manufacturing](https://voxel51.com/blog/tag/manufacturing) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) ![](https://cdn.sanity.io/images/h6toihm1/production/ad9fb967c5455e0f763411fb81956767d7f26482-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Jason Corso Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/6af33def6d297e2382d387e224e16451c95876af-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Why the “Annotate Everything” Era in Automotive AI Is Over\\ \\ Computer Vision\\ \\ • \\ \\ Jul 24, 2025](https://voxel51.com/blog/smarter-automotive-datasets-selection) [![](https://cdn.sanity.io/images/h6toihm1/production/3a661345dfbc596f7118a7b8ec8375ea1f73d138-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 3\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-3) [![](https://cdn.sanity.io/images/h6toihm1/production/047b21a97f6c858334f9f35ed89fa7655ebf5767-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ State-of-the-Art Object Detection with YOLO-NAS & FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ May 4, 2023](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-122-lllmstxt|> ## Visual AI for Defense [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Visual AI in Defense Streamline defense AI development by turning complex visual data into mission-ready machine learning models and vision AI systems. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/082655963a87aa267b9f3fb8ccf1a3bf6face91f-628x395.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/85895c230e09bf6ad4af5d3c687a4d867d843690-628x813.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/499f02bc413b4586e0fe4e1245212848949e2218-628x817.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/01181c9b4ee1b200aba8937aa092d7ebf5deb312-628x393.png?auto=format&dpr=2&fit=max&q=75&w=314) Benefits ## Build defense solutions with Voxel51 FiftyOne helps AI builders manage visual data throughout the development process. Instead of using fragile, homegrown solutions, or inflexible, bloated offerings, AI teams get precisely the tools they need to explore, visualize, and curate high-quality datasets while evaluating computer vision models—in one secure, flexible, and scalable platform. 0% improve productivity 0% increase accuracy & reliability Use cases ## Mission-critical defense applications of visual AI FiftyOne streamlines how teams handle data in visual AI development. By automating data management and providing tools to explore, visualize, and evaluate computer vision models, it helps developers build reliable AI solutions faster – even when working with massive datasets from various sources. ![](https://cdn.sanity.io/images/h6toihm1/production/7ca7407385ade5887632ec169015165af58b7429-768x432.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Reconnaissance & threat detection Build solutions that process multispectral data from satellites, drones, cameras, and sensors to identify, alert, and track potential risks and changes. ![](https://cdn.sanity.io/images/h6toihm1/production/0c505091d77eed31e92df923f107a30614213c2a-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Damage detection & assessment Build visual AI applications that can quickly assess and characterize damage to assets and targets from a combination of geospatial, imaging, and other data. ![](https://cdn.sanity.io/images/h6toihm1/production/4ce9984136e3d44d8e7b12e4a99ce5c8b154d346-768x576.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Autonomous navigation Develop computer vision and AI applications that enable autonomous navigation for drones and other vehicles that can understand and adapt to real-world conditions, from the base to the battlefield. ![](https://cdn.sanity.io/images/h6toihm1/production/02c0b877ff0082154ec7a3d9fa2f9dbcd0fd8fbc-600x400.jpg?auto=format&dpr=2&fit=max&q=75&w=300) Security monitoring Power applications that process surveillance video, images and scans to detect intrusions and suspect behavior to identify and alert on risks and threats. ![](https://cdn.sanity.io/images/h6toihm1/production/9bb258ecad1d3819ebc32300adff8fb6ae535073-768x549.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Search and rescue Deploy applications that process multiple data sources to rapidly locate and assess conditions of people and resources in dynamic environments. ![](https://cdn.sanity.io/images/h6toihm1/production/fd63bd8635f8a5f47cdeda6868bb8ab35a65bce5-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Product assembly & manufacturing Computer vision brings high precision to automated manufacturing quality controls, as well as the loading, preparation, and assembly of raw materials processed by machines. features ## Supports all popular vision tasks and media types Use FiftyOne to store metadata about your samples, like annotations and model predictions, so you can easily pinpoint and optimize scenarios of interest whenever needed. ![](https://cdn.sanity.io/images/h6toihm1/production/4c7bfcece3ca6b5e8eddc75e5194e9be81e0ebe1-3840x1440.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,3840,1440&w=1920) - Classification - Detection - Segmentation - Polygons and polylines - Keypoints - Pointclouds - Heatmaps - Geolocation - Embeddings - Multiview datasets - Images, videos, and 3D data Why FiftyOne ## Key capabilities of FiftyOne ### Fully secure Deploy in networked or fully isolated, air-gapped environments and integrate with your preferred auth solution to support the highest levels of security ### Multimodal data Bring together actual and synthetic image, video, radar, lidar, SAR, and other visual and spatial data for refinement and curation in an integrated environment ### Data exploration Search, view, slice, and dice data to quickly and easily pinpoint samples and scenarios that match your criteria and needs ### Dataset curation Create and refine high-quality datasets for use throughout model development for tuning and evaluation across the full range of scenarios and conditions ### Workflow automation Automate AI/ML operations with prebuilt workflows and custom plugins, including scheduling on a connected GPU/CPU cluster ### Flexibility & extensibility Utilize our library of existing plugins or create your own to fully customize and extend capabilities to fit your specific needs ### Natural language interaction Use the VoxelGPT plugin to find answers, explore, and organize your data, and get helpful AI/ML tips and docs, using natural language ### Team collaboration Safely collaborate at all stages of development, within and across teams, by sharing curated datasets, views, and analyses securely among team members ## Data eats models for lunch Talk to our computer vision experts to start building better datasets and models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-123-lllmstxt|> ## FiftyOne Workshop for Manufacturing [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/286af5bdd8c96bc26657a4ed7234c5eb6565f793-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/286af5bdd8c96bc26657a4ed7234c5eb6565f793-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event Virtual Americas Webinars & Workshops Manufacturing Getting Started with FiftyOne for Manufacturing Use Cases - Sept 30, 2025 Sep 30, 2025 9 AM Pacific Online. Register for the Zoom! Host ![](https://cdn.sanity.io/images/h6toihm1/production/97219ce91c701d6c28cb5dc462c55648cc906607-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Paula Ramos Voxel51 Bio Are you working with computer vision in manufacturing and need deeper visibility into your datasets and models? Join us for a free 90-minute hands-on workshop and learn how to leverage the open-source FiftyOne toolset to optimize your visual AI workflows, from anomaly detection on the production line to worker safety and quality assurance in additive manufacturing. **In this session, you'll learn how to:** - Visualize and audit complex manufacturing datasets. - Explore visual embeddings for failure mode analysis. - Identify and fix labeling issues affecting production models. - Perform advanced data curation for specialized use cases. - Integrate with annotation tools, model pipelines, and plugins. We'll take a data-centric approach to computer vision, starting with importing and exploring industrial visual data, including defects, wear patterns, and worker posture. You'll learn to query and filter datasets to surface edge cases, then use plugins and native integrations to streamline workflows. We'll walk through generating candidate ground truth labels and evaluating fine-tuned foundational models — particularly relevant to manufacturers using pre-trained models for tasks like defect segmentation or object localization in dynamic environments. By the end, you'll see how the FiftyOne App and SDK work together to enable more profound insight into visual AI systems. We'll conclude with a demo showcasing 3D view reconstruction for industrial inspection, revealing how Visual AI can bridge physical and digital layers of your production process. **Prerequisites:** Basic knowledge of Python and computer vision fundamentals. **Resources Provided:** All attendees will receive access to tutorials, videos, and the workshop codebase. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-124-lllmstxt|> ## Exploring Uncommon 3D Objects [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets), [Event Recaps](https://voxel51.com/blog/category/event-recaps) UnCommon Objects in 3D Jun 18, 2025 • 6 min read Article content In this article [Meet the Most Comprehensive Real-World 3D Dataset Ever Created](https://voxel51.com/blog/uncommon-objects-in-3d#9119ef40c661) [Diversity That Reflects Reality](https://voxel51.com/blog/uncommon-objects-in-3d#c6bdbe0d37bf) [Technical Innovations That Set uCO3D Apart](https://voxel51.com/blog/uncommon-objects-in-3d#bcda49f4ec63) [A New Standard in 3D Data](https://voxel51.com/blog/uncommon-objects-in-3d#4d7821442cba) [Getting Started with the Manageable Preview Subset](https://voxel51.com/blog/uncommon-objects-in-3d#832a594c870e) [Understanding the Data Structure](https://voxel51.com/blog/uncommon-objects-in-3d#016581e4492f) [Converting PointClouds to FiftyOne 3D Scenes](https://voxel51.com/blog/uncommon-objects-in-3d#109045baf95b) [Building a Unified Dataset with Proper Relationships](https://voxel51.com/blog/uncommon-objects-in-3d#1ee69bd19ae5) [Running the Complete Parsing Process](https://voxel51.com/blog/uncommon-objects-in-3d#a2a9716407a2) [Exploring Your Processed Dataset](https://voxel51.com/blog/uncommon-objects-in-3d#91dfd6fe7a6c) [Why This Dataset Matters for Your Research](https://voxel51.com/blog/uncommon-objects-in-3d#6ce6d44c3524) [Load the dataset in FiftyOne format directly from the Hugging Face Hub](https://voxel51.com/blog/uncommon-objects-in-3d#3eb79f7548f2) [Unleashing the Full Potential of uCO3D](https://voxel51.com/blog/uncommon-objects-in-3d#ac7184eb7079) [A New Era for 3D Computer Vision](https://voxel51.com/blog/uncommon-objects-in-3d#197b48b6e6dc) In this article [Meet the Most Comprehensive Real-World 3D Dataset Ever Created](https://voxel51.com/blog/uncommon-objects-in-3d#9119ef40c661) [Diversity That Reflects Reality](https://voxel51.com/blog/uncommon-objects-in-3d#c6bdbe0d37bf) [Technical Innovations That Set uCO3D Apart](https://voxel51.com/blog/uncommon-objects-in-3d#bcda49f4ec63) [A New Standard in 3D Data](https://voxel51.com/blog/uncommon-objects-in-3d#4d7821442cba) [Getting Started with the Manageable Preview Subset](https://voxel51.com/blog/uncommon-objects-in-3d#832a594c870e) [Understanding the Data Structure](https://voxel51.com/blog/uncommon-objects-in-3d#016581e4492f) [Converting PointClouds to FiftyOne 3D Scenes](https://voxel51.com/blog/uncommon-objects-in-3d#109045baf95b) [Building a Unified Dataset with Proper Relationships](https://voxel51.com/blog/uncommon-objects-in-3d#1ee69bd19ae5) [Running the Complete Parsing Process](https://voxel51.com/blog/uncommon-objects-in-3d#a2a9716407a2) [Exploring Your Processed Dataset](https://voxel51.com/blog/uncommon-objects-in-3d#91dfd6fe7a6c) [Why This Dataset Matters for Your Research](https://voxel51.com/blog/uncommon-objects-in-3d#6ce6d44c3524) [Load the dataset in FiftyOne format directly from the Hugging Face Hub](https://voxel51.com/blog/uncommon-objects-in-3d#3eb79f7548f2) [Unleashing the Full Potential of uCO3D](https://voxel51.com/blog/uncommon-objects-in-3d#ac7184eb7079) [A New Era for 3D Computer Vision](https://voxel51.com/blog/uncommon-objects-in-3d#197b48b6e6dc) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Meet the Most Comprehensive Real-World 3D Dataset Ever Created UnCommon Objects in 3D loaded into FiftyOne Meta AI’s [uCO3D dataset](https://uco3d.github.io/), which I saw at CVPR 2025, is the most significant advancement in real-world 3D object data collection we’ve seen in years. With 170,000 meticulously captured objects across more than 1,000 categories, [uCO3D](https://arxiv.org/abs/2501.07574) finally solves the persistent dilemma that has plagued 3D vision research: choosing between scale and quality. Previous datasets like CO3Dv2 offered decent quality but limited diversity (just 50 categories), while MVImgNet provided more objects but sacrificed the crucial 360° coverage needed for complete 3D understanding. Meta’s approach combines the best of both worlds while eliminating their respective limitations. This dataset will undoubtedly become the new gold standard for training 3D vision models. ## Diversity That Reflects Reality The 1,000+ object categories in uCO3D deliver an unprecedented breadth of real-world objects that previous datasets simply couldn’t match. By adopting the [LVIS](https://huggingface.co/datasets/Voxel51/LVIS) taxonomy, which deliberately includes long-tail categories, uCO3D captures the true diversity of objects in our world rather than just focusing on common items. This matters enormously because real-world applications need to recognize and understand the full spectrum of objects we encounter, not just the most frequent ones. The category distribution is thoughtfully organized into 50 super-categories, each containing approximately 20 subcategories, making it both comprehensive and structured. Researchers no longer need to accept limited category coverage as an inevitable constraint. ## Technical Innovations That Set uCO3D Apart Every aspect of uCO3D’s creation process incorporates state-of-the-art techniques that previous datasets simply couldn’t match. Instead of using the industry-standard COLMAP for structure-from-motion, Meta AI employed [VGGSfM](https://github.com/facebookresearch/vggsfm) to achieve significantly more accurate camera parameters and point clouds. The segmentation pipeline combines [text-conditioned Segment-Anything (langSAM)](https://github.com/luca-medeiros/lang-segment-anything) with XMem to achieve temporal consistency, thereby resolving the flickering mask issues that plagued previous datasets. Perhaps most impressively, each object includes a complete 3D Gaussian Splat reconstruction, enabling photorealistic novel-view synthesis and canonical-view rendering that was previously impossible with real-world data. These technical choices aren’t just incremental improvements — they’re revolutionary advances. ## A New Standard in 3D Data uCO3D fundamentally redefines what’s possible in 3D computer vision. The combination of unprecedented scale, diversity, and quality addresses all the major limitations that have held back progress in this field. The additional innovations like 3D Gaussian Splat reconstructions and rigorous quality control set a new standard that future datasets will struggle to match. The performance improvements demonstrated across multiple benchmarks provide empirical validation that this dataset delivers on its promises. This is quite simply the dataset that the 3D vision community has been waiting for. ## Getting Started with the Manageable Preview Subset Meta AI brilliantly provides a 52-video preview subset that lets you dive in without downloading the full 19TB dataset. At just 9.6 GB, this subset is ideal for developing your processing pipeline while still representing the dataset’s diversity. You can grab it with a simple command from their GitHub repository: ```python 1git clone git@github.com:facebookresearch/uco3d.git 2python dataset_download/download_dataset.py --download_small_subset --download_folder ./uco3d_subset ``` The preview subset contains the same rich data structure as the full dataset, making it invaluable for initial experimentation. ## Understanding the Data Structure Each object in uCO3D follows a consistent organization that makes programmatic processing straightforward. The dataset organizes objects in a category/subcategory/object\_id hierarchy, with each object directory containing high-resolution video ( `*_video.mp4` ), segmentation masks ( `mask_video.mp4` ), point clouds ( `segmented_point_cloud.ply` ), and camera parameters. The full dataset includes an additional 3D Gaussian Splat reconstructions that enable photorealistic novel view synthesis. This consistent structure makes it easy to iterate through objects and extract the data you need. This logical organization is crucial for efficiently working with such a large-scale dataset. ## Converting PointClouds to FiftyOne 3D Scenes Our first step in parsing uCO3D is converting the point cloud PLY files into interactive FiftyOne 3D scenes. ```python 1def process_point_cloud(point_cloud_path): 2 """Process a PLY point cloud file and create a FiftyOne 3D scene.""" 3 # Convert relative path to absolute path for consistency 4 abs_point_cloud_path = os.path.abspath(point_cloud_path) 5 6 # Extract the directory path where we'll save the output 7 dir_path = os.path.dirname(abs_point_cloud_path) 8 9 # Create a new FiftyOne 3D scene 10 scene = fo.Scene() 11 # Set camera to use Y-up coordinate system which is standard for 3D 12 scene.camera = fo.PerspectiveCamera(up="Y") 13 14 material = fo.PointCloudMaterial( 15 shading_mode="rgb", 16 attenuate_by_distance=True) 17 18 # Create a mesh from the point cloud with RGB coloring 19 mesh = fo.PlyMesh( 20 "mesh", # Name of the mesh in the scene 21 abs_point_cloud_path, 22 is_point_cloud=True, # Treat as point cloud rather than mesh 23 default_material=material, # Use RGB colors from PLY file 24 center_geometry=False, 25 ) 26 # Add the mesh to our scene 27 scene.add(mesh) 28 29 # Save the scene as a .fo3d file in the same directory 30 output_path = os.path.join(dir_path, "scene.fo3d") 31 scene.write(output_path) ``` This function transforms raw PLY files into interactive 3D scenes that you can rotate, zoom, and explore in your browser. ## Building a Unified Dataset with Proper Relationships The heart of our parsing script establishes relationships between different views of the same object. ```python 1def create_dataset(name): 2 """Create a FiftyOne dataset with appropriate schema for UCO3D.""" 3 dataset = fo.Dataset( 4 name=name, 5 persistent=True, # Save dataset to disk 6 overwrite=True # Replace existing dataset if it exists 7 ) 8 9 # Define group field to organize different views of the same object 10 dataset.add_group_field("group", default="rgb") 11 return dataset 12 13def process_objects(root_dir, dataset): 14 """Process all object directories and add them to the FiftyOne dataset.""" 15 processed_count = 0 16 17 # Walk through all category/subcategory/object_id directories 18 for dirpath, _, filenames in os.walk(root_dir): 19 # Skip directories missing required files 20 if not all(any(f.endswith(ext) for f in filenames) 21 for ext in ["scene.fo3d", "mask_video.mp4"]): 22 continue 23 24 # Find RGB video file - must end with _video.mp4 but not be mask_video.mp4 25 rgb_videos = [f for f in filenames if f.endswith("_video.mp4") and f != "mask_video.mp4"] 26 if not rgb_videos: 27 continue 28 29 # Extract category/subcategory/object_id from relative path 30 rel_path = os.path.relpath(dirpath, root_dir) 31 path_parts = rel_path.split(os.path.sep) 32 33 # Ensure we have category/subcategory/object_id structure 34 if len(path_parts) < 3: 35 continue 36 37 # Extract metadata from path components 38 category = path_parts[0] 39 subcategory = path_parts[1] 40 object_id = path_parts[2] 41 42 # Create a group to link different views of the same object 43 group = fo.Group() 44 45 # Create samples for RGB, mask, and point cloud views 46 samples = [\ 47 fo.Sample(\ 48 filepath=os.path.join(dirpath, rgb_videos[0]),\ 49 group=group.element("rgb"),\ 50 category=fo.Classification(label=category),\ 51 subcategory=fo.Classification(label=subcategory),\ 52 object_id=object_id\ 53 ),\ 54 fo.Sample(\ 55 filepath=os.path.join(dirpath, "mask_video.mp4"),\ 56 group=group.element("mask"),\ 57 category=fo.Classification(label=category),\ 58 subcategory=fo.Classification(label=subcategory),\ 59 object_id=object_id\ 60 ),\ 61 fo.Sample(\ 62 filepath=os.path.join(dirpath, "scene.fo3d"),\ 63 group=group.element("point_cloud"),\ 64 category=fo.Classification(label=category),\ 65 subcategory=fo.Classification(label=subcategory),\ 66 object_id=object_id\ 67 )\ 68 ] 69 70 # Add all samples for this object to the dataset 71 dataset.add_samples(samples) 72 processed_count += 1 ``` This powerful approach links RGB videos, segmentation masks, and 3D point clouds into a unified dataset with meaningful relationships. ## Running the Complete Parsing Process The main function ties everything together into a streamlined end-to-end process. ```python 1if __name__ == "__main__": 2 # Create dataset 3 DATASET_NAME = "UCO3D_Dataset" 4 dataset = create_dataset(DATASET_NAME) 5 6 # Process all objects in the workspace 7 root_dir = os.path.abspath(".") 8 processed_count = process_objects(root_dir, dataset) 9 10 # Compute metadata and save 11 dataset.compute_metadata() 12 dataset.save() 13 14 print(f"Created dataset '{DATASET_NAME}' with {processed_count} objects ({len(dataset)} samples)") ``` This code creates a persistent FiftyOne dataset that you can load in future sessions, making it a one-time processing effort. ## Exploring Your Processed Dataset Once parsed, FiftyOne transforms how you interact with uCO3D’s rich multimodal data. ```python 1import fiftyone as fo 2 3# Load your processed dataset 4dataset = fo.load_dataset("UCO3D_Dataset") 5 6# View dataset in the FiftyOne App 7session = fo.launch_app(dataset) 8 9# Filter to see only objects from a specific category 10kitchenware = dataset.match(F("category").label == "beverages") 11session.view = kitchenware 12 13# View only 3D point clouds 14point_clouds = dataset.match(F("group").name == "point_cloud") 15session.view = point_clouds 16 ``` FiftyOne’s interactive visualization tools let you explore videos, masks, and 3D point clouds with unprecedented ease. ## Why This Dataset Matters for Your Research Models trained on uCO3D consistently outperform those trained on previous datasets across all benchmarks. Few-view reconstruction models, such as [LightplaneLRM](https://github.com/facebookresearch/lightplan), show PSNR improvements of more than a full point when trained on uCO3D instead of [CO3Dv2](https://github.com/facebookresearch/co3d) or [MVImgNet](https://gaplab.cuhk.edu.cn/projects/MVImgNet/). Novel view synthesis using CAT3D-like approaches sees LPIPS error reductions of 5–20%. Most impressively, uCO3D’s 3D Gaussian Splat reconstructions enable the training of text-to-3D models, such as Instant3D, using real-world data for the first time, resulting in dramatically more realistic generations. These performance gains translate directly to better applications across AR/VR, robotics, and e-commerce. ## Load the dataset in FiftyOne format directly from the Hugging Face Hub I’ve already parsed this dataset for you! To download the parsed dataset directly from Hugging Face: ```python 1import fiftyone as fo 2from fiftyone.utils.huggingface import load_from_hub 3 4# Load the dataset 5dataset = load_from_hub("Voxel51/uco3d") 6 7# Launch the App 8session = fo.launch_app(dataset) ``` ## Unleashing the Full Potential of uCO3D The combination of uCO3D’s revolutionary data quality and FiftyOne’s exploration capabilities creates a development environment that accelerates 3D vision research. By structuring the dataset with proper relationships between different views, our parsing approach makes it simple to build multi-view training pipelines. The integration with FiftyOne lets you visually inspect reconstruction quality, helping you understand edge cases and failure modes. For anyone working in 3D computer vision, this parsed dataset becomes an invaluable resource that dramatically shortens the path from idea to implementation. This integration represents the new gold standard for working with large-scale 3D datasets. ## A New Era for 3D Computer Vision uCO3D fundamentally redefines what’s possible in 3D vision research, and our parsing pipeline makes it accessible. With its unprecedented combination of scale, diversity, and quality, uCO3D enables a new generation of 3D models that can generalize across thousands of object categories. By converting this dataset into FiftyOne format, we’ve created a structured, interactive environment that makes it easy to explore the data, understand its characteristics, and build sophisticated training pipelines. For researchers and developers in 3D vision, this parsed dataset eliminates countless hours of preprocessing, allowing you to focus on the actual innovation. The future of 3D vision is here — it’s just waiting for you to parse it. [multi-modal AI](https://voxel51.com/blog/tag/multi-modal-ai) [3D point cloud](https://voxel51.com/blog/tag/3d-point-cloud) [datasets](https://voxel51.com/blog/tag/datasets) [CVPR](https://voxel51.com/blog/tag/cvpr) ![](https://cdn.sanity.io/images/h6toihm1/production/a41a0477c7a98264f600772e9568607d070eea59-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Sahota Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/7ec61c80f387b16f24b4b2fe33f804264a38b486-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Rethinking How We Evaluate Multimodal AI\\ \\ Event Recaps\\ \\ • \\ \\ Jun 12, 2025](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/e468545aa08daf6c6829d2593ffd8b5457c7dee5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ NVIDIA’s C-RADIOv3 is the Vision Encoder You Should Be Using\\ \\ Event Recaps, Integrations\\ \\ • \\ \\ Jun 23, 2025](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using) [![](https://cdn.sanity.io/images/h6toihm1/production/3f54c19e72060ff2fa99673840a472f4d6e85c8e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Van der Maaten’s Three-System Roadmap to AGI Is Brilliantly Pragmatic\\ \\ Event Recaps\\ \\ • \\ \\ Jun 17, 2025](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-125-lllmstxt|> ## Madrid AI & ML Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/63e76cde1a45aed8d10a5bcdec5a34ceb15268d3-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/63e76cde1a45aed8d10a5bcdec5a34ceb15268d3-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event In-person EMEA Meetups Madrid AI, ML and Computer Vision Meetup - September 26, 2025 Sep 26, 2025 6:30 - 10:00 PM Google For Startups Campus C. de Moreno Nieto, 2, Arganzuela 28005, Madrid Spain Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/546ad7c42bcfd19086efabfcedcba2eb383ec3fe-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Sergio Paniego Blanco Hugging Face Bio ![](https://cdn.sanity.io/images/h6toihm1/production/b3973bbcce3bcb66a5ed3ad6f343a6dd33bf4e8a-481x481.png?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Voxel51 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/ae0baff6b16d895cd5b929c1380a795f9ac5073e-481x481.png?auto=format&dpr=2&fit=max&q=75&w=42) Máximo Fernández Núñez Machine Learning Engineer Bio ![](https://cdn.sanity.io/images/h6toihm1/production/33fc40d663f663f4e54b2e2b1a16580956d7124b-481x481.png?auto=format&dpr=2&fit=max&q=75&w=42) Hind Azegrouz Intel Bio About this event Acompáñanos para escuchar charlas de expertos en IA, ML y Visión por Computadora. Abrimos puertas a las 18:15 y empezar a las 18:30. Hasta las 20:30 las charlas y luego de 20:30 a 21:30-22:00 networking con el catering de Google. **Evento patrocinado por:** Voxel51 y [Arcasiles Group](https://lu.ma/arcasilesgroup?period=past) Schedule Multimodality at Hugging Face ![](https://cdn.sanity.io/images/h6toihm1/production/546ad7c42bcfd19086efabfcedcba2eb383ec3fe-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Sergio Paniego Blanco Hugging Face Bio In this talk, we’ll explore the latest advances in multimodal AI within the Hugging Face ecosystem. From vision-language models to emerging Omni models, we’ll dive into cutting-edge architectures powering this space. We’ll also take a look at the tools and libraries Hugging Face provides to support end-to-end multimodal workflows. Tus Datos te Están Mintiendo: Búsqueda Semántica Para Encontrar la Verdad ![](https://cdn.sanity.io/images/h6toihm1/production/b3973bbcce3bcb66a5ed3ad6f343a6dd33bf4e8a-481x481.png?auto=format&dpr=2&fit=max&q=75&w=96) Paula Ramos Voxel51 Bio Los modelos de alto rendimiento comienzan con datos de alta calidad, pero encontrar muestras ruidosas, mal etiquetadas o casos límite dentro de conjuntos de datos masivos sigue siendo un gran obstáculo. En esta sesión, exploraremos un enfoque escalable para curar y refinar conjuntos de datos visuales a gran escala utilizando búsqueda semántica impulsada por embeddings basados en transformers. Al aprovechar la búsqueda por similitud y el aprendizaje de representaciones multimodales, aprenderás a descubrir patrones ocultos, detectar inconsistencias y encontrar casos límite. También discutiremos cómo estas técnicas pueden integrarse en lagos de datos y canalizaciones a gran escala para facilitar la depuración de modelos, la optimización de conjuntos de datos y el desarrollo de modelos fundacionales más robustos en visión por computadora. Únete a nosotros para descubrir cómo la búsqueda semántica está transformando la manera en que construimos y refinamos sistemas de inteligencia artificial. Agentes del Mañana: Descifrando los Enigmas de Planificación, UX y Memoria ![](https://cdn.sanity.io/images/h6toihm1/production/ae0baff6b16d895cd5b929c1380a795f9ac5073e-481x481.png?auto=format&dpr=2&fit=max&q=75&w=96) Máximo Fernández Núñez Machine Learning Engineer Bio Los agentes IA, impulsados por LLMs, prometen transformar aplicaciones. Pero, ¿son hoy simples ejecutores o futuros colaboradores inteligentes? Para alcanzar su verdadero potencial, debemos superar barreras críticas. Esta charla se adentra en los 3 enigmas que definirán la próxima generación de agentes: 1\. Planificación Avanzada (El Cerebro): Los agentes actuales a menudo tropiezan con tareas complejas. Exploraremos cómo, más allá de las llamadas a funciones básicas, las arquitecturas cognitivas permiten trazar planes robustos, anticipar problemas y razonar con profundidad. ¿Cómo hacerlos "pensar" varios pasos adelante? 2: UX Revolucionaria (El Alma): La interacción con un agente no puede ser una fuente de frustración. Analizaremos cómo trascender el chat tradicional hacia interfaces "human-on-the-loop", UX colaborativas, generativas y accesibles. ¿Cómo diseñar experiencias que enganchen? 3\. Memoria Persistente (El Legado): Un agente que olvida lo aprendido está condenado a la ineficiencia. Veremos técnicas para dotarlos de memoria significativa que vaya más allá del historial, permitiendo que aprendan y cada interacción sea más inteligente. Llevaremos ideas concretas y una visión clara para contribuir a construir los agentes del mañana: más inteligentes, más intuitivos y verdaderamente capaces. ¿Te unes a la expedición para descifrar el siguiente capítulo de los agentes IA? Desplegando modelos de vision eficicentes: tecnicas de cuantizacion y optimizacion ![](https://cdn.sanity.io/images/h6toihm1/production/33fc40d663f663f4e54b2e2b1a16580956d7124b-481x481.png?auto=format&dpr=2&fit=max&q=75&w=96) Hind Azegrouz Intel Bio Abordamos técnicas para acelerar la inferencia en modelos de visión por computadora mediante optimización y cuantización. Se analizarán estrategias como la reducción de precisión, fusión de operaciones, poda y uso de toolkits como OpenVINO. Se presentarán benchmarks que demuestran mejoras en latencia y rendimiento, manteniendo una precisión aceptable, especialmente en despliegues edge y en tiempo real. ## Sponsors [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-126-lllmstxt|> ## CVPR 2025 Opening Remarks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Opening Remarks from CVPR 2025 Jun 19, 2025 • 5 min read Article content In this article [CVPR 2025 Insights #1: Setting the Stage for Innovation](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#ccd49cb85d74) [Welcome to CVPR 2025](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#336388f338ad) [Raising the Bar: Mandatory Author Reviewing](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#2f2f11e1ad7f) [Quality Over Quantity: Incentives and Enforcement](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#a9519d31604d) [Desk Rejections and Policy Enforcement](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#3656463cf550) [Celebrating Excellence: CVPR 2025 Paper Awards](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#8372c9b149a5) [🎨 Art, Legacy, and Recognition](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#82f543d9631c) [🤝 Community Collaboration](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#ce910e865f9e) [What is next?](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#0b0e9c52ad7f) In this article [CVPR 2025 Insights #1: Setting the Stage for Innovation](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#ccd49cb85d74) [Welcome to CVPR 2025](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#336388f338ad) [Raising the Bar: Mandatory Author Reviewing](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#2f2f11e1ad7f) [Quality Over Quantity: Incentives and Enforcement](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#a9519d31604d) [Desk Rejections and Policy Enforcement](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#3656463cf550) [Celebrating Excellence: CVPR 2025 Paper Awards](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#8372c9b149a5) [🎨 Art, Legacy, and Recognition](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#82f543d9631c) [🤝 Community Collaboration](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#ce910e865f9e) [What is next?](https://voxel51.com/blog/opening-remarks-from-cvpr-2025#0b0e9c52ad7f) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## _**CVPR 2025 Insights \#1**: Setting the Stage for Innovation_ #### [_Paula Ramos, PhD_](https://www.linkedin.com/in/paula-ramos-phd/) The Computer Vision and Pattern Recognition (CVPR) conference continues to evolve, and 2025 marked a pivotal moment in shaping the future of peer review and scholarly contribution in our field. During the opening session, the organizers introduced reforms to improve review quality, community engagement, and fairness. Here’s a summary of what’s new and noteworthy. ![](https://cdn.sanity.io/images/h6toihm1/production/2808c8143cc24e0c027d2ee574418361d3c66781-1290x1345.webp?auto=format&dpr=2&fit=max&q=75&w=1290) ## Welcome to CVPR 2025 With over **13,008 valid submissions** — a **13% increase** from last year — CVPR continues to scale, drawing talent from across the globe. Held primarily in person with virtual components for accessibility, the event features keynotes, tutorials, posters, oral presentations, demos, workshops, and a curated AI Art Gallery. ![](https://cdn.sanity.io/images/h6toihm1/production/fbe766a39e864c9c97b881a2a67d17c33d9daa17-1400x1867.webp?auto=format&dpr=2&fit=max&q=75&w=1400) From a reviewer pool of **12,593 experts**, each paper received at least **three independent reviews**. Ultimately, **2,878 papers were accepted**, resulting in a **22.1% acceptance rate**, reflecting the community’s high bar for excellence. ## Raising the Bar: Mandatory Author Reviewing The conference instituted mandatory author reviewing for the first time in CVPR history. This policy was introduced to ensure that every contributor also participates in reviewing, fostering a culture of responsibility and reciprocity. ### Key Criteria: Authors had to either: - Submit two papers, - Submit a first-author paper, or - Have had at least one paper accepted in a previous top-tier ML conference. - Eligibility extended to current PhD students and those already earning their PhDs. To support a smooth transition, CVPR allowed authors to **opt out** via a simple form or email to the program chairs; no justification was required. The initiative helped reclaim valuable reviewing time and led to a more balanced workload distribution. ## Quality Over Quantity: Incentives and Enforcement To further boost reviewing standards, CVPR 2025 introduced a dual approach combining incentives and compliance: - **Incentives**: Reviewers of nominated or awarded papers were highlighted, and the percentage of reviewers receiving “Outstanding Reviewer” recognition increased from **2% to 6%**. - **Compliance Measures**: Reviewers who failed to submit or delivered poor-quality reviews risked having their submissions rejected by the desk. This accountability measure was applied judiciously, with final decisions vetted by both area and program chairs. As a result, “below expectations” reviews dropped from **9% to 6%**, and PhD students stood out as the **highest quality reviewers**, outperforming their academic and industry peers. ![](https://cdn.sanity.io/images/h6toihm1/production/685cff5274b943421a460a26b470c71f79548494-1400x1867.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ## Desk Rejections and Policy Enforcement Transparency was a recurring theme. Organizers shared that **over 200 papers** were desk rejected this year, for reasons including: **Irresponsible reviewing behavior** (19 cases) **Incomplete author profiles** (18 cases) **Policy violations** such as: - Duplicate submissions, - Use of generative AI without proper verification, - Reference fabrication, - Anonymity breaches, - Exceeding page limits. Notably, a new policy capping authors to **25 submissions** also saw full compliance. Despite some grumbling over filling out OpenReview metadata (e.g., country and institution), organizers stressed its importance for maintaining **geographic diversity and reviewer fairness**. ![](https://cdn.sanity.io/images/h6toihm1/production/4a14304cab13e003982bee18b1d93f477b1e0ef1-1400x1867.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ## Celebrating Excellence: CVPR 2025 Paper Awards A curated list of groundbreaking contributions emerged from rigorous peer review. The award segment of the opening session honored exceptional research across multiple categories: ### 🥇 Best Paper - [**VGGT: Visual Geometry Grounded Transformer**](https://cvpr.thecvf.com/virtual/2025/oral/35294) — \[ _Authors: Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, David Novotny_\] ![](https://cdn.sanity.io/images/h6toihm1/production/7b32b27b022efe235147c547175a57bbcc9e0731-1400x1867.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ### 🌟 Best Student Paper - [**Neural Inverse Rendering from Propagating Light**](https://cvpr.thecvf.com/virtual/2025/oral/35315) — \[ _Authors: Anagh Malik, Benjamin Attal, Andrew Xie, Matthew O’Toole, David B. Lindell_\] ![](https://cdn.sanity.io/images/h6toihm1/production/07d9d06df8be6ab4dec555fb1795fce8d31b85ed-1400x1867.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ### 🏅 Honorable Mentions - [**MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videos**](https://cvpr.thecvf.com/virtual/2025/oral/35311) — \[ _Authors: Zhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang, Linyi Jin, Vickie Ye, Angjoo Kanazawa, Aleksander Holynski, Noah Snavely_\] - [**Navigation World Models**](https://cvpr.thecvf.com/virtual/2025/oral/35338) — \[ _Authors: Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, Yann LeCun_\] - [**Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models**](https://cvpr.thecvf.com/virtual/2025/oral/35281) — \[ _Authors: Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, YenSung Chen, Ajay Patel, Mark Yatskar, Chris Callison-Burch, Andrew Head, Rose Hendrix, Favyen Bastani, Eli VanderBilt, Nathan Lambert, Yvonne Chou, Arnavi Chheda, Jenna Sparks, Sam Skjonsberg, Michael Schmitz, Aaron Sarnat, Byron Bischoff, Pete Walsh, Chris Newell, Piper Wolters, Tanmay Gupta, Kuo-Hao Zeng, Jon Borchardt, Dirk Groeneveld, Crystal Nam, Sophie Lebrecht, Caitlin Wittlif, Carissa Schoenick, Oscar Michel, Ranjay Krishna, Luca Weihs, Noah A. Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, Aniruddha Kembhavi_\] - [**3D Student Splatting and Scooping**](https://cvpr.thecvf.com/virtual/2025/oral/35367) — \[ _Authors: Jialin Zhu, Jiangbei Yue, Feixiang He, He Wang_\] ### 🧑‍🎓 **Best Student Paper Honorable Mention** - [**Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens**](https://cvpr.thecvf.com/virtual/2025/oral/35376) — \[ _Authors: Kaihang Pan, Wang Lin, Zhongqi Yue, Tenglong Ao, Liyu Jia, Wei Zhao, Juncheng Li, Siliang Tang, Hanwang Zhang_\] ## 🎨 Art, Legacy, and Recognition The CVPR Art Gallery Awards highlighted creative brilliance: - **“Green Diffusion” by Masaru Mizuochi**: A poetic parallel between natural decomposition and AI diffusion models, emphasizing the dual forces of creation and destruction through microbial decay and generative noise processes. - **“Learning to Move, Learning to Play, Learning to Animate” by Mingyong Cheng, Sophia Sun, and Han Zhang**: A multimedia performance integrating custom-built robots, real-time AI, motion tracking, and biofeedback-driven sound to explore movement and animation as an embodied experience. - **“Atlas of Perception” by Tom White**: A sculptural exploration of how neural networks understand visual information, revealing the underlying “visual grammar” in the latent space of machine perception. The Longuet-Higgins Prize, recognizing influential papers from CVPR 2015, was awarded to works that stood the test of time, reminding us that the impact of research often blossoms years after publication. ## 🤝 Community Collaboration In closing, organizers emphasized that these changes were made in **collaboration with other top-tier conferences**, sharing policies and lessons to build a stronger reviewing ecosystem across the AI and CV communities. CVPR’s leadership is not just about innovation in research — it’s about cultivating a sustainable, responsible scholarly environment. > “We hope that these continuous efforts to improve reviewing quality will benefit not only the CVPR community but also our sister conferences and the broader field.” ## What is next? If you’re interested in following along as I dive deeper into the world of AI and continue to grow professionally, feel free to connect or follow me on [LinkedIn](https://www.linkedin.com/in/paula-ramos-phd/). Let’s inspire each other to embrace change and reach new heights! You can find me at some [Voxel51 events](https://voxel51.com/events), or if you want to join this fantastic team, it’s worth taking a look at [this page](https://voxel51.com/careers). [CVPR](https://voxel51.com/blog/tag/cvpr) ![](https://cdn.sanity.io/images/h6toihm1/production/e926c07c7d1426c0fde8fdefa637c528d47b16f4-512x512.webp?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a558b86370f2f17212fb2f2c894d590101458a85-5760x3241.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Multimodal Frontier in Computer Vision, Medicine, and Agriculture— CVPR 2025 Reflections\\ \\ Event Recaps, Industry Solutions\\ \\ • \\ \\ Jun 24, 2025](https://voxel51.com/blog/the-multimodal-frontier-in-computer-vision-medicine-and-agriculture-cvpr-2025-reflections) [![](https://cdn.sanity.io/images/h6toihm1/production/97cf3d887735ab9574b2b3e2d3825146016d875e-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Embodied Computer Vision at CVPR 2025: The Next AI Frontier\\ \\ Event Recaps\\ \\ • \\ \\ Jun 30, 2025](https://voxel51.com/blog/embodied-computer-vision-at-cvpr-2025-the-next-ai-frontier) [![](https://cdn.sanity.io/images/h6toihm1/production/e647f4490ed6ad75d63c9f28b67105aebc819e37-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Motion Prompting: Generalized Motion Control for Video Generation\\ \\ Event Recaps\\ \\ • \\ \\ Jun 26, 2025](https://voxel51.com/blog/motion-prompting-generalized-motion-control-for-video-generation) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-127-lllmstxt|> ## Raytheon Technologies Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/d286fcf14376e5da670d2fee8f7d5ab8b89fe09e-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Raytheon Technologies Raytheon Technologies Research Center relies on FiftyOne to visualize large computer vision datasets Apr 27, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) The [Raytheon Technologies Research Center](https://www.rtx.com/who-we-are/what-we-do/transformative-technologies/rtrc) serves as the innovation hub for RTX and its businesses. RTRC’s engineers, scientists and researchers anticipate the discoveries destined to change everything, and they transform that research into the solutions and products that help the company’s businesses shape the future. > "We use FiftyOne to organize large research datasets. My favorite feature is the ability to view distributions over image attributes in the dataset, and filter the dataset by those attributes." – Brett Israelsen, Principal Research Scientist at RTRC ![](https://cdn.sanity.io/images/h6toihm1/production/08e8ddf32e58475f5df1ae64f94921b0f24fb3f2-1048x727.png?auto=format&dpr=2&fit=max&q=75&w=1048) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-128-lllmstxt|> ## Data Blind Spots Webinar [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/fef3495c2bcb5a92c5df355e49f9bc24b98f5d5d-2880x1620.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/fef3495c2bcb5a92c5df355e49f9bc24b98f5d5d-2880x1620.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event Virtual Webinars & Workshops Autonomous Vehicles Exposing Your Data's Blind Spots: Scenario Mining for Safer AV Aug 27, 2025 9 AM Pacific Online. Register for the Zoom! About this event Stress test AV models by mining real-world edge cases using multimodal data and data-centric tools like FiftyOne. Host ![](https://cdn.sanity.io/images/h6toihm1/production/c98c65f91ce1141210966b7a7b090b446e51c7bb-512x512.jpg?auto=format&dpr=2&fit=max&q=75&w=96) Nick Lotz Technical Marketing Engineer Bio Autonomous-vehicle perception stacks routinely miss “unknown unknowns” that sit outside their operational-design domain. This webinar shows how to mine those long-tail, safety-critical scenarios hidden inside massive amounts of raw sensor data. We’ll walk through real examples using datasets like Waymo Open, NuScenes, BDD, and KITTI, and help you build data pipelines for training and validating vision models that actually work in the field. ### **What You Will Learn:** - Why traditional scenario coverage metrics fall short for AV testing - How to use tools like FiftyOne to visualize data gaps and extract failure-prone samples - How to mine safety-critical scenarios from large-scale, multimodal datasets - How to turn natural language descriptions into reusable test cases - How these workflows reduce development cycles and improve AV system robustness ### **Who This Is For:** - AV perception, planning, and test engineers working on real-world validation - Simulation teams looking to scale up coverage - Machine learning engineers and data scientists focused on model robustness - Safety and QA engineers interested in scenario-based testing frameworks - Anyone exploring data-centric approaches to autonomous vehicle development [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-129-lllmstxt|> ## Vivint Customer Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/a988d7ae16303abd2bfd52fd474f020a02aa235e-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Vivint Vivint relies on FiftyOne Teams for intelligent data insights and smarter home security Apr 24, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Vivint](https://www.vivint.com/), a leading smart home company in the United States, delivers an integrated smart home system with in-home consultation, professional installation, and support delivered by Smart Home Pros, as well as 24/7 customer care and monitoring. Dedicated to redefining the home experience with intelligent products and services, Vivint serves more than 1.9 million customers across the US. > "We use [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) to organize, select, display, and share our data which has led to better collaboration with and understanding of our large volume of data. FiftyOne Teams enables us to gain insights such as identifying and understanding data problems early, hypothesis validation, and dataset management overall. This has led to better solution engineering and better testing for the products and services we deliver to our customers." – Lanny Lin, Sr. Director of AI and Data Science at Vivint ![](https://cdn.sanity.io/images/h6toihm1/production/ad1f72c7ea2bb3cebd04179f2ff29ce559ff49d4-996x666.png?auto=format&dpr=2&fit=max&q=75&w=996) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-130-lllmstxt|> ## Medical AI Training Workshop [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/a661cc499d5017c91a6c7e1924d78c8a6aeea29c-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/a661cc499d5017c91a6c7e1924d78c8a6aeea29c-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) In-person Americas Webinars & Workshops Workshop: Train a Medical AI Model in One Day - July 25, 2025 Jul 25, 2025 9:00 AM – 1:00 PM (followed by light lunch and networking) Qualcomm Building 9940 Barnes Canyon Rd San Diego, CA About this event Presented by [EyePop.ai](http://eyepop.ai/) and Voxel51, hosted at Qualcomm. Join us for a hands-on workshop designed to showcase how engineers can build, train, and deploy a computer vision model—fast. Whether you’re a machine learning novice or computer vision veteran, this event is your backstage pass into real-world AI development for medical applications. Host ![](https://cdn.sanity.io/images/h6toihm1/production/c103d784a0a5ae05e6ae21c27b2435d1efe08708-512x512.png?auto=format&dpr=2&fit=max&q=75&w=96) Paula Ramos, PhD Voxel51 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/1fd1c9012dd477f625102d2a3bc5d77cf2b35b7f-512x512.png?auto=format&dpr=2&fit=max&q=75&w=96) Andy Ballester Eyepop.ai Bio ![](https://cdn.sanity.io/images/h6toihm1/production/a7c570082a05e70aa0e3a409e94e4430dffdf753-512x512.png?auto=format&dpr=2&fit=max&q=75&w=96) Blythe Towal, PhD Eyepop.ai Bio **What to Expect** This session brings together two powerful platforms — [EyePop.ai](http://eyepop.ai/) and [FiftyOne](https://docs.voxel51.com/) \- for a one-of-a-kind workflow demo. You’ll leave with a trained model in your account, new tools under your belt, and a clear understanding of how vision AI can be applied to real medical datasets like such as ARCADE, for Stenosis detection **You’ll Learn How To** - Inspect and explore a real-world medical dataset with FiftyOne - Push data to EyePop.ai’s Self-Service Training system. - Label, train, and evaluate your own AI model—on the spot - Auto-label new data using your model - Deploy to the cloud or edge (Snapdragon-compatible!) - Compare quantized vs. original models with FiftyOne **What’s Provided** - A pre-curated medical dataset (limited to 5,000 images per user) - Notebooks and starter scripts (via USB or cloud download) - Pre-provisioned EyePop.ai accounts with API keys - Live support from EyePop.ai and Voxel51 engineers - Optional Qualcomm devices for edge deployment testing **Workshop Highlights** This is a real-time AI training lab with hands-on support. You’ll walk out with your own model, fully trained and ready to deploy. Reserve your spot now and get your API key ahead of time. - Dataset visualization using FiftyOne - Step-by-step model training walkthrough on EyePop.ai - Augmentation quiz and final label review - Real-time model performance evaluation - Deployment and inference testing - Social time, networking, and bonus product demos **Who Should Attend** Engineers, developers, data scientists, product managers, and AI-curious builders—especially those working in or exploring the medical, life sciences, or healthcare space. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-131-lllmstxt|> ## Women in AI Event [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/7a2015ac7fa2e7fc9be86672715c936918544b69-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/7a2015ac7fa2e7fc9be86672715c936918544b69-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Meetups Women in AI - July 24 This event has ended, but you can still catch up! Watch the on-demand recordings and register for our [future events.](https://voxel51.com/events) Jul 24, 2025 9 - 11 AM Pacific Online. Register for the Zoom! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/c55d1fc77c1f39a062dd524f274adf6d7f10f87b-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Shreya Sharma Meta Reality Labs Bio ![](https://cdn.sanity.io/images/h6toihm1/production/bae67b55648f931259e0f1ba534f55e20163e326-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Helena Klosterman Intel Bio ![](https://cdn.sanity.io/images/h6toihm1/production/de68835bd1518bfbb91b4916c97d7e7b5c6c22ba-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Milica Cvetkovic AI @ Google Bio ![](https://cdn.sanity.io/images/h6toihm1/production/947fa55b584178fd8c1fbe04e8be59257f9d50ff-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos Voxel51 Bio About this event Hear talks from experts on cutting-edge topics in AI, ML, and computer vision on July 24. Schedule Exploring Vision-Language-Action (VLA) Models: From LLMs to Embodied AI ![](https://cdn.sanity.io/images/h6toihm1/production/c55d1fc77c1f39a062dd524f274adf6d7f10f87b-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Shreya Sharma Meta Reality Labs Bio This talk will explore the evolution of foundation models, highlighting the shift from large language models (LLMs) to vision-language models (VLMs), and now to vision-language-action (VLA) models. We'll dive into the emerging field of robot instruction following—what it means, and how recent research is shaping its future. I will present insights from my 2024 work on natural language-based robot instruction following and connect it to more recent advancements driving progress in this domain. Multi-modal AI in Medical Edge and Client Device Computing ![](https://cdn.sanity.io/images/h6toihm1/production/bae67b55648f931259e0f1ba534f55e20163e326-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Helena Klosterman Intel Bio In this live demo, we explore the transformative potential of multi-modal AI in medical edge and client device computing, focusing on real-time inference on a local AI PC. Attendees will witness how users can upload medical images, such as X-Rays, and ask questions about the images to the AI model. Inference is executed locally on Intel's integrated GPU and NPU using OpenVINO, enabling developers without deep AI experience to create generative AI applications. Business of AI ![](https://cdn.sanity.io/images/h6toihm1/production/de68835bd1518bfbb91b4916c97d7e7b5c6c22ba-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Milica Cvetkovic AI @ Google Bio The talk will focus on the importance of clearly defining a specific problem and a use case, how to quantify the potential benefits of an AI solution in terms of measurable outcomes, evaluating technical feasibility in terms of technical challenges and limitations of implementing an AI solution, and envisioning the future of enterprise AI. Farming with CLIP: Foundation Models for Biodiversity and Agriculture ![](https://cdn.sanity.io/images/h6toihm1/production/947fa55b584178fd8c1fbe04e8be59257f9d50ff-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Paula Ramos Voxel51 Bio Using open-source tools, we will explore the power and limitations of foundation models in agriculture and biodiversity applications. Leveraging the BIOTROVE dataset. The largest publicly accessible biodiversity dataset curated from iNaturalist, we will showcase real-world use cases powered by vision-language models trained on 40 million captioned images. We focus on understanding zero-shot capabilities, taxonomy-aware evaluation, and data-centric curation workflows. We will demonstrate how to visualize, filter, evaluate, and augment data at scale. This session includes practical walkthroughs on embedding visualization with CLIP, dataset slicing by taxonomic hierarchy, identification of model failure modes, and building fine-tuned pest and crop monitoring models. Attendees will gain insights into how to apply multi-modal foundation models for critical challenges in agriculture, like ecosystem monitoring in farming. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-132-lllmstxt|> ## Sales and Support [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # Get started with FiftyOne Chat with our team about your ML projects. We'll share practical insights and tips to help streamline your workflows. **Flexible hosting and scalability** Work with enterprise-level workloads spanning billions of samples across multimodal data in the cloud, on-premise, or air gapped. **Extensible** Extend capabilities and build custom experiences and front-ends. **Secure and resilient** SSO, role-based access controls, plus audit logging to meet organizational requirements [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-133-lllmstxt|> ## AGI Insights and Discussions [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) AGI [![](https://cdn.sanity.io/images/h6toihm1/production/3f54c19e72060ff2fa99673840a472f4d6e85c8e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Van der Maaten’s Three-System Roadmap to AGI Is Brilliantly Pragmatic\\ \\ Event Recaps\\ \\ • \\ \\ Jun 17, 2025](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic) ## Enough data wrangling.
 Request a demo. [Get started](https://voxel51.com/link-catcher) [Explore the Demo](https://voxel51.com/link-catcher) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-134-lllmstxt|> ## Keypoint Detection Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) Comprehensive Guide to Keypoint Detection for Object Recognition May 5, 2025 • 9 min read Article content In this article [Keypoints Explained](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#9076194fe8d6) [Keypoint Detection Techniques](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#f52ebcd934ba) [Putting Keypoint Detection to Use](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#24ca6ba09fdb) [Keypoint Detection Workflows with FiftyOne](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#ac6f421e044e) [Keypoint Detection Boosts Robotic Grasping: A Real-World Case Study](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#3bc5575949c3) [Future of Keypoint Detection in Object Recognition](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#4aad295d1bf2) [Conclusion](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#e73a9e9e58bb) In this article [Keypoints Explained](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#9076194fe8d6) [Keypoint Detection Techniques](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#f52ebcd934ba) [Putting Keypoint Detection to Use](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#24ca6ba09fdb) [Keypoint Detection Workflows with FiftyOne](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#ac6f421e044e) [Keypoint Detection Boosts Robotic Grasping: A Real-World Case Study](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#3bc5575949c3) [Future of Keypoint Detection in Object Recognition](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#4aad295d1bf2) [Conclusion](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition#e73a9e9e58bb) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/d21be652eacc5b9dd5545861dd41dae1ba22d822-2376x1758.png?auto=format&dpr=2&fit=max&q=75&w=1600) Object recognition sits at the heart of modern AI, powering everything from unlocking your smartphone with a glance to helping self-driving cars navigate safely through traffic. It’s what enables machines to see, understand, and interact with the world around them making it one of the most critical tasks in today’s AI landscape. But as powerful as it is, traditional object recognition methods come with their own set of challenges. Traditional AI methods usually rely on **bounding boxes**, the little rectangles that identify an important object. They are good at finding an object, but less adept at understanding what is really happening inside the same box. That’s where keypoint detection comes in. **Keypoint detection** is like a skeleton that understands and pinpoints precise locations on objects to understand shape, orientation and specific details. ## **Keypoints Explained** Keypoints act as landmarks of an object's distinctive features, like eyes and nose on a face. They are precise, identifiable points that AI uses to create an accurate map of an object. In short, keypoints give AI a concise set of dots to understand and track, making detection tasks easier and reliable. Computer vision and deep learning models learn to recognize and precisely pinpoint the object keypoints, even when objects are twisted, turned or even partially hidden, which bounding boxes fail to achieve. Bounding boxes provide simpler metadata, but they also simplify reality into basic shapes. Keypoints on other hand adapt to an object's flexibility. They capture detailed information on an object's exact shape and orientation. As an example, the following Python code loads an image, runs a pre-trained SuperPoint model to find 2-D keypoints and their confidence scores, then overlays those points on the image. ```python 1from transformers import AutoImageProcessor, SuperPointForKeypointDetection 2import torch 3import matplotlib.pyplot as plt 4from PIL import Image 5import requests 6 7image_path = "~/myimage.png” # Set the image path here 8image = Image.open(image_path) 9 10# Initialize the model and processor 11processor = AutoImageProcessor.from_pretrained("magic-leap-community/superpoint") 12model = SuperPointForKeypointDetection.from_pretrained("magic-leap-community/superpoint") 13 14inputs = processor(image, return_tensors="pt").to(model.device, model.dtype) 15outputs = model(**inputs) 16 17# Postprocess - change model_outputs to outputs here 18image_sizes = [(image.size[1], image.size[0])] 19outputs = processor.post_process_keypoint_detection(outputs, image_sizes) 20keypoints = outputs[0]["keypoints"].detach().numpy() 21scores = outputs[0]["scores"].detach().numpy() 22image_width, image_height = image.size 23 24plt.figure(figsize=(10, 10)) 25plt.axis('off') 26plt.imshow(image) 27plt.scatter( 28 keypoints[:, 0], 29 keypoints[:, 1], 30 s=scores * 100, 31 c='cyan', 32 alpha=0.4 33) 34plt.show() ``` ## **Keypoint Detection Techniques** Keypoint detection is critical to computer vision because it extracts a sparse set of highly repeatable anchor points that can be tracked, matched, or triangulated across frames. Modern systems achieve this by taking three complementary approaches. ### **Heatmap Regression** In **heatmap regression**, a convolutional network outputs a probability map for every key-point class. Each map is a grayscale image whose brighter pixels indicate a higher likelihood that the true key-point is located there. During training, the target map is a 2-D Gaussian “bump” centered on the ground-truth coordinate. The network learns to convert a single point into a smooth probability surface that can later be collapsed back to sub-pixel accuracy. [Image Source](https://user-images.githubusercontent.com/15215732/70846954-6b72d380-1e91-11ea-9499-0f66d9c0ae9c.png) ### **Pose Estimation** **Pose estimation** extends keypoint detection to full skeletons. The model finds joints such as elbows and knees and infers how they connect, recovering the subject’s spatial pose. Pose estimation is vital for augmented reality (AR) filters, motion capture, and robotics. By chaining joints into kinematic graphs, the model tracks complex movements frame-by-frame with high temporal consistency. ### **Part Detection** Part detection decomposes an object into semantically meaningful components (e.g., wheel, door, handle). This approach allows downstream models to reason about each part’s geometry rather than treating the object as a single blob. After it localizes sub-regions, the neural network can predict additional landmarks, boost overall key-point coverage, reduce ambiguity when objects overlap. Using separate validation datasets to tune thresholds for each module further improves precision and ensures the three techniques generalize reliably in production. [Image Source](https://blog.tensorflow.org/2019/11/updated-bodypix-2.html) ## Putting Keypoint Detection to Use Let's now examine why these techniques matter in production systems. Keypoints provide geometric priors that models use to reason about human posture, object orientation, and fine-grained shapes. This information unlocks a spectrum of real-world capabilities that range from safer industrial automation to immersive consumer experiences. ### **Action Recognition** Skeletal keypoints help models to classify complex body movements in real time. Interactive gaming platforms, such as Kinect-style consoles. map a player’s joints to recognize dance steps, yoga poses, and other gestures. The same pipelines extend to surveillance and workplace safety, pose dynamics can flag falls, aggressive behavior, or incorrect ergonomic form. ### **Object Pose Estimation** As mentioned previously, pose estimation matches predicted keypoints to their 2D image projections to recover an object’s full pose in 3D space. Accurate orientation estimates are essential for robotic grasping, bin-picking, automated inspection, and augmented reality. In AR especially, even a few degrees of error can cause a gripper miss or visual drift. ### **Facial Landmark Detection** High-resolution landmark models like [HRNet](https://huggingface.co/docs/timm/en/models/hrnet) locate dozens of reference points along the eyes, nose, mouth, and jawline. These landmarks drive autofocus and exposure control in smartphone cameras, power biometric verification systems, and support driver-attention and fatigue-monitoring solutions in automotive safety. They're responsible for anchoring virtual sunglasses and other fun facial overlays in social apps. ## **Keypoint Detection Workflows with FiftyOne** [FiftyOne](https://voxel51.com/fiftyone) is a computer-vision platform developed by [Voxel51](https://voxel51.com/). It includes tooling for each stage of a keypoint detection pipeline, from dataset exploration and annotation review to view-based error analysis and production monitoring. The result is faster iteration, better model generalization, and simpler hand-offs between data engineers, labelers, and ML engineers. ### **Simplifying Model Development and Deployment** - **Visual dataset exploration**: Filter, search, and sort millions of images or video frames to audit class balance or locate samples where a model confused an elbow for a knee. - **Integrated annotation management**: Connect to popular labeling tools like CVAT and maintain version-controlled keypoint annotations without manual file juggling. - **Smooth production transition**: The same view layouts and evaluation panels used during R&D remain available when validating models on live data, and ensure a consistent feedback loop after deployment. ### **Keypoint-Specific Visualization Features** - **Customizable skeletons** that connect detected joints for cleaner pose inspection. - **Frame-by-frame tracking** of keypoints in video to verify temporal stability. - **Instant anomaly surfacing**: color-code low-confidence points or missing landmarks to spot label noise and model failure modes. For example, the expression below surfaces every sample in which the predicted left\_eye keypoint has confidence < 0.7: ```python 1dataset.filter_labels( 2 "predictions", 3 F("keypoints.detections.points.left_eye.confidence") < 0.7 4) 5 ``` ### **Facilitating Training, Testing, and Error Analysis** - **Balanced dataset curation**: Quantify under-represented poses or body parts, then sample additional images to prevent bias. - **Error diagnostics**: Measure Euclidean distance, PCK, or OKS for each landmark and visualize per-part histograms to trace systematic drift. - **Granular evaluation slicing**: Compare overall mAP to performance on hard subsets (e.g., occluded faces, motion blur) to better target new data acquisition. For example, the code snippet below computes class-specific mAP while restricting the metric to three facial landmarks: ```python 1results = dataset.evaluate_detections( 2 "predictions", 3 gt_field="ground_truth", 4 eval_key="eval", 5 compute_mAP=True, 6 classes=["person"], 7 keypoint_types=["nose", "left_eye", "right_eye"], 8) 9print(results.mAP()) 10 ``` ## Keypoint Detection Boosts Robotic Grasping: A Real-World Case Study In pharmaceutical distribution, item-picking robots must handle blister packs, pill bottles, cartons, and tubes that arrive in every orientation. McKesson’s first generation of KNAPP **Pick-it-Easy Robots** relied on bounding-box detection. The system could find an object in the tote but could not judge the angle of a bottle cap or the position of a tiny blister-pack tab. Misaligned grasps led to slips and re-picks, interrupting the high-throughput flow the warehouse needed. To eliminate those blind spots, McKesson, KNAPP, and Covariant [retrained the vision stack](https://covariant.ai/insights/mckesson-delivers-for-customers-with-the-knapp-pick-it-easy-robot-powered-by-the-covariant-brain) around keypoint detection. A CenterNet-style network now predicts a sparse set of landmarks, giving the motion planner an exact pose for each SKU. In effect, every item carries its own grasp “road-map,” allowing the robot to choose both _where_ and _how_ to grip instead of merely _where_ it is. The gripper’s path is further refined in real time with depth data, so even objects wedged at odd angles in a cluttered bin are approached along a collision-free route. Internal KPIs collected after the upgrade show a **first-attempt pick success rate above 90 percent for single items and roughly 85 percent in cluttered totes**, cutting repeated attempts almost in half. Because fewer picks are repeated, the cell maintains continuous flow. McKesson reports round-the-clock operation without extra staffing, and similar Covariant installations demonstrate throughputs of **up to 515 picks per hour with under 0.1 percent human intervention.** The vision makeover also future-proofs the line: when new medications arrive, the model adapts with a short fine-tuning cycle rather than weeks of rule writing. ## Future of Keypoint Detection in Object Recognition Keypoint detection is moving well beyond classical CNN pipelines. Recent research pairs landmark extraction with [**vision transformers (ViTs)**](https://en.wikipedia.org/wiki/Vision_transformer) that model long-range relationships across an image, allowing the network to reason jointly about spatial context and fine-grained pose. These transformer-keypoint hybrids achieve higher accuracy on crowded-scene benchmarks while maintaining real-time speed when distilled or quantized. Another direction is **edge deployment**. Hardware-efficient backbones now let factories, traffic cameras, and consumer IoT devices run full landmark models locally. Processing on the device trims latency, safeguards privacy, and reduces cloud bandwidth. These benefits are driving rapid adoption in retail analytics, in-cab driver monitoring, and smart-city sensing. Keypoints are also becoming a bridge to **3-D scene understanding**. By predicting landmark coordinates in world space, networks can reconstruct an object’s geometry, estimate scale, and recover complete poses. These capabilities underpin robotic bin-picking, AR object insertion, and digital-twin pipelines, where depth-aware perception is mandatory. Finally, the field is tackling the annotation bottleneck with **self-supervised and weakly supervised learning**. Techniques such as contrastive pre-text tasks and equivariance constraints let models discover stable landmarks without exhaustive human labels. This not only slashes labeling cost but often yields more transferable representations for downstream tasks. Taken together, these advances position keypoint detection as a core primitive for the next generation of vision systems. ## Conclusion Keypoint detection has matured from a research benchmark into a foundational building block for modern computer-vision systems. By localizing precise anatomical or structural landmarks, it supplies the geometric cues required for downstream tasks such as human-pose analysis, 3D object pose, fine-grained face alignment, and 3-D reconstruction. These capabilities now support practical deployments in areas as diverse as surgical navigation, robotic picking, driver-monitoring, motion-capture sports analytics, and AR content creation. Yet the technique remains data-hungry. Engineers must curate balanced datasets, visualize dense landmark annotations, and verify performance on edge cases before moving to production. FiftyOne addresses these pain points directly. Its unified interface lets teams import multimodal datasets, overlay keypoint skeletons on images or video, filter for low-confidence landmarks, and compute metrics such as PCK or OKS at scale. Integrated connectors to common labeling tools and export pipelines shorten iteration loops, lowering both cost and time to deployment. In short, keypoint detection reshapes how machines infer structure and intent, and platforms like FiftyOne make that power accessible to practitioners at any stage. As annotation workflows, self-supervised pretraining, and edge-optimized models continue to advance, the barrier to building landmark-aware applications will fall even further, opening new opportunities across medicine, manufacturing, security, and interactive media. [object recognition](https://voxel51.com/blog/tag/object-recognition) [Keypoints](https://voxel51.com/blog/tag/keypoints) [object detection](https://voxel51.com/blog/tag/object-detection) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/50ab5d62a585ad15e0d6c26c224c48e40f345266-2258x1264.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Implementing Mask R-CNN: Advanced Object Detection and Segmentation\\ \\ Learn\\ \\ • \\ \\ Apr 3, 2025](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation) [![](https://cdn.sanity.io/images/h6toihm1/production/53c3b307e202d573a6fccf19494ba35c43a3dc7e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Best Practices for Evaluating AI Models Accurately\\ \\ Learn\\ \\ • \\ \\ Dec 17, 2024](https://voxel51.com/blog/best-practices-for-evaluating-ai-models-accurately) [![](https://cdn.sanity.io/images/h6toihm1/production/b83b51550d5f3f3fc2f98dbacb2996313896e147-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Quality Dataset Annotation Is Key to Machine Learning\\ \\ Learn\\ \\ • \\ \\ Feb 17, 2025](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-135-lllmstxt|> ## Learn AI Techniques [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Learn [![](https://cdn.sanity.io/images/h6toihm1/production/c69415df557facd00411250d071e976a2bbff233-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ A Comprehensive Guide to Working with Point Cloud Data\\ \\ Learn\\ \\ • \\ \\ Jul 25, 2025](https://voxel51.com/blog/comprehensive-guide-point-cloud-data) [![](https://cdn.sanity.io/images/h6toihm1/production/c845648a6e64e00a3cac3d875280cb462c63633d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Complete Guide to Auto Labeling\\ \\ Learn\\ \\ • \\ \\ Jun 16, 2025](https://voxel51.com/blog/the-complete-guide-to-auto-labeling) [![](https://cdn.sanity.io/images/h6toihm1/production/45b6b66f2f7c3ba6e83db74c270be16d8087ee94-2344x1306.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Comprehensive Guide to Keypoint Detection for Object Recognition\\ \\ Learn\\ \\ • \\ \\ May 5, 2025](https://voxel51.com/blog/comprehensive-guide-to-keypoint-detection-for-object-recognition) [![](https://cdn.sanity.io/images/h6toihm1/production/6c1d7db05a5408e0e53644b3cdb35fc5d46c8389-2340x1302.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How Automated Data Labeling Enhances Computer Vision Efficiency and Accuracy\\ \\ Learn\\ \\ • \\ \\ Apr 17, 2025](https://voxel51.com/blog/how-automated-data-labeling-enhances-computer-vision-efficiency-and-accuracy) [![](https://cdn.sanity.io/images/h6toihm1/production/e0c50d51f1558ec65daf139f10ce4bb099e088ce-2338x1294.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ AI for Predictive Maintenance Using Computer Vision\\ \\ Learn\\ \\ • \\ \\ Apr 17, 2025](https://voxel51.com/blog/ai-for-predictive-maintenance-using-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/3b444e623b2da05079528b845ab98e384b62e003-2330x1306.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ AI Data Modeling for Visual AI: Key Metrics to Build Precise Models\\ \\ Learn\\ \\ • \\ \\ Apr 17, 2025](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models) [![](https://cdn.sanity.io/images/h6toihm1/production/7537eb48892f3c5652be5a82144dfac76ed9a18d-2334x1302.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Image Preprocessing Best Practices To Optimize Your AI Workflows\\ \\ Learn\\ \\ • \\ \\ Apr 17, 2025](https://voxel51.com/blog/image-preprocessing-best-practices-to-optimize-your-ai-workflows) [![](https://cdn.sanity.io/images/h6toihm1/production/04eff439f21847f1068ea3f38b72f8bc7ec77f0e-2340x1308.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Image Similarity Search: Unlocking Pattern Detection in Visual Data\\ \\ Learn\\ \\ • \\ \\ Apr 16, 2025](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data) [![](https://cdn.sanity.io/images/h6toihm1/production/50ab5d62a585ad15e0d6c26c224c48e40f345266-2258x1264.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Implementing Mask R-CNN: Advanced Object Detection and Segmentation\\ \\ Learn\\ \\ • \\ \\ Apr 3, 2025](https://voxel51.com/blog/implementing-mask-r-cnn-advanced-object-detection-and-segmentation) [![](https://cdn.sanity.io/images/h6toihm1/production/e54b1e5afaaede40db246a7681a611856ebd8782-2260x1268.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Enhancing YOLOv8 Segmentation: Precision, Efficiency, and Robustness\\ \\ Learn\\ \\ • \\ \\ Apr 3, 2025](https://voxel51.com/blog/enhancing-yolov8-segmentation-precision-efficiency-and-robustness) [![](https://cdn.sanity.io/images/h6toihm1/production/ca4f84addae3f8e97daad00cf856754301cbb5e0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Are Image Segmentation Maps Superior to Bounding Boxes?\\ \\ Learn\\ \\ • \\ \\ Feb 26, 2025](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes) [![](https://cdn.sanity.io/images/h6toihm1/production/b83b51550d5f3f3fc2f98dbacb2996313896e147-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Quality Dataset Annotation Is Key to Machine Learning\\ \\ Learn\\ \\ • \\ \\ Feb 17, 2025](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning) Load more ## Enough data wrangling.
 Request a demo. [Get started](https://voxel51.com/link-catcher) [Explore the Demo](https://voxel51.com/link-catcher) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-136-lllmstxt|> ## Join the AI Community [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Build better AI, together Step into a world where developers and scientists aren’t just dreaming about AI’s future — they’re building it. Join one of the largest computer vision communities to share, learn, and turn bold ideas into reality. [Join upcoming events](https://voxel51.com/events) [Find us on Discord](https://discord.com/invite/fiftyone-community) ![](https://cdn.sanity.io/images/h6toihm1/production/3f0f981ccf2245904b0ef1f2a9ca6b8fe202648e-1673x1255.png?auto=format&dpr=2&fit=max&q=75&w=837) ![](https://cdn.sanity.io/images/h6toihm1/production/3374802b5468ab950ff74be40448380407d9343b-699x729.png?auto=format&dpr=2&fit=max&q=75&w=350) ![](https://cdn.sanity.io/images/h6toihm1/production/96713e7e2ba3c752720a85fdd3846b277e76e541-768x594.jpg?auto=format&dpr=2&fit=max&q=75&w=384) ![](https://cdn.sanity.io/images/h6toihm1/production/666dfdd1a78cc373ac7741f18256fc5980d971df-967x639.png?auto=format&dpr=2&fit=max&q=75&w=484) Connect ## Join the FiftyOne Community [Join us on Discord](https://community.voxel51.com/) [Check out FiftyOne on GitHub](https://github.com/voxel51/fiftyone) [Follow us on LinkedIn](https://www.linkedin.com/company/voxel51/) [Follow us on X](https://x.com/voxel51) [Find us on Hugging Face](https://huggingface.co/Voxel51) ## Become a member of the AI, ML, and Data Science Meetup Join more than 20,000 developers and scientists around the world who are leveling up their AI skills in
GenAI, Computer Vision, LLMs, MLOps, vector search, and more. [Join upcoming events](https://voxel51.com/events) ![](https://cdn.sanity.io/images/h6toihm1/production/dea067d4cded2a986cc0f4173ec818c64e0d57bc-925x444.svg) Learn ## Free visual AI courses Unlock the power of visual AI with free, hands-on courses on top learning platforms. Gain practical skills in data handling and model evaluation to enhance your AI/ML projects. ### Data-Centric Visual AI on Coursera This comprehensive course is your hands-on guide to developing and maintaining high-quality datasets for visual AI applications. Enroll for free! [Learn more & enroll](https://www.coursera.org/learn/hands-on-data-centric-visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/23ce154d99ae6c8266d222c6e0fd3bc3da1c2e28-1261x600.png?auto=format&dpr=2&fit=max&q=75&w=631) ### Data-Centric Visual AI on LinkedIn Learning This course provides a concise yet comprehensive hands-on experience in implementing data curation methodologies, focusing data feedback loops. [Go to course](https://www.linkedin.com/learning/data-centric-visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/249890d9580fc06338cdc9435c793478e81cc59e-1261x600.png?auto=format&dpr=2&fit=max&q=75&w=631) ### Practical Computer Vision with PyTorch Practical Computer Vision in PyTorch is a comprehensive, hands-on course designed for developers and practitioners eager to explore computer vision using PyTorch. [Go to openHPI](https://open.hpi.de/courses/computervision2025) ![](https://cdn.sanity.io/images/h6toihm1/production/455e4d60a1089a8333bfe4e500525b8bf9369c7f-1261x601.png?auto=format&dpr=2&fit=max&q=75&w=631) By the numbers ## Come build with us 0M Downloads of FiftyOne 0K+ Stars on GitHub 0K Slack+Discord Members 0K+ Meetup Members > “We use FiftyOne to organize large research datasets. My favorite feature is the ability to view distributions over image attributes in the dataset, and filter the dataset by those attributes.” > > **Brett Israelsen** > > Principal Research Scientist, AI, Raytheon ![](https://cdn.sanity.io/images/h6toihm1/production/293dc34ac3ccdd7973ac04b5a27e66a0c0862260-158x61.svg) > “At Allstate, my team works on auto vehicle damage inspection. Verifying the damage to a vehicle can take an insurance claim agent hours to verify, but using computer vision and FiftyOne, we can segment the parts of vehicles first, then detect the damages, and finally match the damage to repair costs and generate reports for the adjusters.” > > **Pavan Nanjundappa** > > Data Science Manager, Allstate India ![](https://cdn.sanity.io/images/h6toihm1/production/876897e6347fb3fae892d3c91ff18031dddb2aef-600x132.svg) > “Smart Eye is the global leader in Human Insight AI, technology that understands, supports and predicts human behavior in complex environments. FiftyOne is leveraged at Smart Eye to efficiently curate the datasets used to train models for human-centric mobility products like Driver Monitoring Systems and Interior Sensing solutions.” > > **Fredrik Walterson** > > Group Manager/Automotive Solutions, Smart Eye ![](https://cdn.sanity.io/images/h6toihm1/production/c7a3b0ccedb2d5f25f7a26f37c47e61ae4782c2c-485x72.png?auto=format&dpr=2&fit=max&q=75&w=100) > “FiftyOne is the backbone of our Data Engine. It helps us to clean and relabel our datasets efficiently, and convert our model predictions into large scale training datasets with a little help from CVAT.” > > **George Pearse** > > Founding Machine Learning Engineer, BinIt ![](https://cdn.sanity.io/images/h6toihm1/production/efebd8fc0ac3b4713e64581916b8f0496b529ae6-154x51.svg) > “As we developed our Florence-2 model, FiftyOne proved invaluable for data management and visualization. Its powerful capabilities helped streamline our workflow, ensuring we built a robust foundation for our models. Now, as we dive into the development of Florence-5B, we're relying on FiftyOne more than ever. The tool's intuitive interface and rich feature set are essential for effectively managing our large datasets and gaining critical insights. If you're working on AI and data-driven projects, I highly recommend checking out FiftyOne. It's made a significant impact on our work, and I'm confident it can do the same for you!” > > **Bin Xiao** > > AI Researcher, Meta (formerly Principal Research Manager, Microsoft GenAI) ![](https://cdn.sanity.io/images/h6toihm1/production/d9b4bdb75662f7f850d7abbf017918438f9f4885-3333x713.png?auto=format&dpr=2&fit=max&q=75&w=100) > “Updata uses FiftyOne to efficiently sample and analyze our computer vision models. Its intuitive UI helps us show clients how machine learning improves manufacturing and sustainability. FiftyOne's Python SDK and visual dataframe concept have also strengthened our tooling for ML experimentation and is an essential component of our stack.” > > **David Cardozo** > > Chief Lead Analyst, Updata ![](https://cdn.sanity.io/images/h6toihm1/production/619a611c506787b71d9209dc34b48b27346c1159-198x47.svg) Contribute ## Contribute to FiftyOne. Here are a few ways to get involved. #### Log an issue If something doesn’t work as
expected or if you think it could be done better, log an issue on GitHub [Visit the Fiftyone repo](https://github.com/voxel51/fiftyone/issues) #### Submit a PR FiftyOne is an Apache 2.0 licensed project, and we always welcome contributions [Contribute to FiftyOne](https://github.com/voxel51/fiftyone/blob/develop/CONTRIBUTING.md) #### Help others on Discord Answer questions in a variety of use case-specific channels on Discord [Join Discord](https://community.voxel51.com/) #### Write a plugin Writing plugins is a great way to contribute to the project and extend FiftyOne to do novel things [Learn more](https://docs.voxel51.com/plugins/developing_plugins.html) #### Give a talk Consider sharing your expertise, research, or lessons learned at an upcoming Meetup [Submit a talk](https://voxel51.com/events) #### Share your success Is your company using FiftyOn to address interesting computer vision challenges? Share your success! [Share your story](https://voxel51.com/events) ### Subscribe to our newsletter Get bi-weekly updates on the latest AI, machine learning, and computer vision news, events, and resources. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-137-lllmstxt|> ## Women in AI Event [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/c5114c78b43a502f91389b79f56e15f0f09aecd7-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/c5114c78b43a502f91389b79f56e15f0f09aecd7-960x540.png?auto=format&dpr=2&fit=max&q=75&w=420) Register for the event Virtual Americas Meetups Women in AI - October 2, 2025 Oct 2, 2025 9 AM Pacific Online. Register for the Zoom! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/8d92c451b31f2b400f5e2d15eb8397d4b9dae693-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Ria Cheruvu NVIDIA Bio ![](https://cdn.sanity.io/images/h6toihm1/production/d8983f9f8fb1fa8819ab9b8966fb8be988a103bf-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Paula Ramos Voxel51 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/e4d2e786d9e6b4fd0e80605a0b3b529be06f1a1d-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Apoorva Joshi MongoDB Bio ![](https://cdn.sanity.io/images/h6toihm1/production/f2e094ec035ad6bb97297b318630a56ff4c55829-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=42&q=75&w=42) Sheena Yap Chan Author Bio About this event Hear talks from experts on the latest topics in AI, ML, and computer vision on October 2. Schedule The Hidden Order of Intelligent Systems: Complexity, Autonomy, and the Future of AI ![](https://cdn.sanity.io/images/h6toihm1/production/8d92c451b31f2b400f5e2d15eb8397d4b9dae693-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Ria Cheruvu NVIDIA Bio As artificial intelligence systems grow more autonomous and integrated into our world, they also become harder to predict, control, and fully understand. This talk explores how complexity theory can help us make sense of these challenges, by revealing the hidden patterns that drive collective behavior, adaptation, and resilience in intelligent systems. From emergent coordination among autonomous agents to nonlinear feedback in real-world deployments, we’ll explore how order arises from chaos, and what that means for the next generation of AI. Along the way, we’ll draw connections to neuroscience, agentic AI, and distributed systems that offer fresh insights into designing multi-faceted AI systems. Managing Medical Imaging Datasets: From Curation to Evaluation ![](https://cdn.sanity.io/images/h6toihm1/production/d8983f9f8fb1fa8819ab9b8966fb8be988a103bf-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Paula Ramos Voxel51 Bio High-quality data is the cornerstone of effective machine learning in healthcare. This talk presents practical strategies and emerging techniques for managing medical imaging datasets, from synthetic data generation and curation to evaluation and deployment. We’ll begin by highlighting real-world case studies from leading researchers and practitioners who are reshaping medical imaging workflows through data-centric practices. The session will then transition into a hands-on tutorial using FiftyOne, the open-source platform for visual dataset inspection and model evaluation. Attendees will learn how to load, visualize, curate, and evaluate medical datasets across various imaging modalities. Whether you're a researcher, clinician, or ML engineer, this talk will equip you with practical tools and insights to improve dataset quality, model reliability, and clinical impact. Building Agents That Learn: Managing Memory in AI Agents ![](https://cdn.sanity.io/images/h6toihm1/production/e4d2e786d9e6b4fd0e80605a0b3b529be06f1a1d-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Apoorva Joshi MongoDB Bio In the rapidly evolving landscape of agentic systems, memory management has emerged as a key pillar for building intelligent, context-aware AI Agents. Different types of memory, such as short-term and long-term memory, play distinct roles in supporting an agent's functionality. In this talk, we will explore these types of memory, discuss challenges with managing agentic memory, and present practical solutions for building agentic systems that can learn from their past executions and personalize their interactions over time. Human-Centered AI: Soft Skills That Make Visual AI Work in Manufacturing ![](https://cdn.sanity.io/images/h6toihm1/production/f2e094ec035ad6bb97297b318630a56ff4c55829-480x481.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=96&q=75&w=96) Sheena Yap Chan Author Bio Visual AI systems can spot defects and optimize workflows—but it’s people who train, deploy, and trust the results. This session explores the often-overlooked soft skills that make Visual AI implementations successful: communication, cross-functional collaboration, documentation habits, and on-the-floor leadership. Sheena Yap Chan shares practical strategies to reduce resistance to AI tools, improve adoption rates, and build inclusive teams where operators, engineers, and executives align. Attendees will leave with actionable techniques to drive smoother, people-first AI rollouts in manufacturing environments. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-138-lllmstxt|> ## Visual AI in Healthcare [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Visual AI in Healthcare & Medicine Deliver better outcomes by using FiftyOne to build powerful solutions for detection, diagnosis, and care. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/b9b8865e1c2a0c63dc781314b4c0104da2b7ed7b-628x393.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/f13d47a258b93b49db217587a52c68e49f8bccb3-628x817.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/e03f2c081e7ccbc72f42001eb9d9177778f1bad8-628x813.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/25eca6139611206869444b056a242197adbefcf0-628x395.png?auto=format&dpr=2&fit=max&q=75&w=314) Benefits ## Build better medical and healthcare solutions with Voxel51 Data is at the heart of medicine and healthcare. AI builders choose FiftyOne to help them refine visual data and models to build robust, reliable visual AI solutions. ### Improve outcomes Increase robustness and reliability of computer vision solutions by bringing quality data and insights to every step of your visual AI development with FiftyOne. ### Deliver solutions faster Shave months off of development time by using FiftyOne to streamline and automate how you curate data and test your models so that you can iterate quickly and get to production faster. ### Increase productivity Avoid repetitive and time consuming computer vision tasks while streamlining data annotation, data quality, and model evaluation workflows with FiftyOne. Use cases ## Help people live healthier, happier lives Companies in medicine and healthcare rely on
FiftyOne across their visual AI development to boost data quality and model performance. ![](https://cdn.sanity.io/images/h6toihm1/production/ff35fdf01624a3a4a3048a77a2c9b7b94d26644a-721x936.png?auto=format&dpr=2&fit=max&q=75&w=361) Computer-aided detection and diagnosis Assist healthcare providers in understanding and evaluating medical data by identifying suspicious regions in images and alerting physicians to concerning areas. ![](https://cdn.sanity.io/images/h6toihm1/production/eec73aab2bf8b596dc2319429be6c8a579e3d32d-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Monitoring disease progression Precisely monitor disease progression over time. Enable doctors to track disease progression by comparing markers in biomedical images across different points in time. ![](https://cdn.sanity.io/images/h6toihm1/production/7c18e4f3fdd58b3e763215ba7599a05cc28ca007-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Preoperative surgical planning Combine digital templating, visualization, and simulation with CT scans and other diagnostic imaging, to help minimize surgery time and its invasiveness. This reduces costs, time in the operating room, and even mortality rates. ![](https://cdn.sanity.io/images/h6toihm1/production/bc76733eb6590f327edc14a16497829b62f47e07-768x425.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Intraoperative surgical guidance Enable surgeons to use real-time visual and spatial information to navigate with enhanced precision. Fuse preoperative and intraoperative images with AR. ![](https://cdn.sanity.io/images/h6toihm1/production/dd84ee3a054f95b56891999700298b475bbffffc-768x512.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Assisting people with vision loss Object and face recognition, monetary denomination verification, and spatial navigation help people with blindness or vision loss map and navigate their surroundings. ![](https://cdn.sanity.io/images/h6toihm1/production/af9e4709c2db6b0e83576c859c73bd15f8026130-768x410.jpg?auto=format&dpr=2&fit=max&q=75&w=384) Slip and fall detection Recognize, analyze, and predict fall incidents by detecting unusual movements or postures, and alert caregivers to ensure help arrives as soon as possible. Features ## How visual AI can help you [Book a demo](https://voxel51.com/sales) Unify multimodal data Model evaluationDe-identification of patient dataData versioning ### Manage millions of multimodal samples through a unified interface Machine learning engineers spend over half of their time wrangling data, but it doesn’t have to be that way. Use FiftyOne’s powerful dataset import and manipulation capabilities to manage millions of diagnostic images and videos. resources ## Learn more about visual AI in healthcare ### Visual AI in healthcare events [View all](https://voxel51.com/events?industry=healthcare) ![](https://cdn.sanity.io/images/h6toihm1/production/1bd4596e43fc708848051c57f484f07ce5f56ef1-1276x754.png?auto=format&dpr=2&fit=max&q=75&w=640) ### Read our blogs [Read healthcare blogs](https://voxel51.com/blog/tag/healthcare) ![](https://cdn.sanity.io/images/h6toihm1/production/eaa44eeb403349df4caffcf9efb3d110b98a6119-1770x2062.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=350&q=75&w=350) ![](https://cdn.sanity.io/images/h6toihm1/production/5691fe278acf8d79c2bb71f6da17f600de33675d-6048x1921.png?auto=format&dpr=2&fit=max&q=75&rect=2387,0,3661,1921&w=640) ### Trusted by experts [Read customer stories](https://voxel51.com/customers?category=health-and-medicine) ![](https://cdn.sanity.io/images/h6toihm1/production/5691fe278acf8d79c2bb71f6da17f600de33675d-6048x1921.png?auto=format&dpr=2&fit=max&q=75&rect=2387,0,3661,1921&w=640) ![](https://cdn.sanity.io/images/h6toihm1/production/1bd4596e43fc708848051c57f484f07ce5f56ef1-1276x754.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=38&q=75&rect=0,0,64,38&w=38) Customer stories ## How SafelyYou detected 350,000 elderly safety events with visual AI using FiftyOne ### With FiftyOne, SafelyYou: - Reclaimed 80 hours of development time - Reduced 77% of images sent for manual verification - Improved Average Precision scores by 10%. [Read the story](https://voxel51.com/customers/safelyyou) ![](https://cdn.sanity.io/images/h6toihm1/production/1a672e54a7fb869beb104c2f70bbdee3ec5e5ff7-912x913.png?auto=format&dpr=2&fit=max&q=75&rect=0,182,912,535&w=456) > “I use FiftyOne on a daily basis to improve the quality of our data and visually inspect model predictions.” > > **Marijn Lems** > > Machine Learning Engineer, Aidence ![](https://cdn.sanity.io/images/h6toihm1/production/749b3b8589ea02792de3cdecd8aba61846bbd261-1024x270.png?auto=format&dpr=2&fit=max&q=75&w=100) > “We use FiftyOne’s CVAT integration to request and load annotations from doctors, and compare them to the models’ predictions. Using FiftyOne’s embeddings visualization, we get a better understanding of our data distribution and identify the most unique images by eliminating near duplicates.” > > **Oğuz Hanoğlu** > > PhD Candidate/Engineer, Multimedia Informatics ![](https://cdn.sanity.io/images/h6toihm1/production/a2e825b3c0669b8e2064489d0bab0d0d5693b698-1024x348.png?auto=format&dpr=2&fit=max&q=75&rect=0,14,1024,327&w=100) Resources ## Developer resources ### Datasets - [Try on FiftyOne: Chest X-ray samples](https://try.fiftyone.ai/datasets/chestx-ray14/samples) - [Kaggle: Diabetic retinopathy detection](https://try.fiftyone.ai/cas/auth/signin?error=InvalidSessionError) - [NVIDIA: VISTA-3D and MedSAM-2](https://dev.to/voxel51/visual-ai-in-healthcare-nvidias-vista-3d-and-medsam-2-medical-imaging-models-18j3) - [Leading models & datasets](https://medium.com/@paularamos_phd/visual-ai-in-healthcare-2025-landscape-8d25f22d43dc) ### Models - [Google MedGemma](https://github.com/harpreetsahota204/medgemma) - [NVIDIA TotalSegmentator foundation model for MRI/CT segmentation](https://developer.nvidia.com/blog/visual-foundation-models-for-medical-image-analysis/) - MedSAM2 Medical segmentation - [Med-sam-2-video-torch](https://docs.voxel51.com/model_zoo/models.html#med-sam-2-video-torch) ### Tutorials - [Evaluating a classifier](https://docs.voxel51.com/tutorials/evaluate_classifications.html) - [Segment CT scans with NVIDIA VISTA-3D](https://voxel51.com/blog/segment-anything-in-a-ct-scan-with-nvidia-vista-3d/) ## Data eats models for lunch Talk to our computer vision experts to start building better datasets and models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-139-lllmstxt|> ## Aisprid Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/6c1873bdb23550a0e00810b5f5246853c39b13db-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Aisprid Aisprid’s high-precision robots for greenhouse farming rely on FiftyOne Apr 13, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) Aisprid is pioneering agricultural robotics to tackle one of humanity’s most pressing challenges: sustainable food production for a growing global population. Their breakthrough technology combines high-precision robotics and artificial intelligence to automate greenhouse and farming tasks, starting with their innovative deleafing robot for tomato plants. Aisprid is helping growers meet increasing food demand while creating more sustainable farming practices for the future. > "Aisprid designs and manufactures AI-driven autonomous and highly accurate robots for fruit and vegetable growers. FiftyOne has been integrated into our MLOps workflow to enable the curation of higher quality data, leading to more accurate models." – Alexandre Cornu, CTO at Aisprid ![](https://cdn.sanity.io/images/h6toihm1/production/0131fba5fbf20202b918a2aac2c9749d3aa2fa3d-2560x1920.jpg?auto=format&dpr=2&fit=max&q=75&w=1600) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-140-lllmstxt|> ## G42 Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/72ec0f31c4005d0d7c86bbecd6519522fec2d88b-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) G42 G42 gets a superior interface for computer vision data with FiftyOne Apr 9, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [G42](https://www.g42.ai/) is a leading AI & Cloud Computing company based in Abu Dhabi, working on projects from molecular medicine to space travel and everything in between. > "FiftyOne provides a superior interface for dealing with computer vision data. Its extensive Python package lets you do almost any data transformation, perform similarity search, and easily evaluate model predictions." – Rustem Galiullin, Data Scientist at G42 ![](https://cdn.sanity.io/images/h6toihm1/production/c97aeada2e993f24c5a07e1d691c65ff4fc1595e-473x313.png?auto=format&dpr=2&fit=max&q=75&w=473) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-141-lllmstxt|> ## Food Waste Hackathon [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/f830109b60f2272c3d450e7f0e814077ef5d1695-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/f830109b60f2272c3d450e7f0e814077ef5d1695-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) In-person EMEA Hackathons Food Waste Estimation Hackathon - Computer Vision for Sustainability - 1 August, 2025 Aug 1, 2025 10:00 AM - 8:00 PM Hasso Plattner Institute Main building H, d-space Prof.-Dr.-Helmert-Str. 2-3, D-14482 Potsdam About this event Join us for a hands-on computer vision hackathon focused on tackling food waste through image analysis. Participants will build systems that estimate food waste by analyzing post-meal images, contributing to sustainability efforts through machine learning. Host ![](https://cdn.sanity.io/images/h6toihm1/production/aa2a4972154ea9e7cfdf522f28d33fbdacfc8e80-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Antonio Rueda-Toicen AI Engineer \| Voxel51 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/dee24044d4aaf635f46b9c6072a422d8b5311be0-400x400.jpg?auto=format&dpr=2&fit=max&q=75&w=96) Felix Boelter AI Engineer \| Hasso Plattner Institute Bio **Target Audience** - Data Scientists - Machine Learning Engineers - Computer vision programmers **Prerequisites:** Participants must understand computer vision fundamentals. Familiarity with the open source dataset curation tool FiftyOne is highly beneficial but not required. We recommend reviewing the documentation of FiftyOne and the following openHPI MOOC link for self-study preparation before the event: - [FiftyOne documentation](https://docs.voxel51.com/) - [Practical Computer Vision with PyTorch](https://open.hpi.de/courses/computervision2025) The particularly relevant sections are neural network fundamentals, image-based regression, image augmentation, object detection, and image segmentation. Tutorial notebooks are available [on GitHub.](https://github.com/andandandand/practical-computer-vision) **Challenge: Food Waste Estimation from Images** Build computer vision systems that analyze food images to estimate waste quantities. Teams will work with post-meal photographs to develop accurate food waste measurement algorithms. **Technical Tasks** **Object Detection and Segmentation** - Zero-shot food item segmentation with YOLO-E and YOLO-W and Grounding DINO + SAM **Dataset Creation and Management** - Image dataset exploration and quality assessment with FiftyOne - Dataset curation using CVAT annotation tools for food images - Custom segmentation dataset creation with food waste labels - Validation set construction for before/after images - Data augmentation techniques for diverse food scenarios **Food Waste Estimation Modeling** - Volume estimation algorithms (primary focus): Calculate remaining food quantities from segmented regions - Waste percentage calculation: Compare before/after states to determine consumption ratios - Multi-food item tracking: Handle plates with multiple food types - Portion size standardization: Account for varying plate sizes **Advanced Features** - Synthetic food image generation for training enhancement - Fine-tuning models for specific food categories **Rewards** - 1st Place: Amazon gift card with a value of 300 EUR - 2nd Place: Amazon gift card with a value of 150 EUR - 3rd Place: Amazon gift card with a value of 50 EUR **Scoring Rubric** Participants earn points based on implementation quality and food waste estimation accuracy: **Technical Implementation (50 points)** - Food Detection/Segmentation (25 points): Accurate identification and boundary detection of food items and use of model ensembles and manual annotation - Waste Estimation Accuracy and Experiment Tracking (25 points): Precision in calculating food waste percentages and volumes, and use of appropriate metrics. **Best Practices Adherence (30 points)** - Dataset curation (30 points): Enrichment of the dataset through labeling, identification and resolution of leaky duplicates across train and validation splits, identification of uniqueness and representativeness of samples, richness of embeddings visualizations, data augmentation strategies, sharing of curated version of the dataset on HuggingFace - Documentation and reproducibility (5 points): quality of public GitHub repository, inclusion of requirements.txt file and / or Docker image with reproducible solution, use of wandb to track experiments **Creative solutions and deployment (15 points)** - Creative Solutions (10 points): Novel approaches to food volume estimation (e.g. augmentation of the dataset with diffusion models). - Deployment (5 points): creation of a deployed version of the app in a Gradio endpoint on HuggingFace **Resources Provided** **Development Platform** - Google Colab: Pre-configured notebooks with quickstart tutorials **Technical Support** - Mentors with expertise in computer vision and food analysis - Access to pre-trained food detection models and baseline implementations - Documentation on food waste estimation techniques **Schedule** - 10:00 AM - Registration and team formation - 10:30 AM - Technical briefing: Food waste estimation challenges - 11:00 AM - Hacking begins - 1:00 PM - Lunch break (observe your own food waste!) - 2:00 PM - Mid-event check-in and mentor consultations - 6:00 PM - Project submissions deadline - 6:30 PM - Team presentations: Demo your food waste estimation system - 7:30 PM - Judging and awards ceremony - 8:00 PM - Event conclusion **Impact** Your work will contribute to understanding and reducing food waste through technology. The best solutions could help restaurants, cafeterias, and households track and minimize their environmental impact. **What to Bring** - Laptop with VSCode installed - Basic knowledge of Python and machine learning frameworks - Interest in sustainability and environmental impact Join us to combine computer vision expertise with environmental consciousness in the heart of Potsdam's tech ecosystem! **Data Protection Notice** Please be aware that we will take photos and videos during the event. If you wish to opt out of this, please let us know in advance. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-142-lllmstxt|> ## Understanding Visual Agents Event [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/1b5ee3eebfb4e43dcbb8b4a0eb78d19855cb00fd-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/1b5ee3eebfb4e43dcbb8b4a0eb78d19855cb00fd-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Meetups Understanding Visual Agents - August 7, 2025 This event has ended, but you can still catch up! Watch the on-demand recordings and register for our [future events.](https://voxel51.com/events) Aug 7, 2025 9 AM Pacific Online. Register for the Zoom! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/659546ed3327c4051837b12078f1a14c1614b5a5-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Raghav Kapoor Adobe Bio ![](https://cdn.sanity.io/images/h6toihm1/production/bf06e48a1c8ef96cb107e0b195c925db60a097a2-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Yixiao Song University of Massachusetts Amherst Bio ![](https://cdn.sanity.io/images/h6toihm1/production/0346a48c2d013123dcfe14f72b76e07ecc073609-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Rasul Osmanbayli Kapital Bank, Baku/Azerbaijan Bio ![](https://cdn.sanity.io/images/h6toihm1/production/1e275aeef4b3b4f0276ec615da45a62abb04e6e5-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Sahota Voxel51 Bio About this event Join the Meetup to hear talks from experts on understanding visual agents. Schedule Foundational capabilities and models for generalist agents for computers ![](https://cdn.sanity.io/images/h6toihm1/production/659546ed3327c4051837b12078f1a14c1614b5a5-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Raghav Kapoor Adobe Bio As we move toward a future where language agents can operate software, browse the web, and automate tasks across digital environments, a pressing challenge emerges: how do we build foundational models that can act as generalist agents for computers? In this talk, we explore the design of such agents—ones that combine vision, language, and action to understand complex interfaces and carry out user-intent accurately. We present OmniACT as a case study, a benchmark that grounds this vision by pairing natural language prompts with UI screenshots and executable scripts for both desktop and web environments. Through OmniACT, we examine the performance of today’s top language and multimodal models, highlight the limitations in current agent behavior, and discuss research directions needed to close the gap toward truly capable, general-purpose digital agents. BEARCUBS: Evaluating Web Agents' Real-World Information-Seeking Abilities ![](https://cdn.sanity.io/images/h6toihm1/production/bf06e48a1c8ef96cb107e0b195c925db60a097a2-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Yixiao Song University of Massachusetts Amherst Bio The talk focuses on the challenges of evaluating AI agents in dynamic web settings, the design and implementation of the BEARCUBS benchmark, and insights gained from human and agent performance comparisons. In the talk, we will discuss the significant performance gap between human users and current state-of-the-art agents, highlighting areas for future improvement in AI web navigation and information retrieval capabilities. Implementing a Practical Vision-Based Android AI Agent ![](https://cdn.sanity.io/images/h6toihm1/production/0346a48c2d013123dcfe14f72b76e07ecc073609-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Rasul Osmanbayli Kapital Bank, Baku/Azerbaijan Bio In this talk, I will share with you practical details of designing and implementing Android AI agents, using deki ( [GitHub - RasulOs/deki: ML model (or several models) to describe the contents of the UI screen](http://github.com/RasulOs/deki)) From theory, we will move to practice and the usage of these agents in industry/production. For end users - remote usage of Android phones or for automation of standard tasks. Such as: - _"Write my friend 'some\_name' in WhatsApp that I'll be 15 minutes late"_ - _"Open Twitter in the browser and write a post about 'something'"_ - _"Read my latest notifications and say if there are any important ones"_ - _"Write a linkedin post about 'something'"_ And for professionals - to enable agentic testing, a new type of test that only became possible because of the popularization of LLMs and AI agents that use them as a reasoning core. Visual Agents: What it takes to build an agent that can navigate GUIs like humans ![](https://cdn.sanity.io/images/h6toihm1/production/1e275aeef4b3b4f0276ec615da45a62abb04e6e5-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Harpreet Sahota Voxel51 Bio We’ll examine conceptual frameworks, potential applications, and future directions of technologies that can “see” and “act” with increasing independence. The discussion will touch on both current limitations and promising horizons in this evolving field. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-143-lllmstxt|> ## Protex AI Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/6034d0ca986e0d226bb5732aecbda2a8a98e6637-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Protex AI How Protex AI built a flexible, scalable ML pipeline for workplace safety with FiftyOne Jul 15, 2025 Protex AI streamlined its computer vision pipeline with FiftyOne, accelerating model iteration and improving safety intelligence across 100+ industrial sites. Sped up model development time by 5x Sites actively processed through FiftyOne 100+ Custom FiftyOne plugins integrated into the pipeline 10 Article content In this article [Introduction: Protex AI](https://voxel51.com/customers/protex-ai#82f179d2c67a) [Challenge: Workflow inefficiencies slowed ML development](https://voxel51.com/customers/protex-ai#77fffc11a88b) [Solution: Operationalizing visual data with FiftyOne](https://voxel51.com/customers/protex-ai#bf44059b0b3e) [Results: 5x faster model iteration enabled by flexible workflows](https://voxel51.com/customers/protex-ai#878c61949b6c) In this article [Introduction: Protex AI](https://voxel51.com/customers/protex-ai#82f179d2c67a) [Challenge: Workflow inefficiencies slowed ML development](https://voxel51.com/customers/protex-ai#77fffc11a88b) [Solution: Operationalizing visual data with FiftyOne](https://voxel51.com/customers/protex-ai#bf44059b0b3e) [Results: 5x faster model iteration enabled by flexible workflows](https://voxel51.com/customers/protex-ai#878c61949b6c) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) ![](https://cdn.sanity.io/images/h6toihm1/production/546439d520ac93f454810b04d85a936465548466-852x480.gif?auto=format&dpr=2&fit=max&q=75&w=852) ### **_Key Results_** - _5x faster model iteration by consolidating fragmented tools into a single visual pipeline_ - _Data from 100+ sites and 1000+ CCTV cameras processed through FiftyOne for data validation, curation, training, and model analysis_ - _~10 customized workflows built using FiftyOne’s plugin framework_ - _Improved collaboration through dataset tagging, URL sharing, and centralized data access_ - _Deeper model insights through side-by-side performance and data comparisons_ > “What really stands out about FiftyOne is the flexibility. The plugin framework lets us customize our workflows based on our unique needs, and the mature SDK lets us consolidate more of our pipeline into one tool, avoiding the cost of stitching together multiple systems. FiftyOne integrates directly into our production pipeline in a reliable, scalable way, which is critical for how we operate.” _— Patrick Rowsome, Head of Computer Vision Operations, Protex AI_ ## **Introduction: Protex AI** [**Protex AI**](https://www.protex.ai/) is a leading provider of AI-powered workplace safety and operations solutions that help organizations build safer, smarter industrial workplaces using real-time computer vision. By leveraging existing CCTV infrastructure, Protex AI enables safety and operations teams to monitor industrial workplace environments and take proactive action. Their technology has driven a significant impact, including an 80% reduction in incidents at major retail, logistics, and manufacturing organizations such as Amazon, DHL, General Motors, and Sysco. ## **Challenge: Workflow inefficiencies slowed ML development** Protex AI’s computer vision team processes video data from more than 1,000 cameras across 100+ customer sites. Their production computer vision models perform tasks such as object detection, classification, and pose estimation to power real-time safety and operational insights. Using Protex AI, EHS and operations teams proactively prevent accidents and enforce compliance across industrial environments. These safety and operational use cases include: - Speeding vehicles in facilities - Improper ergonomics (e.g., unsafe posture and actions such as manual lifting) - Missing protective equipment (e.g., helmets, gloves) - Unauthorized access to restricted zones - Site hazards like spills and clutter that could lead to accidents Before adopting FiftyOne, the team relied heavily on manual scripts to train models and iterate on experiments. Over time, their home-grown processes became unwieldy, and the need for a central solution for data and model work became evident. Without a centralized system in place, sharing data, debugging model outputs, and maintaining version control added complexity to every iteration. > _“It was very clumsy to use a script to generate a video with predictions on it, and then share it around with the team. The lack of a central system for data and models not only slowed model development but also added operational overhead.” — Patrick Rowsome, Head of Computer Vision Operations, Protex AI_ ## **Solution: Operationalizing visual data with FiftyOne** Protex AI slowly transitioned from a fragmented, script-heavy workflow to a cohesive visual data engine that powers their daily ML operations using FiftyOne. The ML engineers initially experimented with open source FiftyOne and soon realized the power of the product. As their scalability and collaboration needs increased, the Enterprise version of FiftyOne was a big draw for the team, and FiftyOne slowly became the central hub for Protex AI’s ML operations ### **How Protex AI uses FiftyOne** The team uses FiftyOne to ingest CCTV footage from the client's warehouses and sites. - **Ingest and inspect:** Ingests raw CCTV footage for visual validation, checking for data issues like corruption or network glitches before it enters the pipeline. - **Curate for training:** They curate large video datasets down to high-quality subsets optimized for model training. - **Analyze and refine:** FiftyOne is then used to analyze model performance, identify strengths and weaknesses, and guide further data and model refinement. Protex AI’s ML team also built ~10 internal plugins using FiftyOne’s Python SDK and flexible plugin framework for their niche operations for data filtering, annotation handoff, and running inference jobs, further tightening the loop between data and model refinement. ## **Results: 5x faster model iteration enabled by flexible workflows** Since integrating FiftyOne, Protex AI has seen measurable improvements in their computer vision workflow: **5x speedup in model iteration**, driven by streamlining data curation, centralized collaboration, and the ability to quickly identify model strengths and weaknesses. - **Improved collaboration** across teams with shared tagging, URL sharing, and centralized access to datasets and annotations - **Intuitive model analysis:** The team found the tool’s ability to view model performance alongside the underlying data extremely helpful. As early testers of this functionality, they found the side-by-side comparison of multiple model versions to be an essential tool for debugging and optimization. It helped the team drill into true positives, false positives, and other performance metrics while staying visually grounded in the actual data. - **Custom plugin framework** supporting ~10 internal customized workflows that accelerate development and eliminate repetitive manual tasks > “What really stands out about FiftyOne is the flexibility. The plugin framework lets us customize our workflows based on our unique needs, and the mature SDK lets us consolidate more of our pipeline into one tool, avoiding the cost of stitching together multiple systems. FiftyOne integrates directly into our production pipeline in a reliable, scalable way, which is critical for how we operate.” _— Patrick Rowsome, Head of Computer Vision Operations, Protex AI_ By integrating FiftyOne into their ML pipeline, Protex AI has accelerated model development, improved debugging efficiency, and delivered safer, more reliable solutions, driving 80 %+ reductions in workplace incidents for their clients. ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-144-lllmstxt|> ## Qinecsa Pharmacovigilance Solutions [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/3191b9ce305cb63311a09968bd394af91f5e837a-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Qinecsa Qinecsa’s pharmacovigilance solutions rely on FiftyOne Apr 6, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Qinecsa](https://qinecsa.com/) is a leading provider of technology-led, end-to-end pharmacovigilance solutions and is a trusted partner to global life science companies. Qinesca brings together best-in-class technology and scientific expertise to connect life science companies to the right safety solutions. > "Qinesca’s Pv-trace project assists with object tracking inside warehouses. FiftyOne is used to create annotations of boxes and build models for edge devices." – Bhageeratha Tanedar, IT Analyst at Qinecsa ![](https://cdn.sanity.io/images/h6toihm1/production/e15d1985aa225fc7d157d7ebb678c5a2b5703425-960x640.png?auto=format&dpr=2&fit=max&q=75&w=960) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-145-lllmstxt|> ## Image Similarity Search [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) Image Similarity Search: Unlocking Pattern Detection in Visual Data Apr 16, 2025 • 10 min read Article content In this article [Importance of Image Similarity Search](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#44f05175051d) [Linking to Practical Applications](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#33a6ca64b3bd) [Benefits of Using Image Similarity Tools](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#293e1d8d25b1) [Natural Language Queries with FiftyOne Brain](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#83b8c771b6dc) [What Is Image Similarity Search?](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#1dc27da2fb87) [Algorithms Used in Image Similarity Search](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#df6e88602e28) [How Image Similarity Search Works](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#e6490623c9b6) [Use Cases for Image Similarity Search](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#e209674d878d) [Challenges and Considerations](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#bf36f32d2b58) [Why Image Similarity Search Matters](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#b0b6aef7d887) [Try It Yourself: Hands-On Similarity Search](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#f5ac1f783ad7) In this article [Importance of Image Similarity Search](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#44f05175051d) [Linking to Practical Applications](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#33a6ca64b3bd) [Benefits of Using Image Similarity Tools](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#293e1d8d25b1) [Natural Language Queries with FiftyOne Brain](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#83b8c771b6dc) [What Is Image Similarity Search?](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#1dc27da2fb87) [Algorithms Used in Image Similarity Search](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#df6e88602e28) [How Image Similarity Search Works](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#e6490623c9b6) [Use Cases for Image Similarity Search](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#e209674d878d) [Challenges and Considerations](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#bf36f32d2b58) [Why Image Similarity Search Matters](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#b0b6aef7d887) [Try It Yourself: Hands-On Similarity Search](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data#f5ac1f783ad7) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/0752fc1024cc346d5f4a4ca9b421040e8f512945-700x700.gif?auto=format&dpr=2&fit=max&q=75&w=700) Modern industries increasingly depend on [object detection](https://en.wikipedia.org/wiki/Object_detection) within visual data, from e-commerce product searches and manufacturing anomaly detection to advanced security systems and medical imaging analysis. As the volume of visual data grows exponentially, efficiently identifying similar images based on content rather than traditional metadata becomes crucial. Image similarity search in combination with [FiftyOne](https://voxel51.com/fiftyone/) addresses this need, significantly enhancing how visual information is analyzed, organized, and utilized across web search engines, specialized data mining platforms, and consumer-facing applications. Image similarity search refers to methods of identifying visually related images within a dataset by comparing their visual content, such as shapes, colors, and textures. It involves converting each image into a numerical representation known as an embedding, typically generated by deep learning models. These embeddings capture visual characteristics of images in a high-dimensional vector space, enabling efficient comparison and retrieval of visually similar items. By mapping images into this embedding space, similar images naturally cluster together, which simplifies searching and analyzing visual data. ![](https://cdn.sanity.io/images/h6toihm1/production/43ae39af3435dce8892cfd3f6aaa588ce0893eb4-1107x875.png?auto=format&dpr=2&fit=max&q=75&w=1107) ## **Importance of Image Similarity Search** By enabling systems to locate similar images based on content, these methods open a wide array of use cases: - [**Anomaly detection**](https://en.wikipedia.org/wiki/Anomaly_detection): Pinpoint subtle production defects by comparing new images to a reference set. - **Object classification**: Recognize objects, shapes, or patterns even under variations in angle or lighting. - **Personalized recommendations**: Suggest visually alike products or designs that align with user preferences. These benefits pave the way for more efficient operations, lower error rates, and faster decision-making across industries. For instance, when applied to e-commerce, an image-based recommender can surface relevant products with minimal user input. ## **Linking to Practical Applications** ![](https://cdn.sanity.io/images/h6toihm1/production/cc716023da38d6c5decee9d90268192ab71dd0f6-760x580.png?auto=format&dpr=2&fit=max&q=75&w=760) When image similarity search aligns with industry-specific goals, it unlocks new possibilities for automation and value creation. This approach is analogous to document similarity techniques used for text. However, in this case, the techniques apply to visual content, including images within documents. Image similarity search can also help maintain large retail catalogs by streamlining product organization and discovery. Tools like FiftyOne help streamline the data structures and [labeling efforts](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox/) needed to build robust query vector indexing within a vector database, letting teams focus on model improvements rather than data headaches. ## **Benefits of Using Image Similarity Tools** FiftyOne actively orchestrates the image similarity workflow. It leverages powerful external libraries (like PyTorch for embeddings) and specialized vector search backends (like FAISS or vector databases) for the core computations. This integration allows FiftyOne to provide crucial tools that: - Streamline labeling, filtering, and comparison workflows. - Integrate results from large-scale nearest neighbor searches. - Offer analytics for quick debugging and iteration on your embeddings or search algorithms. Here’s how you can tag the top-k neighbors directly in FiftyOne, so you can then filter or visualize them in the UI: ```python 1def tag_neighbors_in_fiftyone(dataset, query_sample, neighbor_idxs): 2 # Remove old tags 3 for s in dataset: 4 if "neighbors" in s.tags: 5 s.tags.remove("neighbors") 6 s.save() 7 # Tag query + neighbors 8 query_sample.tags.append("neighbors") 9 query_sample.save() 10 for idx in neighbor_idxs: 11 s = dataset[idx] 12 s.tags.append("neighbors") 13 s.save() 14 ``` ## **Natural Language Queries with FiftyOne Brain** Beyond image-based similarity, FiftyOne’s Brain API integrates powerful multimodal embedding models like CLIP (Contrastive Language-Image Pretraining), enabling users to perform intuitive natural language queries. CLIP generates embeddings for both images and text prompts within the same high-dimensional vector space, allowing users to search datasets simply by entering descriptive text phrases. For instance, queries such as "blue sedan on highway," "dog playing fetch," or "sunset at the beach" will instantly surface visually matching images, making data exploration significantly more accessible. This capability provides considerable advantages for dataset curation, annotation validation, and model debugging. Users can seamlessly execute natural language searches through FiftyOne’s interactive UI or programmatically via its Python SDK, streamlining workflows and greatly reducing the reliance on manual tagging or metadata. As a result, dataset management becomes more intuitive, efficient, and aligned with real-world scenarios. ## **What Is Image Similarity Search?** Image similarity search retrieves images similar to a given query vector (the image’s feature representation) from a collection or index. Unlike text-based methods, it focuses on visual content such as colors, textures, and shapes, using deep learning or engineered features to measure image proximity in a metric space. ### **Defining Image Similarity Search** Fundamentally, similarity search uses a numerical representation (commonly a vector) for each image. You provide a query image (or query vector), and the search system computes distances (e.g., cosine similarity, Euclidean distance) to find images that have minimal distance to the query. Smaller distance equates to higher similarity. ### **Key Functionalities** - **Feature Comparison**: Images are often converted into embeddings via Convolutional Neural Networks (CNNs). These embeddings are then compared for resemblance in a high-dimensional space. - **Content-based Retrieval**: Instead of matching metadata or keywords, the system matches underlying visual patterns, capturing subtle nuances—like the texture of a fabric or the contours of a product. ## **Algorithms Used in Image Similarity Search** **Search Strategy** **Core Concept** **Speed** **Accuracy** **Scalability** **Key Consideration** _Exact Search (e.g., Brute Force)_ Compare query vector to every other vector using a chosen metric. Very Slow Exact (100%) Poor Only feasible for very small datasets. _Tree-based (e.g., KD-Tree, Ball Tree)_ Partition feature space hierarchically to prune search space. Moderate Exact/Approx. _Moderate_ Performance degrades significantly in high dimensions. _Hashing-based (e.g., LSH)_ Hash similar items to the same buckets for faster candidate selection. Fast Approximate Good Accuracy depends heavily on hash function tuning. _Quantization/Graph ANN (e.g., FAISS)_ Cluster/index vectors (often using compressed codes or graph structures). Very Fast Approximate Very High Tunable speed/accuracy trade-off; Index training needed. ## **How Image Similarity Search Works** ### **Feature Extraction** Images are transformed into feature vectors through ML models. For example, a deep CNN might output a 512-dimensional embedding for each image. These embeddings capture vital details like edges, shapes, and textures: - **Deep Learning Models**: ResNet or MobileNet can serve as pretrained backbones, with final dense layers used as the feature representation. - **Hand-Engineered Features**: In some simpler tasks, SIFT or HOG might still suffice, though they usually underperform modern neural nets. ![](https://cdn.sanity.io/images/h6toihm1/production/225f3332e6f70ce3c9de91b0315905183318007f-1464x813.png?auto=format&dpr=2&fit=max&q=75&w=1464) Below is a minimal example of using a pretrained ResNet (minus its classification layer) to produce a 512-dimensional embedding for each image: ```python 1import torch 2import torch.nn as nn 3import torchvision.models as models 4import torchvision.transforms as T 5from PIL import Image 6# 1) Load a pretrained ResNet and remove its final classification layer 7base_model = models.resnet18(weights=models.ResNet18_Weights.IMAGENET1K_V1) 8embedder = nn.Sequential(*list(base_model.children())[:-1]).eval() 9# 2) Define a simple transform (resize, center-crop, normalize) 10transform = T.Compose([\ 11 T.Resize(256),\ 12 T.CenterCrop(224),\ 13 T.ToTensor(),\ 14 T.Normalize([0.485, 0.456, 0.406],\ 15 [0.229, 0.224, 0.225]),\ 16]) 17def get_embedding(img_path): 18 pil_img = Image.open(img_path).convert("RGB") 19 tensor_img = transform(pil_img).unsqueeze(0) 20 with torch.no_grad(): 21 feats = embedder(tensor_img) # [1, 512, 1, 1] 22 return feats.flatten().numpy() # shape: (512,) 23 ``` ### **Similarity Measurement** After feature extraction, embeddings are compared to the query vector using metrics like cosine similarity, which measures orientation-independent resemblance and handles scale variations well, or Euclidean (L2) distance, emphasizing absolute distances especially when embeddings are normalized. The choice of metric depends on domain-specific considerations, with cosine similarity commonly preferred for orientation-invariant matching. Here's a concise function to compute cosine similarity between embeddings: ```python 1import numpy as np 2def cosine_similarity(emb1, emb2): 3 # Normalize both embeddings 4 emb1_norm = emb1 / (np.linalg.norm(emb1) + 1e-8) 5 emb2_norm = emb2 / (np.linalg.norm(emb2) + 1e-8) 6 # Dot product of normalized vectors = cosine similarity 7 return float(np.dot(emb1_norm, emb2_norm)) ``` ### **Result Retrieval** ![](https://cdn.sanity.io/images/h6toihm1/production/d1712ab7fbf9bed46af8cdf825a00397e8d600ce-1176x875.png?auto=format&dpr=2&fit=max&q=75&w=1176) Once distances are computed, the system performs a nearest neighbor search to retrieve the most similar images for the query this can be done via: - **KD-Trees or Ball Trees**: Traditional data structures for lower-dimensional data, though they can degrade in high dimensions. - **Approximate Nearest Neighbor (ANN) Search**: Methods like FAISS (Facebook AI Similarity Search) or Annoy (Approximate Nearest Neighbors Oh Yeah) handle massive datasets by approximating the exact nearest neighbors, enabling sub-linear retrieval times. - **Inverted File Indexing (IVF)**: Combines clustering-based partitioning with local searching in each cluster to optimize large-scale queries. ```python 1import faiss 2# Suppose we have embeddings in a NumPy array: all_embs.shape = (N, 512) 3dimension = 512 4index = faiss.IndexFlatL2(dimension) 5index.add(all_embs) 6 7def find_neighbors_l2(query_emb, k=5): 8 # Convert to float32 and reshape for FAISS 9 q_float = query_emb.astype('float32').reshape(1, -1) 10 distances, indices = index.search(q_float, k) 11 return distances[0], indices[0] 12 ``` ## **Use Cases for Image Similarity Search** ### **Manufacturing** By comparing new product images against a reference library, the system can quickly spot blemishes or anomalies (e.g., scratches on a phone's casing). This streamlines data mining for production anomalies and reduces recall costs. ### **Retail and E-Commerce** If a shopper uploads an image of a designer bag, dress, or shoe, visual search tools can quickly identify and recommend similar products from a store's inventory, matching color, shape, style, or pattern, thus streamlining the shopping experience without relying on descriptive keywords. ### **Medical Image Analysis** Hospitals accumulate immense volumes of X-rays, CT scans, and MRI images. Image similarity frameworks can compare new scans to known examples of ailments (e.g., cancerous tumors) to flag potential diagnoses. This aids radiologists and speeds up the diagnostic pipeline. ### **Face Recognition** Many security systems rely on nearest neighbor search of face embeddings to confirm identities or detect persons of interest. Cosine similarity is often used to measure how close a face embedding is to a known ID. ### **Object Tracking** In video surveillance, tracking a suspect's clothing color or object shape might rely on repeated similarity checks across frames or across multiple camera feeds. ### **E-Commerce & Visual Search** ![](https://cdn.sanity.io/images/h6toihm1/production/3f71a97c89975c682822f3888a403c201253daa7-1270x1143.png?auto=format&dpr=2&fit=max&q=75&w=1270) Consider a consumer who snaps a photo of a shoe they like on the street. They upload it to an online store's app. The platform extracts the query vector from the photo and runs an approximate nearest neighbor search in its vector database, retrieving the top 5–10 matches. The user can then pick from those visually similar models. This approach simplifies user journeys and can boost sales conversion by presenting relevant recommendations quickly. ## **Challenges and Considerations** ### **High Computational Cost** Running vector search on large-scale image catalogs can be resource-intensive: - **GPU Acceleration**: Many systems rely on GPUs to embed images in real time or perform large-scale matrix multiplications for distance computations. - **Cloud Solutions**: Outsourcing storage and on-demand compute can simplify scaling while controlling overhead costs. ### **Handling Large Datasets** As image libraries soar into the millions or billions: - **Approximate Nearest Neighbor Search**: Partitioning or clustering data (e.g., IVF, multi-index hashing) drastically speeds up queries. - **Indexing Structures**: Well-optimized indexes can help avoid naive O(n) searches. For large-scale collections (millions of images), exact searches can be slow. Below is a snippet using an IVF index (inverted file) to speed up queries at the cost of a slight approximation: ```python 1nlist = 100 # number of clusters 2quantizer = faiss.IndexFlatL2(dimension) 3ivf_index = faiss.IndexIVFFlat(quantizer, dimension, nlist, faiss.METRIC_L2) 4# Train on sample embeddings, then add them 5ivf_index.train(all_embs) 6ivf_index.add(all_embs) 7def find_neighbors_ivf(query_emb, k=5): 8 q_float = query_emb.astype('float32').reshape(1, -1) 9 distances, indices = ivf_index.search(q_float, k) 10 return distances[0], indices[0] 11 ``` ### **Variability in Images** Real-world images can vary widely in lighting, angle, resolution, or background: - **Data Augmentation**: Training or fine-tuning embedding models on augmented samples (e.g., rotations, flips, color shifts) makes them more robust. - **Invariant Features**: Metric learning sometimes includes invariance to certain transformations, ensuring retrieval remains accurate. ## **Why Image Similarity Search Matters** Image similarity search is essential for quickly detecting visual patterns in large datasets across industries like manufacturing, healthcare, and e-commerce. Advancements in machine learning, vector databases, approximate nearest neighbor retrieval, and platforms like FiftyOne improve precision, scalability, and automation. Future integration of natural language processing (NLP) and vector search promises richer, multimodal retrieval, bridging text and visuals to drive innovative solutions. ## **Try It Yourself: Hands-On Similarity Search** We’ve provided a companion [Jupyter notebook](https://colab.research.google.com/drive/1d5cdFh3Z8GCMN5PLO3es-E0zxyxGAutW?usp=sharing) demonstrating: 1. **Loading Data**: Loading a sample dataset (COCO subset) using the FiftyOne Dataset Zoo. 2. **Embedding Generation**: Extracting meaningful feature vectors (embeddings) from images using a pre-trained ResNet model via PyTorch/TorchVision. 3. **Indexing**: Building an efficient search index from the computed embeddings using FAISS (\`IndexFlatL2\`). 4. **Similarity Search**: Querying the FAISS index to find the nearest neighbors (most visually similar images) for a selected query image. 5. **Visualization & Integration**: Displaying the query and neighbor images side-by-side using Matplotlib, and tagging these samples within the FiftyOne dataset for interactive exploration in the FiftyOne App. By working through this notebook, you can gain hands-on experience implementing the core components of an image similarity search pipeline, from data loading and embedding to indexing, searching, and integrating results using FiftyOne. **Image Citations** - Lin, Tsung-Yi, et al. _"Microsoft COCO: Common Objects in Context."_ COCO Dataset 2017 Validation Split, cocodataset.org, 2017, [https://cocodataset.org/#home](https://cocodataset.org/#home). Accessed 24 Mar. 2025. [similarity search](https://voxel51.com/blog/tag/similarity-search) [images](https://voxel51.com/blog/tag/images) [Visual AI](https://voxel51.com/blog/tag/visual-ai) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/d2e24d0a14de508f8ccd36eaffd3b909c9f193b9-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Image Embeddings Transform Computer Vision Capabilities\\ \\ Learn\\ \\ • \\ \\ Nov 25, 2024](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities) [![](https://cdn.sanity.io/images/h6toihm1/production/ca4f84addae3f8e97daad00cf856754301cbb5e0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Are Image Segmentation Maps Superior to Bounding Boxes?\\ \\ Learn\\ \\ • \\ \\ Feb 26, 2025](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes) [![](https://cdn.sanity.io/images/h6toihm1/production/9252e8bb5db5c4805f4a6f315b51527ee5c22072-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Guide to AI Image Segmentation\\ \\ Learn\\ \\ • \\ \\ Dec 19, 2024](https://voxel51.com/blog/a-guide-to-ai-image-segmentation) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-146-lllmstxt|> ## Link Catcher Tool [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) about [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-147-lllmstxt|> ## Model Evaluation Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # FiftyOneModel Evaluation From aggregate performance metrics to sample-level diagnostics, FiftyOne allows you to diagnose failure modes and edge cases preventing your models from reaching optimal performance in production. [Book a demo](https://voxel51.com/sales) [View docs](https://docs.voxel51.com/user_guide/evaluation.html) ![](https://cdn.sanity.io/images/h6toihm1/production/3e795ee996356246388c3bad919ae3363db07ec4-2954x1206.png?auto=format&dpr=2&fit=max&q=75&rect=237,0,2522,1206&w=1261) Model Evaluation ## Instantly correlate metrics with data samples Quickly move between aggregate performance metrics and specific data samples to identify exactly what’s driving model failures or successes. [Explore model evaluation](https://voxel51.com/blog/unified-model-insights-with-fiftyone-model-evaluation-workflows) ### Built-in evaluation methods Run analyses on regression, classification, detection, polygon, instance, and semantic segmentation tasks. Access standard aggregate metrics like classification reports, confusion matrices, and precision-recall curves directly within FiftyOne. ![](https://cdn.sanity.io/images/h6toihm1/production/47cc933d695e44db182e05ceb9a9e7015e0bcd16-2552x1508.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,2552,1508&w=640) ### Fine-grained, sample-level insights Go beyond aggregate metrics by capturing detailed statistics like accuracy and false-positive counts at the sample level to reveal annotation errors, training gaps, or data biases. ![](https://cdn.sanity.io/images/h6toihm1/production/b8a6b8077d19b9c26375c5396fc52fb46d5d04aa-2552x1508.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,2552,1508&w=640) Product Demo ## See model evaluation in action Quickly move between aggregate performance metrics and specific data samples to identify exactly what’s driving model failures or successes. ## Model vs. reality Compare model predictions against ground truth labels directly on your images. Quickly identify where your model excels — and exactly where it needs improvement. [Explore model predictions](https://docs.voxel51.com/user_guide/dataset_creation/#model-predictions) ![FiftyOne allows you to visualize model predictions against ground truth labels](https://cdn.sanity.io/images/h6toihm1/production/7fb4c1b0fbb00cc0b0b9183788f6ef2b8d1d86e6-1258x1200.png?auto=format&dpr=2&fit=max&q=75&w=600) ## Compare models with scenario analysis Benchmark models across key metrics and granular data slices to pinpoint performance gaps, identify edge-case failures, and guide targeted model improvements. [Learn more](https://docs.voxel51.com/user_guide/evaluation.html) ![](https://cdn.sanity.io/images/h6toihm1/production/3ca1b330c63e537c1217fd35d2e7cdf338c9637c-1600x1600.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1600,1600&w=600) ## Enough data wrangling.
 Request a demo. [Book a demo](https://voxel51.com/sales) [Explore Verified Auto Labeling](https://voxel51.com/annotation) ![](https://cdn.sanity.io/images/h6toihm1/production/ea42e9b26f49f1cb54bb8aca31dc10e7f74fe11f-3024x960.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-148-lllmstxt|> ## Smart Eye Partnership [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/5b420def37b04af0cc5d0167fe91492f7c7362ac-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Smart Eye Smart Eye relies on FiftyOne to build Human Insight AI Apr 20, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Smart Eye](https://www.smarteye.se/) is the leading provider of Human Insight AI, technology that understands, supports and predicts human behavior in complex environments. In automotive, Smart Eye’s driver monitoring systems and interior sensing solutions improve road safety and the mobility experience. > "Smart Eye is the global leader in Human Insight AI, technology that understands, supports and predicts human behavior in complex environments. FiftyOne is leveraged at Smart Eye to efficiently curate the datasets used to train models for human-centric mobility products like Driver Monitoring Systems and Interior Sensing solutions." – Fredrik Walterson, Group Manager/Automotive Solutions at Smart Eye ![](https://cdn.sanity.io/images/h6toihm1/production/90b2a6ce5ddad126bcfec8b2ce3ace18431b0905-500x400.jpg?auto=format&dpr=2&fit=max&q=75&w=500) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-149-lllmstxt|> ## Verified Auto Labeling [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Verified-Auto-Labeling [![](https://cdn.sanity.io/images/h6toihm1/production/990a31a85d850d41fb482ae29960a8fa10b18ecd-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Behind the math: how we built the annotation savings calculator\\ \\ Product & News\\ \\ • \\ \\ Jun 4, 2025](https://voxel51.com/blog/how-we-built-annotation-savings-estimator) ## Enough data wrangling.
 Request a demo. [Get started](https://voxel51.com/link-catcher) [Explore the Demo](https://voxel51.com/link-catcher) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-150-lllmstxt|> ## SafelyYou Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/1a672e54a7fb869beb104c2f70bbdee3ec5e5ff7-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) SafelyYou How SafelyYou detected 350,000 elderly safety events with visual AI using FiftyOne May 4, 2025 SafelyYou overcame data management bottlenecks by adopting FiftyOne’s privacy-compliant model development workflow, enabling accurate fall detection and effective healthcare management. Reclaimed development time per month 80 hours Reduction in images sent for manual verification 77% Improved Average Precision scores 10%+ Article content In this article [Challenge](https://voxel51.com/customers/safelyyou#df8a23c90d5c) [Solution](https://voxel51.com/customers/safelyyou#4128f658310c) [Key Results](https://voxel51.com/customers/safelyyou#afbdf48fab58) In this article [Challenge](https://voxel51.com/customers/safelyyou#df8a23c90d5c) [Solution](https://voxel51.com/customers/safelyyou#4128f658310c) [Key Results](https://voxel51.com/customers/safelyyou#afbdf48fab58) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [SafelyYou](https://www.safely-you.com/) AI-powered solutions improve elderly patient safety and streamline care delivery in senior care facilities. By leveraging computer vision models and data pipelines that process millions of videos, their technology has enabled clinicians to analyze over 350,000 real events—more than any other company globally—resulting in an 80% reduction in fall-related ER visits and driving both patient safety and cost savings for healthcare facilities. > “The efficiencies we’ve brought in with FiftyOne have saved our team 60-80 hours of development time per month.” – Ryan Szeto, Senior Computer Vision Engineer at SafelyYou ## Challenge ### Data management and scalability limitations slowed AI-driven care innovation SafelyYou’s AI solution provides patient safety and continuous care management by leveraging computer vision models that process real-time data from millions of hours of video footage. Before implementing FiftyOne as part of their ML data management and model development pipeline, SafelyYou’s AI team faced challenges in scaling their AI-driven solutions for fall detection and healthcare monitoring. - **Custom data analysis:** The AI team spent at least half their time maintaining custom-built scripts for reviewing AI predictions, false positives, and negatives for their [Respond](https://www.safely-you.com/safelyyou-respond/) solution, considerably slowing their model development workflow. - **Data management dependencies:** The team had limited control over their data and relied on their internal Case & Notification team (which was responsible for managing the clinical transactional database) to implement queries and visualizations for AI model evaluation. The workflow wasn’t ideal and hindered development and quick iterations. - **Scalability:** Their transactional data pipeline was difficult to scale, which created bottlenecks in processing the growing volume of video data required for training AI models. As the team moved to an event-based architecture to reduce bottlenecks, the AI team needed a data solution that was scalable and flexible, and which allowed them to: - Explore and filter large volumes of video data quickly and share interesting samples without downloading them - Instantly visualize object annotations and identify mistakes - Protect the personally identifiable information (PII) associated with healthcare ## Solution ### Scalable, privacy-compliant model development workflow for accurate fall detection and care management using FiftyOne SafelyYou found **FiftyOne** as the perfect fit to address their data management and scalability challenges for their [Respond](https://www.safely-you.com/safelyyou-respond/) and [Clarity](https://www.safely-you.com/safelyyou-clarity/) solutions. With FiftyOne as the primary resource for ML work, the team was able to: - **Streamline the ML workflow** by visualizing and exploring videos and metadata such as start/end timestamps, sensor IDs, and AI scores in production. They were able to cross-reference videos with their data warehouse to construct datasets for evaluation and explore their data visually and dynamically, without downloading the data to a remote server or incurring dependencies on other teams. - **Easily find annotation mistakes** with FiftyOne field visualizations and tagging to distinguish incorrect labels and predictions - **Fine-tune and validate AI models** by examining model performance to track false positives/negatives and create quality datasets for training and evaluation. **With FiftyOne**, SafelyYou was able **to meet their privacy compliance needs by ensuring** personally identifiable information **(PII) in healthcare data was protected** during the annotation and model training process. ## Key Results ### Streamlined ML workflow reclaimed 60-80 hours of development time per month; 77% reduction in images that require manual verification FiftyOne has helped SafelyYou process over 60 million minutes of video footage, totalling 3TB of data. Since implementing the tool in their ML pipeline, the team saw the impact on their workflow almost immediately. - The team was able to identify AI model failures and annotation mistakes easily. Through the process, they achieved a **77% reduction in images sent for manual verification**. - Efficiency in workflows **saved** **60 – 80 hours of development time per month**. - Model **Average Precision scores went up by 10%**, improving fall detection accuracy in production. These results have delivered significant business benefits to SafelyYou. Their technology has reduced falls by 40% and improved patient safety. It has also helped extend residents’ stays by an average of 4+ months, enabling care facilities to retain long-term residents and provide consistent care. > “FiftyOne is our primary resource for machine learning research. Thanks to FiftyOne's convenient field visualizations and filtering capabilities, we can easily distinguish incorrect labels and predictions, and therefore iterate on models faster than ever. As a result, we've achieved a 77% reduction in images sent for manual verification.”– Ryan Szeto, Senior Computer Vision Engineer at SafelyYou ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-151-lllmstxt|> ## Best of CVPR Event [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/13e655d277629be8c1b9e0c056902773db6b9621-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/13e655d277629be8c1b9e0c056902773db6b9621-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Meetups Best of CVPR – July 11, 2025 This event has ended, but you can still catch up! Watch the on-demand recordings and register for our [future events.](https://voxel51.com/events) Jul 11, 2025 9 AM Pacific Online. Register for the Zoom! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/2004f6377450a4bed1b21dc055ef6c3c89721e36-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Max Gutbrod OTH Regensburg Bio ![](https://cdn.sanity.io/images/h6toihm1/production/61264c83ea9cef5d2c92f7e00b77086783499c93-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Maregu Assefa Khalifa University Bio ![](https://cdn.sanity.io/images/h6toihm1/production/96e9eed25d0357627fbd471c8266a07afc05585f-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Aayush Dhakal Washington University in St. Louis Bio ![](https://cdn.sanity.io/images/h6toihm1/production/3b94241be288bdd975fc72cb5cefaa0f7eb8ac79-240x240.png?auto=format&dpr=2&fit=max&q=75&w=42) Rui Xiao Technical University of Munich Bio About this event Welcome to the Best of CVPR series, your virtual pass to some of the groundbreaking research, insights, and innovations that defined this year’s conference. Live streaming from the authors to you. Schedule OpenMIBOOD: Open Medical Imaging Benchmarks for Out-Of-Distribution Detection ![](https://cdn.sanity.io/images/h6toihm1/production/2004f6377450a4bed1b21dc055ef6c3c89721e36-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Max Gutbrod OTH Regensburg Bio As AI becomes more prevalent in fields like healthcare, ensuring its reliability under unexpected inputs is essential. We present OpenMIBOOD, a benchmarking framework for evaluating out-of-distribution (OOD) detection methods in medical imaging. It includes 14 datasets across three medical domains and categorizes them into in-distribution, near-OOD, and far-OOD groups to assess 24 post-hoc methods. Results show that OOD detection approaches effective in natural images often fail in medical contexts, highlighting the need for domain-specific benchmarks to ensure trustworthy AI in healthcare. DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation ![](https://cdn.sanity.io/images/h6toihm1/production/61264c83ea9cef5d2c92f7e00b77086783499c93-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Maregu Assefa Khalifa University Bio Semi-supervised medical image segmentation often suffers from class imbalance and high uncertainty due to pathology variability. We propose DyCON, a Dynamic Uncertainty-aware Consistency and Contrastive Learning framework that addresses these challenges via two novel losses: UnCL and FeCL. UnCL adaptively weights voxel-wise consistency based on uncertainty, initially focusing on uncertain regions and gradually shifting to confident ones. FeCL improves local feature discrimination under imbalance by applying dual focal mechanisms and adaptive entropy-based weighting to contrastive learning. RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings ![](https://cdn.sanity.io/images/h6toihm1/production/96e9eed25d0357627fbd471c8266a07afc05585f-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Aayush Dhakal Washington University in St. Louis Bio The choice of representation for geographic location significantly impacts the accuracy of models for a broad range of geospatial tasks, including fine-grained species classification, population density estimation, and biome classification. Recent works learn such representations by contrastively aligning geolocation\[lat,lon\] with co-located images. While these methods work exceptionally well, in this paper, we posit that the current training strategies fail to fully capture the important visual features. We provide an information-theoretic perspective on why the resulting embeddings from these methods discard crucial visual information that is important for many downstream tasks. To solve this problem, we propose a novel retrieval-augmented strategy called RANGE. We build our method on the intuition that the visual features of a location can be estimated by combining the visual features from multiple similar-looking locations. We show this retrieval strategy outperforms the existing state-of-the-art models with significant margins in most tasks. FLAIR: Fine-Grained Image Understanding through Language-Guided Representations ![](https://cdn.sanity.io/images/h6toihm1/production/3b94241be288bdd975fc72cb5cefaa0f7eb8ac79-240x240.png?auto=format&dpr=2&fit=max&q=75&w=96) Rui Xiao Technical University of Munich Bio CLIP excels at global image-text alignment but struggles with fine-grained visual understanding. In this talk, I present FLAIR—Fine-grained Language-informed Image Representations—which leverages long, detailed captions to learn localized image features. By conditioning attention pooling on diverse sub-captions, FLAIR generates text-specific image embeddings that enhance retrieval of fine-grained content. Our model outperforms existing methods on standard and newly proposed fine-grained retrieval benchmarks, and even enables strong zero-shot semantic segmentation—despite being trained on only 30M image-text pairs. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-152-lllmstxt|> ## Wildlife Conservation with AI [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ad5c1f638ed4aa77e2de95e816fded79e4b3b0dd-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Wildlife.ai Wildlife.ai uses FiftyOne to accelerate wildlife conservation Apr 1, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Wildlife.ai](https://wildlife.ai/) a New Zealand charity that uses artificial intelligence and machine learning for environmental education and protection. The organization works with grassroots wildlife conservation projects and its work is open source, adaptable, and scalable. > "Wildlife.ai is developing the Wildlife Watcher, an open source smart camera trap that captures images of a much wider variety of animals than is possible today, while also automatically producing valuable data insights using machine learning. The cameras have already outperformed traditional monitoring tools to record the presence of invasive and native species in New Zealand. > > To make the Wildlife Watchers available to the wider community, Wildlife.ai is using FiftyOne and [other open source software](https://github.com/wildlifeai/wai_data_tools) to enable users to easily analyze the camera data and create their own models." – Victor Anton, Founder and CEO at Wildlife.ai ![](https://cdn.sanity.io/images/h6toihm1/production/5a44fe8eef4a3215bf900455f6636a95b4739159-1190x824.png?auto=format&dpr=2&fit=max&q=75&w=1190) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-153-lllmstxt|> ## Berkshire Grey Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/72a7793467b8ea2576dceb4ce7dce9331cdc561c-2128x2129.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Berkshire Grey 3x Faster robotics dataset investigations: How Berkshire Grey drives productivity with FiftyOne Jun 6, 2025 Berkshire Grey accelerated robotic vision model development and improved training data quality by adopting FiftyOne for multimodal data. Sped up robotics data investigation and curation by 3x Curation workflows resulted in datasets with robot picks \> 1M Team collaboration features resulted in faster feedback loops improving efficiency Article content In this article [Challenge: Fragmented data management and slow investigation speed hampered development](https://voxel51.com/customers/berkshire-grey#e12c6aa30abb) [Solution: Faster, more flexible data exploration with FiftyOne](https://voxel51.com/customers/berkshire-grey#4ed5ca53f2e8) [Results: 3x Faster investigations and data curation](https://voxel51.com/customers/berkshire-grey#22a34e9db768) In this article [Challenge: Fragmented data management and slow investigation speed hampered development](https://voxel51.com/customers/berkshire-grey#e12c6aa30abb) [Solution: Faster, more flexible data exploration with FiftyOne](https://voxel51.com/customers/berkshire-grey#4ed5ca53f2e8) [Results: 3x Faster investigations and data curation](https://voxel51.com/customers/berkshire-grey#22a34e9db768) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [**Berkshire Grey**](https://www.berkshiregrey.com/) is a leading robotics company delivering AI-powered robotic solutions to automate picking, sorting, and transferring tasks for optimizing material handling in warehouses. By utilizing accurate computer vision models that enable millions of robotics picks, Berkshire Grey has enabled leading retailers, e-commerce providers, and logistics companies such as Maersk and FedEx to achieve over 50% reduction in labor requirements while improving productivity and reducing errors. > “FiftyOne has really helped us speed up investigations. For example, if we see a wrong suction cup grasping an item, we can quickly visualize the issue across all data sources and identify what went wrong.” — Dimitry Pechyoni, Senior Principal Machine Learning Engineer at Berkshire Grey ## **Challenge: Fragmented data management and slow investigation speed hampered development** Before adopting FiftyOne, Berkshire Grey’s team relied on an internally developed system for visualizing and analyzing multimodal data, mostly images, video, and point clouds from their robotic systems. The data came from a variety of sensors, with hundreds of thousands of robotic pick actions being tracked daily. While this data was invaluable for troubleshooting and model training, the existing tool had several drawbacks: - **Slow issue investigation:** The internal system lacked the visualization capabilities and interactivity for quick data exploration. - **Lack of interactivity:** The internal tool only allowed viewing one example at a time, further limiting their ability to analyze issues effectively. - **Complex integration and customization:** Adding visualizations for new sensors took months, and filtering based on multiple fields simultaneously was painfully slow. ## **Solution: Faster, more flexible data exploration with FiftyOne** Berkshire Grey discovered FiftyOne as the perfect fit for their data management and model investigation needs for their robotic systems. The platform’s ability to handle large, multimodal datasets with high performance and its user-friendly interface made it the perfect choice for their team’s needs. FiftyOne’s interactive features allowed Berkshire Grey to visualize large datasets in a fraction of the time, with improved filtering and querying capabilities. > “The speed at which FiftyOne fetches data from our system is incredibly fast. When you're dealing with millions of data samples, that’s a game-changer. Most data management systems just can’t handle that scale.” — Dimitry Pechyoni, Senior Principal Machine Learning Engineer at Berkshire Grey - **Data management:** The team developed custom data ingestion scripts, continuously importing data into FiftyOne, resulting in datasets with greater than 1M robot picks. - **Interactive filtering and tagging:** The intuitive filtering system of FiftyOne allowed teams to slice data effectively, reducing troubleshooting time for identifying problematic robotic picks. **Enhanced collaboration:** Tagging and URL-sharing features helped team collaboration, allowing members across multiple teams to quickly pinpoint and share interesting scenarios. ## **Results: 3x Faster investigations and data curation** After deploying FiftyOne, Berkshire Grey saw immediate improvements in their data workflows. New members were able to start using the tool’s features almost immediately after onboarding. - **3x Faster investigations**: The team was able to significantly reduce the time for issue investigation. What used to take more than a day could now be done in a fraction of the time. Filtering complex datasets became instant and interactive, boosting the team’s productivity and ability to diagnose problems. > “FiftyOne has really helped us speed up investigations. For example, if we see a wrong suction cup grasping an item, we can quickly visualize the issue across all data sources and identify what went wrong.” —Dimitry Pechyoni, Senior Principal Machine Learning Engineer at Berkshire Grey - **Improved data quality for model training**: By using FiftyOne for data curation, the team was able to tag, filter, and eliminate poor-quality samples quickly, which improved the overall quality of the data used for training, leading to better model performance. > “Without FiftyOne, our investigation capabilities would be severely limited. Model training would grind to a halt as it entirely depends on the tool’s data management. FiftyOne is foundational to our robotics operations.”—Dimitry Pechyoni, Senior Principal Machine Learning Engineer at Berkshire Grey ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-154-lllmstxt|> ## AI Data Modeling Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) AI Data Modeling for Visual AI: Key Metrics to Build Precise Models Apr 17, 2025 • 8 min read Article content In this article [Key Metrics for Evaluating Visual AI Data Models](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#7525505b5cae) [Beyond Pixel-Level Accuracy](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#73f371747269) [Leveraging FiftyOne for Data-Driven Metric Analysis](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#45ba808a5a54) [Final Insights](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#efca8dd615b4) [The Future of Visual Discovery](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#99fca5475542) [Explore the Jupyter Notebook](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#1271df1195f4) In this article [Key Metrics for Evaluating Visual AI Data Models](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#7525505b5cae) [Beyond Pixel-Level Accuracy](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#73f371747269) [Leveraging FiftyOne for Data-Driven Metric Analysis](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#45ba808a5a54) [Final Insights](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#efca8dd615b4) [The Future of Visual Discovery](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#99fca5475542) [Explore the Jupyter Notebook](https://voxel51.com/blog/ai-data-modeling-for-visual-ai-key-metrics-to-build-precise-models#1271df1195f4) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/e31020e93662e6c93b1cecd27f56fe3f80382cc5-640x426.gif?auto=format&dpr=2&fit=max&q=75&w=640) Modern artificial intelligence (AI) applications rely on increasingly complex data – often vast amounts of data generated by sensors, cameras, and user interactions – to power tasks like image classification, [object detection](https://en.wikipedia.org/wiki/Object_detection), and semantic segmentation. Yet collecting and labeling large volumes of raw data is only the beginning. To fully harness the capabilities of neural networks and deep learning, an effective data modeling process is essential. This means going beyond simple annotation to carefully structure input data, measure relevant metrics, and iteratively refine predictive models. In this article, we'll delve into AI data modeling for visual AI, explore key metrics for evaluating data models, and showcase how tools like [FiftyOne](https://voxel51.com/fiftyone/) can be used to drive iterative improvements. ### **The Importance of AI Data Modeling in Visual AI** [Data modeling](https://en.wikipedia.org/wiki/Data_modeling), the process of structuring, organizing, and preparing data for training and evaluating AI models is central to developing accurate visual AI systems, such as object detection or semantic segmentation networks, directly influencing how effectively insights can be extracted from diverse data types like images, videos, and associated metadata. It extends beyond merely labeling input-output variables, requiring a deep understanding of data structures, distributions, biases, class imbalances, labeling inconsistencies, and critical success metrics. Effective modeling involves not only capturing objects and bounding boxes but also documenting the business processes underlying data acquisition, preprocessing steps, and their potential biases, thus preserving meaningful variations while reducing noise and skew. Ultimately, structured AI data modeling bridges the gap between training datasets and the inputs required by advanced or popular AI models like Mask R-CNN or YOLO, enabling the creation of precise and reliable visual AI solutions. ### **Beyond Collection and Labeling** Even for deep learning approaches, simply gathering diverse data and applying bounding boxes or class labels is not enough. Data modeling digs deeper: - **Data Quality Checks:** Are classes well-represented, or do some dominate the dataset? Are annotations consistent, or do boundaries vary drastically between labelers? - **[Potential Biases](https://en.wikipedia.org/wiki/Bias_(statistics)):** Could the input data disadvantage certain categories or demographics? Are certain object classes underrepresented, leading to skewed predictions? - **Relevant Metrics:** Selecting the right key metrics ensures you accurately capture model performance across tasks like classification, detection, and segmentation. By evaluating these factors early, you can correct issues before they propagate, ultimately reducing confusion and creating models that generalize more effectively. ## **Key Metrics for Evaluating Visual AI Data Models** Evaluating visual AI systems demands more than a single accuracy number. Here are some core metrics that help gauge AI models across different tasks: ### **Object Detection** ![](https://cdn.sanity.io/images/h6toihm1/production/e7902fe3e61d1a326979a71587befe20769cce64-882x588.png?auto=format&dpr=2&fit=max&q=75&w=882) Object detection requires precise bounding boxes and accurate class labels. To evaluate detection accuracy, metrics such as Average Precision (AP) are commonly calculated at different Intersection over Union (IoU) thresholds, including AP@0.50 for strict bounding box overlap and AP@0.75 for a more lenient overlap. These metrics help determine how effectively a model localizes objects under varying degrees of precision. Additionally, localization errors can be evaluated by measuring the Center Point Distance, which quantifies how far a predicted bounding box center deviates from the ground truth, revealing difficulties with offset predictions. Analyzing Bounding Box Overlap separately from AP can further clarify whether bounding boxes systematically overshoot or undershoot actual object boundaries. ### **Image Classification** For image classification, you often rely on class probabilities: - **Confusion Matrix Analysis** - A confusion matrix shows correct predictions along the diagonal and misclassifications elsewhere. This helps identify which classes are most confused with one another and highlights potential model biases or labeling issues. - **Calibration Curves** - Calibration curves map predicted probabilities against actual outcomes. If your data modeling pipeline produces models that are consistently overconfident, calibration analysis can reveal where to adjust thresholding or weighting. \[caption id="attachment\_9804" align="aligncenter" width="800"\] Confusion matrix with model predictions across classes, with correct predictions along the diagonal\[/caption\] ### **Semantic Segmentation** [Semantic segmentation](https://docs.voxel51.com/user_guide/using_datasets.html#semantic-segmentation) assigns a class label to every pixel, making precise metric evaluation essential to gauge model quality, especially in scenarios involving imbalanced datasets or complex object shapes. Metrics like Mean Intersection-over-Union (mIoU) assess overall segmentation performance by averaging the IoU across all classes, while Weighted IoU addresses class imbalance by assigning higher importance to critical, underrepresented classes. Similarly, evaluating boundary accuracy through metrics such as the Boundary F1 Score ensures the precision of object delineation, particularly for thin or irregular shapes. The following Python functions provide compact implementations for these key metrics: ```python 1import numpy as np 2# Mean IoU (mIoU) 3def mean_iou(cm): 4 return np.nanmean(np.diag(cm) / (cm.sum(1) + cm.sum(0) - np.diag(cm))) 5# Weighted IoU 6def weighted_iou(cm, weights): 7 return np.nansum((np.diag(cm) / (cm.sum(1) + cm.sum(0) - np.diag(cm))) * weights) 8# Boundary F1 Score 9def boundary_f1(pred, true): 10 tp = (pred & true).sum() 11 precision = tp / (pred.sum() + 1e-6) 12 recall = tp / (true.sum() + 1e-6) 13 return 2 * precision * recall / (precision + recall + 1e-6) 14 ``` Using these metrics during data modeling allows for precise quantification of segmentation performance, enabling targeted refinements in datasets and predictive models. ## **Beyond Pixel-Level Accuracy** Even advanced metrics like IoU or AP only partially reflect real-world performance. Data modeling for artificial intelligence should also address explainability and fairness: \[caption id="attachment\_9805" align="aligncenter" width="640"\] A saliency map highlights which everyday items on the desk capture the model’s attention most.\[/caption\] ### **Explainability Metrics** - **Feature Importance** - Techniques can identify which input data regions or attributes drive the model's decisions. For instance, saliency maps can highlight critical portions of an image. - **Attribution Methods** - Methods like Grad-CAM or SHAP reveal which parts of the image the model weights most heavily. This is particularly relevant in high-stakes domains (e.g., medical imaging) where it's crucial to understand why a model flagged a specific region. **Fairness and Bias Metrics** AI systems should be designed to perform equitably and consistently across various demographic groups or classes. When a model achieves good results for certain groups but demonstrates poorer performance for others, it may reinforce existing biases present within the underlying data structures or highlight critical omissions in the data collection process. Evaluating model performance using group-based metrics is essential to addressing this issue. Specifically, analyzing results such as object detection accuracy across different demographics or geographic locations can reveal hidden biases within the model and data, allowing for corrective measures that promote fairness and inclusivity. ## **Leveraging FiftyOne for Data-Driven Metric Analysis** One of the most effective ways to operationalize the above metrics is to use data modeling tools like FiftyOne. FiftyOne helps you manage your valuable data assets by allowing you to interactively visualize results and query subsets for targeted analysis. ### **Interactive Visualization of Metrics** \[caption id="attachment\_9808" align="aligncenter" width="1024"\] Precision-Recall curves illustrating object detection performance, based on evaluation data in FiftyOne.\[/caption\] FiftyOne supports: - **[Precision-Recall](https://en.wikipedia.org/wiki/Precision_and_recall) Curves:** Adjust thresholds to see how your model's false positives trade off against false negatives in real time. - **Confusion Matrices:** Identify persistent misclassification patterns and evaluate how well each label is represented. - **Segmentation Masks:** Overlay predicted masks on ground truth to see pixel-level discrepancies, especially at object boundaries. ### **Filtering and Querying for In-Depth Analysis** \[caption id="attachment\_9809" align="aligncenter" width="1024"\] FiftyOne's filtering interface displays examples with borderline IoU values (0.50-0.52), highlighting detections barely meeting the threshold for user review.\[/caption\] FiftyOne lets you slice and filter data in powerful ways: - **Difficult Examples:** Quickly isolate images with poor IoU or bounding box overlap. These "failure samples" often highlight the biggest gaps in your data modeling approach. - **Low-Confidence Predictions:** Examine predictions where the model is uncertain. Such cases might indicate insufficient training data or overly complex object classes. ### **Iterative Refinement of Data Models** \[caption id="attachment\_9810" align="aligncenter" width="864"\] Diagram of a simple refinement loop that improves visual AI performance through analysis, metrics, and targeted updates.\[/caption\] Because business processes evolve and new input data continuously arrives, data modeling is never "one and done." FiftyOne promotes: 1. **Data Analysis:** Explore how your model behaves on specific subsets or problem classes. 2. **Metric Evaluation:** Re-run metrics like AP, IoU, or confusion matrices to see if changes in the dataset or model architecture lead to real improvements. 3. **Model Refinement:** Incorporate newly discovered insights into revised training pipelines (e.g., balancing underrepresented classes or adjusting thresholding logic.) ## **Final Insights** ### The Importance of Selecting the Right Metrics Building precise visual AI solutions goes far beyond generating bounding boxes or pixel labels. Thorough AI data modeling involves selecting key metrics, from Average Precision for detection to IoU for segmentation and beyond, to capture model behavior accurately. It requires continual monitoring of model calibration, boundary accuracy, and fairness across different segments. Simply put, the data modeling process shapes how AI models perceive the world. ### FiftyOne as an Essential Tool FiftyOne transforms raw metrics into actionable insights, enabling data scientists to quickly pinpoint failure cases, analyze low-confidence predictions, and refine data models iteratively. By leveraging interactive visualizations and querying capabilities, you can streamline the creating models process, leading to more effective data modeling and improved performance over time. ## **The Future of Visual Discovery** As future trends in artificial intelligence bring new complexities—richer image data, larger neural networks, and evolving business processes, the need for robust data modeling only grows. The combination of metrics-driven strategies and advanced platforms like FiftyOne ensures your predictive models remain accurate, equitable, and ready for the next wave of challenges in visual AI. ## **Explore the Jupyter Notebook** We’ve provided a companion [Jupyter notebook](https://colab.research.google.com/drive/16VPj6bDeGCnzO9lwmZPd2G6A0cdkG7sp) __ that demonstrates: - Classification (CIFAR-10, ImageNet-sample) with confusion matrices - Object detection ( [COCO](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation/) \+ YOLOv8) and IoU≥0.5 analysis - Semantic segmentation (VOC + FCN-ResNet50) with binary IoU - Iterative improvements via manual label fixes - Embedding visualization (CIFAR-10) using PCA - [Dataset exports](https://docs.voxel51.com/user_guide/export_datasets.html) to multiple formats with FiftyOne’s CLI Follow these examples to refine datasets, diagnose model errors, and streamline your AI workflow. **Image Citations** - Lin, Tsung-Yi, et al. "Microsoft COCO: Common Objects in Context." COCO Dataset 2017 Validation Split, cocodataset.org, 2017, [https://cocodataset.org/#home](https://cocodataset.org/#home). Accessed 25 Mar. 2025. [Visual AI](https://voxel51.com/blog/tag/visual-ai) [data-centric AI tooling](https://voxel51.com/blog/tag/data-centric-ai-tooling) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/d2e24d0a14de508f8ccd36eaffd3b909c9f193b9-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Image Embeddings Transform Computer Vision Capabilities\\ \\ Learn\\ \\ • \\ \\ Nov 25, 2024](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities) [![](https://cdn.sanity.io/images/h6toihm1/production/04eff439f21847f1068ea3f38b72f8bc7ec77f0e-2340x1308.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Image Similarity Search: Unlocking Pattern Detection in Visual Data\\ \\ Learn\\ \\ • \\ \\ Apr 16, 2025](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data) [![](https://cdn.sanity.io/images/h6toihm1/production/ca4f84addae3f8e97daad00cf856754301cbb5e0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Are Image Segmentation Maps Superior to Bounding Boxes?\\ \\ Learn\\ \\ • \\ \\ Feb 26, 2025](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-155-lllmstxt|> ## LanceDB Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/31cd5d87baf1e384ed51106820bab1b6de42f927-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) LanceDB LanceDB finds FiftyOne an invaluable developer tool for building its vector search database Apr 2, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [LanceDB](https://lancedb.com/) is a developer-friendly, serverless vector database for AI applications. Easy to start, easy to deploy, LanceDB is an embedded database that makes data management for LLMs frictionless. > "At LanceDB we are building a database for vector search and multi-modal data. We regularly use computer vision datasets, which requires inspecting them and validating them – to ensure the performance, quality, and generally verify that our product is functional. FiftyOne has proved to be an invaluable tool for our development work." – Jai Chopra, Head of Product at LanceDB ![](https://cdn.sanity.io/images/h6toihm1/production/1bd90c592b36222521fedb02603db6fbc92276cd-2186x1754.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-156-lllmstxt|> ## Annotation Savings Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Annotation-Savings [![](https://cdn.sanity.io/images/h6toihm1/production/990a31a85d850d41fb482ae29960a8fa10b18ecd-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Behind the math: how we built the annotation savings calculator\\ \\ Product & News\\ \\ • \\ \\ Jun 4, 2025](https://voxel51.com/blog/how-we-built-annotation-savings-estimator) ## Enough data wrangling.
 Request a demo. [Get started](https://voxel51.com/link-catcher) [Explore the Demo](https://voxel51.com/link-catcher) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-157-lllmstxt|> ## Visual Agents at CVPR 2025 [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision) Visual Agents at CVPR 2025 May 28, 2025 • 25 min read Article content In this article [AI Systems that See, Understand, and Act](https://voxel51.com/blog/visual-agents-at-cvpr-2025#57359b6c0bae) [Why This Research Wave Matters Now for Agentic AI](https://voxel51.com/blog/visual-agents-at-cvpr-2025#f7db74b50ffd) [From Research to Practical Breakthroughs](https://voxel51.com/blog/visual-agents-at-cvpr-2025#52700b9fbb8d) [The Visual Agent papers from CVPR I’m most excited about are:](https://voxel51.com/blog/visual-agents-at-cvpr-2025#39123337608c) [How are Visual Agents Different from Vision Language Models](https://voxel51.com/blog/visual-agents-at-cvpr-2025#3473398f23eb) [The Action-Perception Gap](https://voxel51.com/blog/visual-agents-at-cvpr-2025#2dbe1efd733d) [The Action Space Challenge](https://voxel51.com/blog/visual-agents-at-cvpr-2025#b199b6c3e821) [The Missing Embodiment](https://voxel51.com/blog/visual-agents-at-cvpr-2025#22922a779196) [From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons](https://voxel51.com/blog/visual-agents-at-cvpr-2025#d2418ecfb5eb) [ShowUI: Advanced Vision-Language-Action for GUI Interactions](https://voxel51.com/blog/visual-agents-at-cvpr-2025#6249fcc58026) [GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration](https://voxel51.com/blog/visual-agents-at-cvpr-2025#fa3591ec9251) [SpiritSight Agent: Advanced GUI Agent with One Look](https://voxel51.com/blog/visual-agents-at-cvpr-2025#c33810ff92a2) [ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems](https://voxel51.com/blog/visual-agents-at-cvpr-2025#9dda3e2a8f62) [The Future of Visual Agents is Moving from Perception to Interaction](https://voxel51.com/blog/visual-agents-at-cvpr-2025#52245ff34ec7) In this article [AI Systems that See, Understand, and Act](https://voxel51.com/blog/visual-agents-at-cvpr-2025#57359b6c0bae) [Why This Research Wave Matters Now for Agentic AI](https://voxel51.com/blog/visual-agents-at-cvpr-2025#f7db74b50ffd) [From Research to Practical Breakthroughs](https://voxel51.com/blog/visual-agents-at-cvpr-2025#52700b9fbb8d) [The Visual Agent papers from CVPR I’m most excited about are:](https://voxel51.com/blog/visual-agents-at-cvpr-2025#39123337608c) [How are Visual Agents Different from Vision Language Models](https://voxel51.com/blog/visual-agents-at-cvpr-2025#3473398f23eb) [The Action-Perception Gap](https://voxel51.com/blog/visual-agents-at-cvpr-2025#2dbe1efd733d) [The Action Space Challenge](https://voxel51.com/blog/visual-agents-at-cvpr-2025#b199b6c3e821) [The Missing Embodiment](https://voxel51.com/blog/visual-agents-at-cvpr-2025#22922a779196) [From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons](https://voxel51.com/blog/visual-agents-at-cvpr-2025#d2418ecfb5eb) [ShowUI: Advanced Vision-Language-Action for GUI Interactions](https://voxel51.com/blog/visual-agents-at-cvpr-2025#6249fcc58026) [GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration](https://voxel51.com/blog/visual-agents-at-cvpr-2025#fa3591ec9251) [SpiritSight Agent: Advanced GUI Agent with One Look](https://voxel51.com/blog/visual-agents-at-cvpr-2025#c33810ff92a2) [ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems](https://voxel51.com/blog/visual-agents-at-cvpr-2025#9dda3e2a8f62) [The Future of Visual Agents is Moving from Perception to Interaction](https://voxel51.com/blog/visual-agents-at-cvpr-2025#52245ff34ec7) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## AI Systems that See, Understand, and Act Visual Agents represent a significant advancement in Visual AI and agentic AI. These AI agents enable systems to perceive, understand, and interact with visual interfaces like humans do. As specialized Visual AI systems built upon Vision Language Models, they’re fundamentally changing how we approach the automation of visual interface interactions. This year’s **CVPR 2025** conference features some exciting research on Visual agents; these papers collectively signal that Visual Agents have moved from theoretical possibility to practical reality, representing one of the most exciting developments in applied AI. ## Why This Research Wave Matters Now for Agentic AI Recent advancements in foundational vision language models have finally provided the perceptual capabilities to tackle the long-standing challenge of GUI automation. This timing is critical as these capabilities align with the growing need for AI systems that can navigate the increasingly complex digital world on our behalf. With the CVPR 2025 deadline approaching and the official CVPR 2025 dates now published, interest in these breakthroughs has surged **.** What makes this CVPR particularly exciting is the complementary nature of the accepted research. Different teams have simultaneously tackled distinct aspects of the visual agent challenge: - Novel architectures specifically designed for the interleaved nature of vision-language-action sequences - Techniques for efficiently processing high-resolution screenshots without losing critical details - Specialized methods for precise element grounding that enable reliable interaction - Approaches for managing interaction histories across multiple observation-action cycles - Systems demonstrating cross-platform compatibility from web to mobile interfaces ## From Research to Practical Breakthroughs Visual Agents are moving from academic curiosity to practical technology. Digital interfaces are a part of nearly every aspect of work and life, and the ability to automate interactions with them becomes increasingly valuable. The timing of these breakthroughs couldn’t be more perfect. Early work in this area struggled with fundamental limitations—imprecise element localization, brittle performance across different interfaces, and limited action capabilities. The progression toward generalist agents capable of working across diverse environments represents a crucial evolutionary step, potentially leading to systems that can handle various visual interface tasks with human-like adaptability. Meanwhile, the emergence of collaborative AI systems indicates a future where visual agents coordinate with other specialized AI components to tackle complex workflows. ## The Visual Agent papers from CVPR I’m most excited about are: [From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons](https://arxiv.org/abs/2412.08442) [ShowUI: One Vision-Language-Action Model for GUI Visual Agent](https://arxiv.org/abs/2411.17465) [GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration](https://arxiv.org/abs/2503.17709) [SpiritSight Agent: Advanced GUI Agent with One Look](https://hzhiyuan.github.io/SpiritSight-Agent/) [ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems](https://xxyqwq.github.io/ComfyBench/) As we explore these papers in greater detail, we’ll see how they collectively represent incremental progress and a fundamental shift in what’s possible at the intersection of computer vision, natural language processing, and interactive systems. Most significantly, these advancements are happening just as the need for such technology is exploding. Before diving into the research in these papers, let’s clearly understand what Visual Agents are and the capabilities that position them at the forefront of agentic AI. If you’re wondering what is agentic ai, these examples provide a compelling answer. ## How are Visual Agents Different from Vision Language Models A **generalist Vision Language Model (VLM)**, or Multimodal Large Language Model (MLLM), is a foundation model primarily trained on vast quantities of paired text and image data. These models excel at tasks like understanding images, answering questions about visual content, generating captions, or processing text found within images. Their strength lies in **understanding and reasoning about visual and textual modalities**. On the other hand, a **Visual Agent** is a more specialized type of AI agent, typically built upon or adapted from a VLM/MLLM. Its specific goal is to perceive visual environments (like Graphic User Interfaces—GUIs) and act within them. These are often referred to as **Vision-Language-Action (VLA) models**. While they leverage the base VLM’s visual and language understanding capabilities, their core function is **interacting with and controlling an environment based on visual observation and natural language instructions**. The key differences and reasons why a general VLM alone is insufficient for GUI agent tasks lie in the **output modality, required specialized capabilities, and the nature of the task itself**. ### Output Modality: Actions vs. Text General VLMs are primarily designed to **output text**. Visual Agents, however, must produce **executable actions** that manipulate the GUI, such as clicking, typing, or scrolling. This fundamental difference in output type means a general VLM, without significant adaptation, cannot directly control a GUI environment. While general VLMs have visual understanding, GUI tasks demand specific skills that they typically lack or perform poorly on. ### Element Grounding A critical capability for a Visual Agent is accurately identifying and locating specific interactive elements (like buttons or input fields) within a GUI screenshot. Even the most powerful, general vision language models **significantly struggle with element grounding** from visual input alone. This is a major limitation because an agent cannot interact with elements it cannot reliably find. SpiritSight explicitly notes that the primary challenge in learning GUI navigation is learning the positional sub-policy needed for accurate grounding. ### Processing High-Resolution GUI Inputs Efficiently GUI screenshots are often high-resolution (e.g., 2K). Processing these high-resolution images results in very long token sequences, which is computationally expensive for models not specifically optimized for this. General VLMs may not have mechanisms like UI-Guided Visual Token Selection or Universal Block Parsing to efficiently handle UI visuals’ redundancy and structured nature while preserving necessary detail for grounding. This can lead to inefficiencies and high computational costs. ### Managing Interleaved Vision-Language-Action History GUI tasks often involve multi-step interactions where the agent needs to understand the context of previous observations (screenshots), user instructions, and actions taken. General VLMs are not inherently structured to effectively manage and reason over this complex, interleaved history of different modalities in a sequence. Techniques like Interleaved Vision-Language-Action Streaming are needed for this. ## The Action-Perception Gap The fundamental distinction between general Vision Language Models and Visual Agents lies in their core design purpose: **understanding versus acting**. This difference shapes everything from their architecture to their training objectives. General VLMs excel at passive analysis — describing what they see, answering questions about visual content, or reasoning based on visual input. Their output is primarily textual interpretation. While impressively capable at understanding visual scenes, they operate fundamentally as sophisticated perception systems, not interactive agents. Visual Agents, by contrast, are built for the dynamic loop of perception, decision, and action. They must understand what they’re seeing and use that understanding to make consequential decisions that change their operating environment. This requires a fundamentally different architecture optimized for: - Processing visual feedback resulting from their own actions - Maintaining state across multiple interaction steps - Generating precise, executable commands rather than descriptive text ## The Action Space Challenge Perhaps the most significant limitation preventing general VLMs from functioning as effective agents is their inability to generate structured, executable actions. Visual Agents require specialized output capabilities that translate understanding into precise commands: `CLICK(x=483, y=217) TYPE("search query") SCROLL_DOWN(amount=0.5)` These are structured commands with exact parameters that control interfaces. The action space varies significantly across platforms — web interfaces offer different interaction possibilities than mobile apps — requiring Visual Agents to adapt to diverse control paradigms. General VLMs lack both the training to generate such structured outputs and the architectural components to precisely locate interactive elements. Without specialized training on interaction trajectories showing the relationship between observations and resulting actions, these models cannot develop the procedural understanding necessary for sequential decision-making. ## The Missing Embodiment At their core, general VLMs lack what we might call “embodied experience” — the fundamental understanding of how actions affect environments and how to leverage those effects to accomplish goals. This gap can’t be addressed through simple adaptations but requires specialized training regimes with interactive data, architectural modifications to support action generation, and mechanisms for maintaining context across interaction sequences. The recent research on Visual Agent at CVPR 2025 is a significant leap forward — they bridge this fundamental action-perception gap that has long separated powerful understanding models from truly interactive AI systems. ## [From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons](https://arxiv.org/abs/2412.08442) This paper introduces the Generalist Embodied Agent (GEA), showcasing how Multimodal Large Language Models can be transformed into versatile agents capable of handling diverse real-world tasks. GEA is a huge advancement in creating Visual Agent systems that seamlessly operate across embodied AI, games, UI control, and planning domains. **Important links:** - [Paper on arXiv](https://arxiv.org/abs/2412.08442) - Dataset: Not yet released - Model: Not yet released This paper was interesting because it demonstrates how to create Visual Agents capable of real-world tasks, from object manipulation to game playing. The GEA model is on LLaVA-OneVision, which the authors picked for its ability to handle long-context interactions. The model has a novel **multi-embodiment action tokenizer** that unifies diverse action types, and employs a two-stage training process combining supervised learning and reinforcement learning 1. **Unified Model Architecture:** GEA adapts a pretrained MLLM to process environmental context and predict appropriate actions across various domains. 2. **Novel Action Tokenizer:** A sophisticated tokenizer based on Residual VQ-VAE enables the model to handle discrete and continuous actions uniformly. 3. **Two-Stage Training:** The model undergoes supervised fine-tuning followed by reinforcement learning, using a massive dataset of 2.2 million trajectories compiled from diverse sources, including human demonstrations, learned policies, and motion planners. ### Impressive Results GEA demonstrates strong cross-domain generalization, achieving competitive or state-of-the-art results across diverse benchmarks: - **Manipulation:** Reaches 90% success rate in [CALVIN](https://github.com/mees/calvin) (10% higher than comparable methods), outperforms baselines in Meta-World and Habitat Pick, though struggles with Maniskill’s challenging camera angles - **Gaming:** Achieves 44% of expert scores in [Procgen](https://github.com/openai/procgen) (outperforming specialist models) and surpasses Gato in Atari - **Navigation:** Matches Gato in [BabyAI](https://github.com/mila-iqia/babyai) despite using only visual inputs and fewer demonstrations - **UI Control:** Outperforms GPT-4o with Set-of-Mark prompting on [AndroidControl](https://github.com/google-research/google-research/tree/master/android_control) - **Planning:** Nearly matches specialist RL systems on LangR tasks Performance gains stem from cross-domain SFT training and targeted RL fine-tuning. Further improvements in [Maniskill](https://github.com/haosulab/ManiSkill), [Atari](https://dev1nw.github.io/atari-gpt/), and AndroidControl could be achieved by extending RL training to these domains. ### Key Lessons for Practitioners This paper gives us key insights and lessons learned specifically for **fine-tuning a general VLM to become a visual agent**. To adapt a general VLM into a capable visual agent for diverse embodied tasks, the authors point towards leveraging a strong pretrained MLLM base, learning a flexible action representation, and training with a combination of large-scale supervised data from multiple domains and subsequent online reinforcement learning. The seven key lessons are as follows: 1. **Start Strong:** A pretrained multimodal language model provides substantial advantages, particularly in visual tasks. In the paper, they used LLaVA-OneVision. However, given the fast pace of model progress, I’d be interested in seeing how [Qwen2.5-VL](https://github.com/harpreetsahota204/qwen2_5_vl)does. 2. **Unify Actions:** A sophisticated action tokenizer is essential for handling discrete and continuous actions across different domains. 3. **Two-Stage Training Works Best:** Supervised finetuning and reinforcement learning create the most capable agents. - **Stage 1: Supervised Finetuning (SFT):** First, fine-tune a pretrained VLM using supervised learning on a large collection of embodied experiences (demonstrations) from diverse domains. This stage adapts the VLM for embodied decision-making and is the process which creates the GEA-Base model. - **Stage 2: Online Reinforcement Learning (RL):** Follow up the SFT with online RL training in interactive simulators for a subset of domains. This stage uses the agent’s interactions to learn and improve. Techniques like LoRA are used for efficient fine-tuning during this stage. 4. **Diversity Matters:** Training with cross-domain data (2.2M+ trajectories) significantly boosts performance compared to single-domain training. 5. **Reinforcement Learning is Crucial:** Online RL is essential for developing robust agents to recover from mistakes. 6. **Build on a Strong Foundation:** RL is most effective after initial supervised training, rather than from scratch. In essence, to adapt a general VLM/MLLM into a capable visual agent for diverse embodied tasks, the paper’s lessons point towards **leveraging a strong pretrained MLLM base, learning a flexible action representation,and training with a combination of large-scale supervised data from multiple domains and subsequent online reinforcement learning**. ## ShowUI: Advanced Vision-Language-Action for GUI Interactions ShowUI introduces a vision-language-action (VLA) model specifically designed for GUI visual agents that operate in the digital world. Unlike traditional GUI automation methods that rely on metadata like HTML, ShowUI takes a more human-like approach by focusing on visual perception and interaction. The research addresses key challenges: - Expensive visual modeling of high-resolution screenshots - Managing complex vision-language-action sequences Effectively utilizing diverse training data **Important links:** - [Paper on arXiv](https://arxiv.org/abs/2411.17465) - [GitHub](https://github.com/showlab/ShowUI) - [Project Demo](https://huggingface.co/spaces/showlab/ShowUI) - [ShowUI Web Dataset on Hugging Face](https://huggingface.co/datasets/Voxel51/ShowUI_Web) - [ShowUI Desktop Dataset on Hugging Face](https://huggingface.co/datasets/Voxel51/ShowUI_desktop) - [Model on Hugging Face](https://huggingface.co/showlab/ShowUI-2B) ### The ShowUI Dataset Rather than using a massive dataset, ShowUI employs a carefully curated corpus of 256K data instances with 2.7M element annotations across various platforms: - **Web data:** 22K screenshots with 576K visual elements, deliberately filtered to focus on interactive elements rather than static text - **Mobile data:** 97K screenshots from the AMEX dataset with valuable functionality descriptions - **Desktop data:** Limited 100 screenshots augmented with GPT-4o assistance to create diverse queries - **Navigation data:** 137K tasks from GUIAct for web and mobile navigation A key innovation was the rebalanced sampling strategy that ensured fair exposure to each data type during training, despite significant differences in dataset sizes. ### The ShowUI Model ShowUI is built on the Qwen2-VL-2B foundation and introduces three technical innovations: 1. **UI-Guided Visual Token Selection:** Reduces computational costs by treating screenshots as connected graphs, identifying redundant visual areas while preserving important elements. This approach reduces visual tokens by 33% and speeds up training by 1.4×. 2. **Interleaved Vision-Language-Action Streaming:** Flexibly handles both multi-step navigation tasks (with visual-action history tracking) and single-screenshot, multi-action tasks through a standardized JSON action format. 3. **Small-scale, High-quality Dataset:** Demonstrates that careful curation and balanced sampling can outperform larger, noisier datasets. Despite its relatively small size and minimal training data, ShowUI achieves 75.1% accuracy in zero-shot screenshot grounding, setting a new standard for lightweight GUI agents. ### Key Lessons for Practitioners 1. **Leverage UI structure:** GUI screenshots contain inherent patterns and redundancies that can be exploited to reduce computational costs without losing important information. 2. **Standardize actions:** Use structured formats like JSON for actions and provide documentation in prompts to encourage systematic behavior. 3. **Focus on data quality:** Carefully analyze and select the most informative data types rather than simply collecting more data. For GUIs, visually rich elements often provide more value than static text. 4. **Augment limited data:** When facing scarce data (like the desktop examples), use large language models to generate diverse queries around existing annotations. 5. **Balance your datasets:** Implement sampling strategies that ensure all data types receive adequate representation during training, regardless of their original size. 6. **Start lightweight:** You don’t need massive models for effective GUI agents. A well-tuned smaller model with thoughtful data curation and efficient visual processing can deliver excellent results. This research demonstrates that through careful design choices and targeted optimizations, even lightweight models can achieve state-of-the-art performance in complex GUI interaction tasks. ## [GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration](https://arxiv.org/abs/2503.17709) GUI-Xplore introduces a novel dataset designed to overcome limitations in existing GUI agent systems. While current solutions often struggle with generalization across different applications, GUI-Xplore addresses this challenge by providing rich exploration context and expanding beyond basic navigation tasks. The research pairs this innovative dataset with Xplore-Agent, a baseline model demonstrating significantly improved performance in unfamiliar app environments through its unique exploration-based approach. **Important links:** - [Paper on arXiv](https://arxiv.org/abs/2503.17709) - [Dataset on Hugging Face](https://huggingface.co/datasets/9211sun/GUI-Xplore) - [GitHub](https://github.com/921112343/GUI-Xplore) ### The GUI-Xplore Dataset The dataset comprises exploration videos from 312 apps spanning 6 primary software domains and 33 sub-categories. With 115 hours of exploratory content (averaging 23.73 minutes per app), it delivers comprehensive coverage of real-world GUI interactions. GUI-Xplore includes over 32,500 question-answer pairs across five carefully designed hierarchical tasks: - **Application Overview & Page Analysis:** Testing understanding of app functions and specific screens - **Application Usage:** Evaluating the ability to infer operation sequences - **Action Recall & Sequence Verification:** Assessing comprehension of temporal and logical relationships What makes GUI-Xplore unique is its exploration-first approach, which provides contextual app knowledge that enables agents to adapt to new environments, similar to how humans explore unfamiliar interfaces. ### The Xplore-Agent Model Xplore-Agent leverages the dataset through a two-stage pipeline: 1. **Action-aware GUI Modeling:** Extracts key frames from exploration videos using luminance difference detection, then converts these frames into structured textual representations of GUI elements and interactions. 2. **Graph-Guided Environment Reasoning:** Constructs a GUI Transition Graph that maps complex page relationships and interaction patterns, then guides an LLM’s reasoning across the five downstream tasks. The authors dubbed this the “exploration-then-reasoning paradigm,” an innovation in how GUI agents are trained. Rather than immediately attempting tasks in unfamiliar environments (as traditional approaches do), this paradigm involves: 1. First, **exploring** the application interface to gather context about its structure, elements, and interaction patterns. 2. Then, **reasoning** about specific tasks using the knowledge gained during exploration This approach mirrors how humans naturally interact with new software — we typically explore an interface before attempting specific tasks. GUI-Xplore facilitates this through its exploration videos, and Xplore-Agent implements it via its two-stage pipeline (Action-aware GUI Modeling followed by Graph-Guided Environment Reasoning). This approach showed a 10% performance improvement over state-of-the-art methods when tested on unfamiliar applications, demonstrating the effectiveness of exploration-based learning. ### Key Lessons for Practitioners - **Context is crucial:** Pre-exploration of interfaces dramatically improves agent performance in unfamiliar environments. - **Move beyond simple automation:** Complex tasks require understanding local interactions and global app structure. - **Structure matters:** Explicitly modeling UI transitions as graphs helps agents navigate complex application flows. - **Efficient processing is essential:** Converting dense visual information into structured representations makes it manageable for language models. - **Operational understanding remains challenging:** While Xplore-Agent improves, understanding temporal and logical relationships in action sequences still presents significant research opportunities. Using the exploration-then-reasoning paradigm, developers can create more adaptable GUI agents that better mimic human approaches to navigating unfamiliar interfaces. This could potentially unlock more natural and effective human-computer interaction. ## SpiritSight Agent: Advanced GUI Agent with One Look SpiritSight is designed to help users interact with graphical interfaces by automatically making decisions based on screenshots. This research tackles a fundamental challenge that has limited previous vision-based agents: **poor element grounding** (the ability to accurately identify and locate GUI elements like buttons and text boxes). While vision-based approaches offer better cross-platform compatibility than methods requiring HTML or XML data, but they’ve historically struggled with precision locating interface elements. **Important links:** - [Paper on arXiv](https://arxiv.org/abs/2503.03196) - [Project page](https://hzhiyuan.github.io/SpiritSight-Agent/) - [Dataset on Hugging Face](https://huggingface.co/datasets/SenseLLM/GUI-Lasagne-L1) - [Models on Hugging Face](https://huggingface.co/collections/SenseLLM/spiritsight-67c7d2bd1149cf108b162009) ### The GUI-Lasagne Dataset The GUI-Lasagne dataset is at the heart of SpiritSight’s capabilities, a meticulously structured collection comprising 5.73 million samples gathered using scalable, cost-effective methods from real-world sources. The GUI-Lasagne dataset gets its name from its layered structure, designed to systematically build agentic capabilities from foundational skills to complex navigation. This hierarchical approach equips SpiritSight with robust GUI understanding and grounding capabilities through three distinct levels: **Level One: Visual-Text Alignment (3M samples)** - **Purpose:** Builds foundational ability to recognize and locate text/icon elements - **Key tasks:** text2bbox (locate elements from text), bbox2text (recognize content within areas), and bbox2dom (understand GUI layout) - **Collection:** Gathered from real-world web and mobile interfaces using automated tools - **Emphasis:** Intentionally contains the most abundant data to develop robust grounding capabilities **Level Two: Visual-Function Alignment (1.5M samples)** - **Purpose:** Teaches locating elements based on their function - **Collection:** Synthesized using powerful vision models to generate functional descriptions - **Validation:** Human-verified with 90.9% acceptance rate - **Output:** Creates function2bbox pairs linking element purposes to their locations **Level Three: Visual Navigation (0.64M samples)** - **Purpose:** Trains on complete navigation trajectories - **Innovation:** Cleaned using GPT-4o with Chain-of-Thought reasoning to filter out incorrect labels - **Quality:** Human validation confirmed 93.7% reliability in the cleaning process This hierarchical approach deliberately builds strong grounding abilities before addressing complex navigation tasks, with 90% of data dedicated to grounding (Levels One and Two). The dataset supports both web and mobile platforms and includes English and Chinese samples, enabling cross-platform and cross-lingual capabilities. Ablation studies confirmed each level’s value, with even mobile navigation data improving web navigation performance through cross-platform knowledge transfer. ### The SpiritSight Model SpiritSight introduces a novel technical approach called Universal Block Parsing (UBP) that solves a fundamental problem in processing high-resolution GUI screenshots: - **Resolves Positional Ambiguity:** UBP replaces traditional global coordinates with block-specific coordinates, creating clear one-to-one mappings between visual inputs and element locations. - **Enhances Spatial Understanding:** Incorporates 2D Block-wise Position Embedding to preserve spatial relationships between interface elements. These innovations enable SpiritSight to: - **Achieve Superior Performance:** Outperforms other advanced methods across diverse GUI benchmarks. - **Work Across Platforms:** Functions effectively on web and mobile interfaces without platform-specific adaptation. - **Scale Effectively:** Available in different model sizes (2B, 8B, 26B parameters) to balance performance and resource requirements. - **Operate End-to-End:** Processes screenshots directly to actions without requiring intermediate tools like OCR or candidate element extraction. By solving the critical element grounding challenge that has limited previous approaches, SpiritSight represents a significant step toward truly versatile, vision-based AI assistants that can help users navigate any graphical interface with unprecedented accuracy and reliability. The model is available in three sizes (2B, 8B, and 26B parameters) and evaluations show it outperforms other advanced methods across diverse GUI navigation benchmarks, demonstrating exceptional cross-platform compatibility. ### Key Lessons for Practitioners This research demonstrates that with carefully designed datasets and innovative methods like UBP, vision-based GUI agents can achieve the accuracy and reliability needed for practical applications across diverse interface environments. If you’re developing your own GUI agent or dataset, consider these critical insights: 1. **Prioritize element grounding** — The ability to locate interface elements accurately is the foundation of effective GUI agents. Dedicate significant training data to this skill. 2. **Structure datasets hierarchically** — Build from fundamental skills (recognition, grounding) to complex tasks (navigation) for stronger learning foundations. 3. **Address coordinate ambiguity** — When working with high-resolution inputs, consider techniques like UBP to resolve positional ambiguity for better grounding accuracy. 4. **Quality trumps quantity** — While scale matters, carefully filtered and cleaned data is more valuable than raw volume. Consider using Chain-of-Thought reasoning to structure and verify navigation data. 5. **Embrace cross-platform training** — Including diverse GUI environments (web, mobile) in training data enhances versatility and generalization. Mobile navigation data can even help with web navigation tasks. 6. **Consider end-to-end approaches** — With the right dataset and methods, end-to-end vision-based approaches can overcome previous limitations and achieve impressive performance without complex multi-stage pipelines. ## ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems **Important links:** - [Paper on arXiv](https://arxiv.org/abs/2409.01392) - [GitHub](https://github.com/xxyQwQ/ComfyBench) - [Project page](https://xxyqwq.github.io/ComfyBench/) This research introduces **ComfyBench** and **ComfyAgent** as contributions to a new frontier in Visual Agent research: using LLM-based agents to design collaborative AI systems autonomously. This approach represents a significant paradigm shift from traditional AI research, which has primarily focused on developing **monolithic models** to maximize performance on specific tasks. Instead, this work explores how agents can design complex systems that integrate multiple specialized models and tools to achieve more sophisticated outcomes. **ComfyBench** stands as the **first-of-its-kind comprehensive benchmark** specifically designed to evaluate agents’ capabilities in designing and executing collaborative AI systems within ComfyUI — an open-source platform where users construct workflows by connecting nodes (representing different models or tools) in directed acyclic graphs (DAGs). Complementing this benchmark is **ComfyAgent**, a novel framework built upon two core innovations: representing workflows with **code** rather than other representations, and employing a sophisticated **multi-agent system** with specialized roles (planning, retrieval, adaptation, refinement) that collaborate to overcome limitations like context constraints and hallucination. ### The ComfyBench Benchmark ComfyBench evaluates agents’ ability to construct ComfyUI workflows through 200 diverse tasks, with documentation for 3,205 nodes and 20 tutorial workflows as resources. Tasks may include auxiliary media requiring specific processing. The benchmark uses three difficulty levels: - 100 “Vanilla” tasks: Basic workflow adaptations - 60 “Complex” tasks: Integration of multiple workflow techniques - 40 “Creative” tasks: Pushing beyond imitation toward innovation These challenges test visual programming logic, natural language translation, tool selection, multi-modal reasoning, and parameter tuning. The most advanced tasks require integrating techniques across domains like image generation and video processing. Unlike benchmarks measuring output quality, ComfyBench assesses agents’ fundamental ability to orchestrate AI components into functional visual systems. ### Evaluation Metrics The research introduces novel evaluation metrics specifically designed for assessing workflow generation. Traditional metrics for image or video generation aren’t applicable here because the focus is on evaluating the generated workflows themselves, not just their outputs. Two progressive evaluation metrics are employed: **1\. Pass Rate** - Measures the ratio of tasks where generated workflows are syntactically and semantically correct - A task is marked as “passed” only if the server successfully executes the workflow and returns a success message - This metric verifies the functional correctness of the designed system **2\. Resolve Rate** - Measures the ratio of tasks where workflows produce results matching the task requirements - Evaluation uses Visual Language Models (VLMs), specifically GPT-4o, to assess alignment between outputs and instructions - The VLM reviews both the task instruction and generated output, providing a True/False judgment on compliance This two-tier evaluation uniquely assesses an agent’s ability to design collaborative AI systems, both in terms of process (can the workflow execute?) and outcome (does it achieve the desired result?). ### The ComfyAgent Framework ComfyAgent enables LLMs to autonomously design collaborative AI workflows in ComfyUI through two key innovations: 1. **Python-like Code Representation** — Workflows are represented in code format rather than JSON or element lists, leveraging LLMs’ code generation abilities and providing richer semantic information. Ablation studies confirm this is the most effective format. 2. **Multi-Agent Architecture** — Addresses single-agent limitations through specialized agents: - **PlanAgent:** Creates and updates the workflow strategy - **CombineAgent:** Integrates multiple workflows - **AdaptAgent:** Adjusts workflow parameters - **RefineAgent:** Checks and fixes errors - **RetrieveAgent:** Gathers relevant knowledge A memory system stores history, reference materials, and the current workspace. Removing any agent component reduces overall performance, confirming each plays a vital role. ### Key Lessons for Practitioners 1. **Representation Matters:** Code-based workflow representation significantly outperforms other formats like JSON or element lists. 2. **Knowledge Retrieval Is Essential:** Effective agents must retrieve and utilize documentation and example workflows rather than relying solely on their inherent knowledge. 3. **Multi-Agent Architecture Works Better:** Breaking down complex workflow design into specialized roles (planning, retrieval, adaptation, refinement) helps overcome limitations like context windows and hallucination. 4. **Creative Tasks Remain Challenging:** While current agents perform reasonably well on simpler tasks, they struggle with novel applications — ComfyAgent resolved only 15% of creative tasks. 5. **Dataset Construction Advice:** When building similar benchmarks, focus on comprehensive documentation, diverse tiered tasks, well-annotated examples, and reliable automated evaluation methods. This research paves the way for more intelligent collaborative AI systems, though significant challenges remain in developing agents that can streamline rather than just imitate existing workflows. ## The Future of Visual Agents is Moving from Perception to Interaction The CVPR 2025 papers mark a revolutionary leap in visual AI. Visual Agents have evolved from theory to reality, solving fundamental challenges that previously limited their capabilities. GEA’s unified action tokenizer, ShowUI’s efficient visual processing, GUI-Xplore’s exploration paradigm, and SpiritSight’s Universal Block Parsing collectively bridge the gap between visual understanding and action. Meanwhile, ComfyBench points toward agents that can orchestrate entire AI systems. #### This transformation arrives just as digital interfaces permeate every aspect of life. The ability to automate visual interactions promises to streamline workflows, enhance accessibility, and enable entirely new capabilities. We’re witnessing the early stages of truly agentic Visual AI systems— **agent AI** that don’t just perceive the world but meaningfully act within it. CVPR 2025 will likely be remembered as the tipping point where visual agents made the crucial transition from possibility to practicality. [CVPR](https://voxel51.com/blog/tag/cvpr) [Visual AI](https://voxel51.com/blog/tag/visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/a41a0477c7a98264f600772e9568607d070eea59-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Sahota Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/db7358784a18a2ffa365704e7d941e73fbdf1fcd-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 1\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-1) [![](https://cdn.sanity.io/images/h6toihm1/production/c63b546423b6cec32ccccc7df6bc4e0fced0b1a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best of CVPR 2025 Series – Day 2\\ \\ Computer Vision\\ \\ • \\ \\ May 29, 2025](https://voxel51.com/blog/the-best-of-cvpr-2025-series-day-2) [![](https://cdn.sanity.io/images/h6toihm1/production/6af33def6d297e2382d387e224e16451c95876af-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Why the “Annotate Everything” Era in Automotive AI Is Over\\ \\ Computer Vision\\ \\ • \\ \\ Jul 24, 2025](https://voxel51.com/blog/smarter-automotive-datasets-selection) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-158-lllmstxt|> ## Seafar Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/76bbb00710f5018bbd3eca2d5e71f4675a4ce26d-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Seafar FiftyOne strengthens Seafar’s solutions for autonomous vessels for shipping Apr 12, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) Seafar NV is an independent ship management company offering services for remote navigation, specifically unmanned, autonomous, and crew-reduced, semi-autonomous vessels for ship owners and shipping companies, with an emphasis on effective and safe operations. > "FiftyOne enabled us to strengthen our data-centric computer vision flows. Through their effective UI, we managed to explore and understand the ingested image data from our large Inland Waterways fleet. This enabled us to build strong baseline object detection models for maritime obstacles detection. On top of that, the Python SDK is a crucial component of our MLOps pipelines, enabling us to automate data pre-processing, smart selection, and continuous improvement of models."​ – Kais Bedioui, Computer Vision Engineer at Seafar ![](https://cdn.sanity.io/images/h6toihm1/production/1b36dae4ce7c841dc84b008e622c2d37b011cc01-2308x1608.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-159-lllmstxt|> ## Van der Maaten's AGI Roadmap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Van der Maaten’s Three-System Roadmap to AGI Is Brilliantly Pragmatic Jun 17, 2025 • 3 min read Article content In this article [System 1: We’ve Nearly Maxed Out “Thinking Fast”](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic#c77436c4b420) [System 2: “Thinking Slow” Is Where the Action Is](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic#3cc04223ff88) [System 3: “Thinking Together” Will Define the AGI Era](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic#b199f184b007) [The Twin Challenges: Scaling and Integration](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic#a89912659303) [Why This Framework Matters](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic#0715fca239d0) In this article [System 1: We’ve Nearly Maxed Out “Thinking Fast”](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic#c77436c4b420) [System 2: “Thinking Slow” Is Where the Action Is](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic#3cc04223ff88) [System 3: “Thinking Together” Will Define the AGI Era](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic#b199f184b007) [The Twin Challenges: Scaling and Integration](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic#a89912659303) [Why This Framework Matters](https://voxel51.com/blog/van-der-maaten-s-three-system-roadmap-to-agi-is-brilliantly-pragmatic#0715fca239d0) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/09cb438b8a2eac79dae3d29e88ec93ff79b709d1-1400x599.webp?auto=format&dpr=2&fit=max&q=75&w=1400) [Laurens Van der Maaten](https://lvdmaaten.github.io/) has articulated a compelling framework for AGI development at CVPR 2025. His vision centers on “human agents intelligence” — a future where networks of specialized AI and human agents collaborate to solve problems. This represents a fundamental shift from thinking about AGI as a single superintelligent system to viewing it as an emergent property of interconnected intelligence. The approach acknowledges what many AI researchers have ignored: true intelligence is inherently social and collaborative. This perspective cuts through the hype to deliver a roadmap that makes sense. ![](https://cdn.sanity.io/images/h6toihm1/production/2cd3b2b3bd71de63e2cab4e57977d03b8db0abeb-1400x756.webp?auto=format&dpr=2&fit=max&q=75&w=1400) ## System 1: We’ve Nearly Maxed Out “Thinking Fast” The reactive, instinctive capabilities of large language models represent the low-hanging fruit of AI development. Current models like Llama 4 have gotten remarkably good at generating immediate responses based on their pre-trained knowledge. The massive gains we’ve seen from scaling up parameters and training data primarily benefit these System 1 capabilities, which essentially rely on pattern recognition and memorization. However, we’re witnessing diminishing returns from this approach, suggesting we’re approaching the ceiling of what pure pre-training can accomplish. The real breakthroughs will come from what happens next. ## System 2: “Thinking Slow” Is Where the Action Is ![](https://cdn.sanity.io/images/h6toihm1/production/7992f2cda6006709ab88937f36bee7000cfda751-1400x584.webp?auto=format&dpr=2&fit=max&q=75&w=1400) Reasoning capabilities represent the current frontier of AI research and development. Today’s most promising advancements come from models that can deliberate, consider multiple options, and work through problems step-by-step using “scratchpads.” These reasoning models utilize reinforcement learning to effectively utilize their working memory, thereby dramatically improving performance on complex tasks such as math and coding. The shift from pure pattern matching to explicit reasoning is an “evolutionary” step that addresses many of the limitations of current AI systems. This is where we’ll see the biggest improvements in the next few years. ## System 3: “Thinking Together” Will Define the AGI Era ![](https://cdn.sanity.io/images/h6toihm1/production/0465a00a9f65dc6f29cb8943c027ad30d2e66c12-1400x584.webp?auto=format&dpr=2&fit=max&q=75&w=1400) The true path to AGI runs through multi-agent collaboration networks. Creating systems where specialized agents can find each other, communicate effectively, and collaborate on complex tasks requires solving entirely new classes of problems. We’ll need to democratize agent creation, build platforms for agent discovery and interaction, develop standardized communication protocols, and implement sophisticated “theory of mind” capabilities. The ultimate challenge, beyond making smarter individual AIs, is orchestrating efficient collaboration between systems with complementary skills. This third system represents the most ambitious and potentially revolutionary aspect of Van der Maaten’s vision. ## The Twin Challenges: Scaling and Integration ![](https://cdn.sanity.io/images/h6toihm1/production/a44b1e1837cc90f96c756c355d5ced396a5f6545-1400x591.webp?auto=format&dpr=2&fit=max&q=75&w=1400) Building these systems isn’t just hard — it’s “rocket science every step along the way.” The AI community faces two fundamental challenges that span all three systems. Scaling must evolve beyond simply increasing parameters to include scaling agent networks and interactions. Integration requires harmonizing vertical capabilities (like code generation or visual understanding) with horizontal capabilities (like consistent tone and instruction-following). These challenges demand innovations that go beyond our current approaches to AI development. We need breakthroughs in multi-agent reinforcement learning and mechanism design that AI researchers have barely begun to explore. ## Why This Framework Matters ![](https://cdn.sanity.io/images/h6toihm1/production/c56ea6d1f12c3656061130994961d35cb8b59d38-1400x587.webp?auto=format&dpr=2&fit=max&q=75&w=1400) Van der Maaten’s three-system model is a pragmatic alternative to the magical thinking that dominates the discourse on AGI. By breaking down the path to AGI into discrete, understandable systems with specific capabilities and challenges, this framework provides a roadmap that feels both ambitious and achievable. It acknowledges the social nature of intelligence and the necessity of collaboration rather than focusing solely on individual superintelligence. The approach also highlights underexplored areas where innovations are needed, pointing researchers in productive directions. This could be the blueprint that finally moves AGI discussions from science fiction to science. [LLMs](https://voxel51.com/blog/tag/llms) [CVPR](https://voxel51.com/blog/tag/cvpr) [AGI](https://voxel51.com/blog/tag/agi) ![](https://cdn.sanity.io/images/h6toihm1/production/a41a0477c7a98264f600772e9568607d070eea59-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Sahota Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/7ec61c80f387b16f24b4b2fe33f804264a38b486-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Rethinking How We Evaluate Multimodal AI\\ \\ Event Recaps\\ \\ • \\ \\ Jun 12, 2025](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/e468545aa08daf6c6829d2593ffd8b5457c7dee5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ NVIDIA’s C-RADIOv3 is the Vision Encoder You Should Be Using\\ \\ Event Recaps, Integrations\\ \\ • \\ \\ Jun 23, 2025](https://voxel51.com/blog/nvidia-c-radiov3-is-the-vision-encoder-you-should-be-using) [![](https://cdn.sanity.io/images/h6toihm1/production/1798d34efc8956a6696377fb7776886ec0e61092-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ UnCommon Objects in 3D\\ \\ Datasets, Event Recaps\\ \\ • \\ \\ Jun 18, 2025](https://voxel51.com/blog/uncommon-objects-in-3d) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-160-lllmstxt|> ## Zero-Shot Auto-Labeling Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [ML@Voxel51](https://voxel51.com/blog/category/mlvoxel51) Zero-shot auto-labeling rivals human performance Jun 4, 2025 • 8 min read Article content In this article [How far can zero-shot auto labeling take us in the quest for labeled datasets?](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#d40e4d5773b8) [Verified Auto Labeling delivers up to 95% model performance on downstream inference](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#c7152bb3163a) [Verified Auto Labeling reduces annotation costs by 100,000x](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#1b3993879a50) [Clean labels aren’t always better: how confidence thresholds impact model performance](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#0cf33a3c45bb) [Engineering tradeoffs at scale: why model selection matters in auto-labeling](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#0dbefb91a5b1) [What is the best auto-labeling tool?](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#18102359a615) In this article [How far can zero-shot auto labeling take us in the quest for labeled datasets?](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#d40e4d5773b8) [Verified Auto Labeling delivers up to 95% model performance on downstream inference](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#c7152bb3163a) [Verified Auto Labeling reduces annotation costs by 100,000x](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#1b3993879a50) [Clean labels aren’t always better: how confidence thresholds impact model performance](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#0cf33a3c45bb) [Engineering tradeoffs at scale: why model selection matters in auto-labeling](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#0dbefb91a5b1) [What is the best auto-labeling tool?](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance#18102359a615) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) One of the biggest bottlenecks in deploying visual AI and computer vision is annotation — the costly, time-consuming process of manually labeling images to train machine learning models. For years, AI hype reinforced the idea that more labels meant better models, enabling annotation providers like Scale AI to become billion-dollar giants. That model no longer fits today’s AI pipelines. Today, we’re introducing Verified Auto Labeling,a new approach to AI-assisted annotation that combines Voxel51’s [expertise in data curation](https://voxel51.com/curation) with automated labeling and QA workflows. Our research paper, _Auto-Labeling Data for Object Detection_, establishes new benchmarks that demonstrate that **Verified Auto Labeling achieves up to 95% of human-level performance — while cutting labeling costs by up to 100,000×.** Yes, you read that right—100,000×. [Read the research paper](https://voxel51.com/whitepapers/auto-labeling-data-for-object-detection) ## How far can zero-shot auto labeling take us in the quest for labeled datasets? As foundation models become increasingly sophisticated, it's widely believed that auto-labeling will significantly reduce the need for [human annotation](https://voxel51.com/blog/why-quality-dataset-annotation-is-key-to-machine-learning). But how effective are today's zero-shot models in practice? To find out, we benchmarked leading vision-language models (VLMs) as foundation models—including [YOLOE](https://docs.voxel51.com/integrations/ultralytics.html#open-vocabulary-segmentation), [YOLO-World](https://docs.voxel51.com/integrations/ultralytics.html#open-vocabulary-detection), and [Grounding DINO](https://huggingface.co/docs/transformers/en/model_doc/grounding-dino) — across four widely-used datasets: [Berkeley Deep Drive](https://docs.voxel51.com/dataset_zoo/datasets.html#bdd100k) (BDD, autonomous driving), [Microsoft Common Objects in Context](https://docs.voxel51.com/dataset_zoo/datasets.html#coco-2014) (COCO), [Large Vocabulary Instance Segmentation](https://www.lvisdataset.org/) (LVIS, high complexity), and [PASCAL Visual Object Classes](https://docs.voxel51.com/dataset_zoo/datasets.html#voc-2007) (VOC, general imagery). These datasets span basic object categories to challenging, long-tail distributions. ![Our research compared auto-labeling and human labeling approaches using object detection tasks. Starting from identical unlabeled images, we generated labels both automatically (using foundation models like YOLO-World and Grounding DINO) and manually (human annotators). We then trained lightweight inference models from each label set without pre-trained weights or additional fine-tuning, allowing a fair and practical comparison of labeling costs, label quality, and downstream validation accuracy (measured via mAP).](https://cdn.sanity.io/images/h6toihm1/production/0f59103a1f669cc3e42d6fe7dcae1015c1bd8802-3841x2160.png?auto=format&dpr=2&fit=max&q=75&w=1600)Our research compared auto-labeling and human labeling approaches using object detection tasks. Starting from identical unlabeled images, we generated labels both automatically (using foundation models like YOLO-World and Grounding DINO) and manually (human annotators). We then trained lightweight inference models from each label set without pre-trained weights or additional fine-tuning, allowing a fair and practical comparison of labeling costs, label quality, and downstream validation accuracy (measured via mAP). While AI-assisted annotation continues to improve, most [auto-labeling methods](https://arxiv.org/abs/2005.04757?utm_source=chatgpt.com) still rely on small human-labeled seed sets (typically 1–10%). Intrigued by this limitation, we sought to [evaluate](https://voxel51.com/blog/tag/evaluation) how effectively auto-labeling systems could perform _without_ _any initial human labels_. Zero. Zip. Zilch. Through this investigation, we also gained critical insights into setup and configuration that unlock auto-labeling performance. We started by measuring F1 scores, which provide a direct assessment of auto-label quality by balancing precision (how accurate labels are) and recall (how many true objects are correctly labeled). High F1 scores indicate auto-labels closely approximate human annotation accuracy. ![](https://cdn.sanity.io/images/h6toihm1/production/ea146b5d3de2bddd723f7a4bcc4b61ac61ff9a0f-1198x425.png?auto=format&dpr=2&fit=max&q=75&w=1198) On simpler datasets like VOC, YOLO-World achieved an impressive F1 score of 0.785, meaning it produced labels nearly as accurate as humans for straightforward object categories. However, performance decreased with dataset complexity: on COCO, the top models achieved approximately 0.640, and on the highly challenging LVIS dataset, scores dropped to 0.215, underscoring the model’s difficulty in accurately labeling rare classes. For specialized or complex classes, we recommend a hybrid approach combining Verified Auto Labeling with [targeted human annotation](https://docs.voxel51.com/user_guide/annotation.html). Given the efficiency of Verified Auto Labeling, this hybrid method still delivers substantial cost savings. We can also integrate proprietary models into Verified Auto Labeling to further improve accuracy on specialized datasets. ## Verified Auto Labeling delivers up to 95% model performance on downstream inference Evaluations based purely on metrics like precision and recall tell only part of the story. To measure the effectiveness of Verified Auto Labeling, we conducted a more practical test: we trained lightweight models (the kind you'd actually deploy on edge devices) directly from the auto labels, without using any pre-trained weights or human-labeled data. This approach allowed us to see whether auto-labels alone could produce high-performing models in real-world scenarios. Using mean Average Precision (mAP), a key real-world metric for object detection accuracy, we found that models trained solely on auto-labels performed just as well—and sometimes even better—than models trained on traditional human labels. ![](https://cdn.sanity.io/images/h6toihm1/production/5626dacc07b8818ef0b71ad808719c79ad06bb28-901x517.png?auto=format&dpr=2&fit=max&q=75&w=901) On VOC, auto-labeled models achieved mAP50 scores of 0.768, closely matching the 0.817 achieved with human-labeled data. On COCO, auto-labeled models reached mAP50 of 0.538 compared to 0.588 for human-labeled counterparts, demonstrating competitive real-world performance. Interestingly, in certain cases—such as detecting rare classes in COCO or VOC—auto-label-trained models occasionally outperformed those trained on human labels. This may occur because foundation models, trained on massive datasets, can generalize better across diverse objects or more consistently label challenging edge cases. In contrast human annotators might occasionally mislabel or overlook subtle object instances, particularly when working at scale – as illustrated by our donut example below. ![An example illustrating how auto-labeling can sometimes outperform human annotations. Here, the human annotator incorrectly labeled the donut count (left), while the auto-labeling model (right) more consistently identified and accurately counted the donuts. This demonstrates the potential advantage of foundation models in labeling repetitive or subtle object instances, which human annotators may overlook or mislabel, especially at scale.](https://cdn.sanity.io/images/h6toihm1/production/aa5c429eefc5a469b79e6a9e6dfb069e4b97e941-2561x1121.png?auto=format&dpr=2&fit=max&q=75&w=1600)An example illustrating how auto-labeling can sometimes outperform human annotations. Here, the human annotator incorrectly labeled the donut count (left), while the auto-labeling model (right) more consistently identified and accurately counted the donuts. This demonstrates the potential advantage of foundation models in labeling repetitive or subtle object instances, which human annotators may overlook or mislabel, especially at scale. Performance was notably weaker on highly complex or specialized datasets such as LVIS and BDD, which contain numerous nuanced or domain-specific classes. On LVIS, for instance, auto-label-trained models yielded very low mAP scores (less than 0.10), highlighting the significant challenges foundation models face when handling rare, ambiguous, or highly specialized object definitions. Foundation models often perform poorly in these scenarios because they're not specifically trained to distinguish extremely rare or specialized classes. ![](https://cdn.sanity.io/images/h6toihm1/production/814f49635b6e6af1fcf411018526f574d32cc2ba-918x582.png?auto=format&dpr=2&fit=max&q=75&w=918) These results indicate that while **auto-labels** **achieve about 90–95% of the performance of human labeling** in many practical scenarios, careful consideration of dataset complexity and class definitions remains essential. For specialized or particularly challenging categories, teams should adopt hybrid annotation strategies, combining auto-labeling’s scalability with targeted human expertise. We’ll have more on that topic soon. > "With Verified Auto Labeling, teams can bootstrap an entire detection dataset with no human-provided seed labels and train edge-friendly detectors that nearly match fully human-supervised results. All at six orders of magnitude lower cost." –Dr. Jason Corso, Chief Science Officer at Voxel51 ## Verified Auto Labeling reduces annotation costs by 100,000x While previous research qualitatively claimed auto-labeling reduces annotation costs, our study provides concrete figures: - Labeling 3.4 million objects on a single [NVIDIA L40S GPU](https://www.nvidia.com/en-us/data-center/l40s/) costs $1.18 and took just over an hour. - Manually labeling the same dataset via AWS SageMaker, which has among the least expensive annotation costs, would cost roughly $124,092 and take nearly 7,000 hours. ![Labeling Cost Comparison between human annotation services and auto-labeling methods across four popular datasets. The table highlights the significant cost and time reductions achieved using auto-labeling, as detailed in Auto-Labeling Data for Object Detection.](https://cdn.sanity.io/images/h6toihm1/production/663531163ca76e28d726514f89f88e0d1acf35c5-1215x430.png?auto=format&dpr=2&fit=max&q=75&w=1215)Labeling Cost Comparison between human annotation services and auto-labeling methods across four popular datasets. The table highlights the significant cost and time reductions achieved using auto-labeling, as detailed in Auto-Labeling Data for Object Detection. **Verified Auto-Labeling is 100,000× cheaper and 5,000× faster than traditional annotation.** These dramatic savings fundamentally alter the economics of bringing computer vision to production, freeing budget for [quality assurance](https://voxel51.com/blog/build-better-visual-ai-datasets-with-the-fiftyone-data-quality-workflow), [edge-case analysis](https://docs.voxel51.com/tutorials/small_object_detection.html#Identifying-Edge-Cases), and [strategic dataset expansion](https://voxel51.com/blog/data-augmentation-is-still-data-curation). ## Clean labels aren’t always better: how confidence thresholds impact model performance Confidence thresholds determine how sure a model must be before accepting a prediction as correct. In auto-labeling, each detected object receives a confidence score (0–1), reflecting how certain the model is about the detection. Practitioners typically set thresholds (e.g., 0.5) to filter lower-confidence predictions, assuming higher thresholds yield cleaner labels and better models. [Choosing the right confidence threshold](https://voxel51.com/blog/finding-the-optimal-confidence-threshold) for auto-labeling seems straightforward: higher confidence should mean cleaner labels and better downstream results. However, our benchmarks revealed a surprising insight: high-confidence labels (0.8–0.9), while appearing cleaner, consistently harmed downstream performance due to reduced recall. Optimal downstream performance (mAP) occurred at moderate confidence thresholds (0.2–0.5), balancing precision and recall effectively. However, although the best performance for all models and datasets tested was within this range, the actual best threshold varies across the study. ![Label precision climbs monotonically with threshold, but downstream mAP peaks at the mid-range (0.2-0.5). High-confidence labels (0.8-0.9) look “clean” but starve the model of recall, degrading performance by up to 8 mAP—even though F1 on the label set is highest. The study quantifies this across three detectors and four datasets. ](https://cdn.sanity.io/images/h6toihm1/production/9110b339d18c164d75c6e5349c06513fe57bcb51-1517x380.png?auto=format&dpr=2&fit=max&q=75&w=1517)Label precision climbs monotonically with threshold, but downstream mAP peaks at the mid-range (0.2-0.5). High-confidence labels (0.8-0.9) look “clean” but starve the model of recall, degrading performance by up to 8 mAP—even though F1 on the label set is highest. The study quantifies this across three detectors and four datasets. Understanding this balance enables better tuning of auto-labeling pipelines, prioritizing overall model effectiveness over superficial label cleanliness. ## Engineering tradeoffs at scale: why model selection matters in auto-labeling Performance and practical usability vary significantly among foundation models, making careful selection critical for auto-labeling workflows. Our experiments revealed substantial trade-offs beyond just accuracy: - While YOLO-World rapidly labeled large-scale datasets in minutes (~3 min for VOC), Grounding DINO was significantly slower (~38 min for VOC), due to computational constraints in handling complex text prompts. - Models like Grounding DINO encountered memory limitations with datasets containing verbose class descriptions (e.g., LVIS), requiring specialized adaptations and increasing labeling time dramatically. Understanding these real-world differences matters because model choice directly impacts operational efficiency, scalability, and overall deployment costs—critical considerations for any AI project. ## What is the best auto-labeling tool? Not all auto-labeling tools are created equal. As our research demonstrates, achieving optimal results requires careful selection of foundation models, confidence thresholds, and other parameters tailored to your use case. Developed by our world-class ML team, FiftyOne’s **Verified Auto Labeling** builds directly on this research, integrating automated labeling with streamlined QA workflows. Traditional auto-labeling solutions typically produce raw, noisy outputs that require extensive human cleanup. While initial costs may seem low, [hidden costs quickly pile up](https://voxel51.com/blog/how-to-tame-your-data-dragon/) — due to intensive manual QA, multiple review cycles, and costly re-annotations. Demo of the Verified Auto Labeling workflow within FiftyOne Verified Auto Labeling uses confidence scoring to automatically highlight labels that are most likely to need human attention, helping annotators efficiently prioritize their QA efforts. Annotators can still review any of the labels generated—even those that the system considers lower priority—providing flexibility to further refine labels. This approach streamlines the annotation workflow, reduces overall annotation costs, and enhances dataset quality and downstream model performance. - **One‑click QA**: Accept or reject auto‑labels instantly with [RER‑backed confidence scoring](https://voxel51.com/blog/understanding-dataset-difficulty-with-class-wise-autoencoders/), drastically cutting down unnecessary human oversight. - **Intelligent ranking for human review**: Direct human effort to the exact samples that most impact model quality. - **Strategic data selection**: Leverage FiftyOne’s powerful data curation capabilities to identify samples most likely to boost model performance — and reduce overall annotation volume and costs. - **Difficulty scoring**: Quantify labeling uncertainty to highlight exactly where human input adds the greatest value. Verified Auto Labeling is currently in beta, rolling out to existing FiftyOne Enterprise customers. Join our [upcoming workshop on June 24](https://voxel51.com/events/verified-auto-labeling-smarter-annotation-at-scale-june-24-2025) to learn more about Verified Auto Labeling. Alternatively, add yourself to our [beta waitlist here](https://voxel51.com/annotation#waitlist). ### **Cite this post** To cite this post in your research: _[Brent Griffin](https://voxel51.com/author/brent), [Jacob Sela](https://voxel51.com/author/jacob-sela), [Manushree Gangwar](https://voxel51.com/author/manushree), [Jason Corso](https://voxel51.com/author/jason). (June 4, 2025). Zero-shot auto-labeling rivals human performance. https://arxiv.org/abs/2506.02359_ [ML research](https://voxel51.com/blog/tag/ml-research) [computer vision research](https://voxel51.com/blog/tag/computer-vision-research) ![](https://cdn.sanity.io/images/h6toihm1/production/768656903b888850cdc5f83e35d55a29723f61e1-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Brent Griffin Bio ![](https://cdn.sanity.io/images/h6toihm1/production/8d9142c33a86d87e5a663508f5b5316a1240642c-400x400.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Jacob Sela Bio ![](https://cdn.sanity.io/images/h6toihm1/production/7c1e360ecabc3eb2a9098fdf879536b1430cea4e-400x400.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Manushree Gangwar Bio ![](https://cdn.sanity.io/images/h6toihm1/production/ad9fb967c5455e0f763411fb81956767d7f26482-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Jason Corso Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/4f4b3ab874b2159c10b6ac360c5f72b3baf7ade4-2500x1406.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Understanding Dataset Difficulty with Class-Wise Autoencoders\\ \\ ML@Voxel51\\ \\ • \\ \\ Feb 4, 2025](https://voxel51.com/blog/understanding-dataset-difficulty-with-class-wise-autoencoders) [![](https://cdn.sanity.io/images/h6toihm1/production/c287dc1bbe28f20f90a4c7256fbd60c7d950707b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ On Leaky Datasets and a Clever Horse\\ \\ ML@Voxel51\\ \\ • \\ \\ Dec 10, 2024](https://voxel51.com/blog/on-leaky-datasets-and-a-clever-horse) [![](https://cdn.sanity.io/images/h6toihm1/production/f78a97baf9ac0dca186397ad635f08adf77dfd00-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Bias in Data: What Embeddings Reveal About Real vs Synthetic Data Distribution\\ \\ ML@Voxel51\\ \\ • \\ \\ Jan 7, 2025](https://voxel51.com/blog/bias-in-data-what-embeddings-reveal-about-real-vs-synthetic-data-distribution) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-161-lllmstxt|> ## Auto-Labeling for Object Detection [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Research Paper # Auto-Labeling Data for Object Detection How far can zero-shot auto labeling take us in the quest for labeled datasets and performant models? To find out, we rigorously benchmarked Verified Auto Labeling across popular computer vision datasets. Get key insights and practical best practices from our latest research, including: - **Real-world performance:** How auto-labeling can achieve up to 95% of human accuracy—at 100,000× lower cost. - **Foundation model comparisons:** See how YOLO-World, YOLOE, and Grounding DINO differ in accuracy, speed, and scalability. - **Choosing confidence thresholds**: Why a higher threshold isn’t always better - **Downstream model training:** When auto-labels outperform human annotations — and why. - **Pitfalls to avoid:** Where auto-labeling falls short, and how hybrid approaches with targeted human annotation can help. ![](https://cdn.sanity.io/images/h6toihm1/production/552aa52a004e7137aba09ab689db8bccbc12809d-3840x2160.jpg?auto=format&dpr=2&fit=max&q=75&w=1600) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-162-lllmstxt|> ## Visual AI for Aviation [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Visual AI in Aviation FiftyOne is mission control for visual AI data. It gives AI/ML teams unprecedented visibility into dataset and model strengths and weaknesses, turning visual AI development from guesswork into precision engineering. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/1b43bfb6fda172741ff9ee15bea4843274b05e52-628x395.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/ecb765ef85f7077a7839f5b3c8de97fb5e57f2f6-628x817.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/c077e7fbccca03c22963f255856ccfde728e20ea-628x813.png?auto=format&dpr=2&fit=max&q=75&w=314) ![](https://cdn.sanity.io/images/h6toihm1/production/a1a9a6abced0535288894fffd50ddb35db3f7d23-628x393.png?auto=format&dpr=2&fit=max&q=75&w=314) Benefits ## FiftyOne helps aviation AI take flight with better data and better models. Computer vision revolutionizes aviation through automated aircraft inspection, plane identification, drone operations, airport security, and more. FiftyOne enhances these applications by unlocking data insights to maximize model performance. 0% improve productivity 0% increase model accuracy features ## Supports all popular vision tasks and media types. Use FiftyOne to store metadata about your samples, like annotations and model predictions, so you can easily pinpoint and optimize scenarios of interest whenever needed. ![](https://cdn.sanity.io/images/h6toihm1/production/0870ccefc24a1487f124ffc9bf87701c5cf6abbe-2560x961.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,2560,961&w=1280) - Classification - Detection - Segmentation - Polygons and polylines - Keypoints - Pointclouds - Heatmaps - Geolocation - Embeddings - Multiview datasets - Images, videos, and 3D data ## Data eats models for lunch Talk to our computer vision experts to start building better datasets and models. [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ac0775f29416480c0d8115ac92f9088eaab372ab-3024x961.png?auto=format&dpr=2&fit=max&q=75&w=1512) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-163-lllmstxt|> ## Upcoming Voxel51 Events [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/849de3e6ac8096687759c612a76c564e100463b1-3024x640.png?auto=format&dpr=2&fit=max&q=75&w=1512) # Voxel51 Events Join Voxel51 and the FiftyOne community at these events to learn all about computer vision, machine learning, and AI. Format Virtual In-person Region EMEA Americas Category Webinars & Workshops Meetups Industry Manufacturing Autonomous Vehicles [![](https://cdn.sanity.io/images/h6toihm1/production/03ebd9c87f7cc01872aaf1d562b1d3cdb5d10754-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ From Research to Reality: Building GUI Agents That Actually Work - August 22, 2025\\ \\ Welcome to the Visual Agents Workshop Series, your virtual pass to learn about visual agents - how they work, how to dev...\\ \\ Aug 22, 2025](https://voxel51.com/events/from-research-to-reality-building-gui-agents-that-actually-work-august-22-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/fef3495c2bcb5a92c5df355e49f9bc24b98f5d5d-2880x1620.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exposing Your Data's Blind Spots: Scenario Mining for Safer AV\\ \\ Stress test AV models by mining real-world edge cases using multimodal data and data-centric tools like FiftyOne.\\ \\ Aug 27, 2025](https://voxel51.com/events/exposing-your-datas-blind-spots-scenario-mining-for-safer-av) [![](https://cdn.sanity.io/images/h6toihm1/production/daf11585925107f46b2c82877eb862444d7058fd-1200x630.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Scaling Computer Vision AI in the Enterprise - 28 August, 2025\\ \\ Online -> Register for the Event\\ \\ Key Takeaways:\\ \\ Understand computer vision workflows and use cases across multiple ind...\\ \\ Aug 28, 2025](https://www.datacamp.com/webinars/scaling-computer-vision-ai-in-the-enterprise) [![](https://cdn.sanity.io/images/h6toihm1/production/e9dd901fc82b4b27b0342db4fe50e1380153c507-960x541.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ AI, ML and Computer Vision Meetup - Aug 28, 2025\\ \\ Hear talks from experts on the latest topics in AI, ML and Computer Vision!\\ \\ Aug 28, 2025](https://voxel51.com/events/ai-ml-and-computer-vision-meetup-aug-28-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/2bc63f0cc384ef11ad9a6fbf2b11217505725ec2-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ From Research to Reality: Building GUI Agents That Actually Work - August 29, 2025\\ \\ Welcome to the Visual Agents Workshop Series, your virtual pass to learn about visual agents - how they work, how to dev...\\ \\ Aug 29, 2025](https://voxel51.com/events/from-research-to-reality-building-gui-agents-that-actually-work-august-29-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/2b278ef1a28ee02fb9ccf8b8d85df092958807c3-1600x840.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Virtual How Porsche Uses Auto-Labeling to Supercharge AV Development - 4 September, 2025\\ \\ Online -> Register for the Event\\ \\ Join us for an exclusive webinar showcasing how Porsche is advancing its autonomous v...\\ \\ Sep 4, 2025](https://events.databricks.com/FY260904-WB-EngineeringRD/registration?scid=701Vp00000U6EaCIAV&utm_medium=Partner&utm_source=n/a) [![](https://cdn.sanity.io/images/h6toihm1/production/e3b40f53b685e10d34816d355272c984eb12a0d1-960x540.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual AI in Manufacturing and Robotics - September 10, 2025\\ \\ Join us for the first in a series of virtual events to hear talks from experts on the latest developments at the interse...\\ \\ Sep 10, 2025](https://voxel51.com/events/visual-ai-in-manufacturing-september-10-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/f738f742abf67718f860efe462f9042b6639ec99-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual AI in Manufacturing and Robotics - September 11, 2025\\ \\ Join us for dat two in a series of virtual events to hear talks from experts on the latest developments at the intersect...\\ \\ Sep 11, 2025](https://voxel51.com/events/visual-ai-in-manufacturing-and-robotics-september-11-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/63d5925d12404c349aa01929b8e5802b4a794f47-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual AI in Manufacturing and Robotics - September 12, 2025\\ \\ Join us for a series of virtual events to hear talks from experts on the latest developments at the intersection of Visu...\\ \\ Sep 12, 2025](https://voxel51.com/events/visual-ai-in-manufacturing-and-robotics-september-12-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/0cd905823479352908c60eb3a07cb4230442a758-960x540.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Boston AI, ML and Computer Vision Meetup - September 25, 2025\\ \\ This hands-on workshop explores computer vision techniques for automotive damage detection using the CarDD dataset - the...\\ \\ Sep 25, 2025](https://voxel51.com/events/boston-ai-ml-and-computer-vision-meetup-workshop-september-25-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/5d8257912f602205a536d40e7c3890b68fe8c187-960x540.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Valencia AI, ML and Computer Vision Meetup - September 25, 2025\\ \\ Acompáñanos para escuchar charlas de expertos en IA, ML y Visión por Computadora.\\ \\ Abrimos puertas a las 4:30 PM , comen...\\ \\ Sep 25, 2025](https://voxel51.com/events/valencia-ai-ml-and-computer-vision-meetup-september-25-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/63e76cde1a45aed8d10a5bcdec5a34ceb15268d3-960x540.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Madrid AI, ML and Computer Vision Meetup - September 26, 2025\\ \\ Acompáñanos para escuchar charlas de expertos en IA, ML y Visión por Computadora.\\ \\ Abrimos puertas a las 18:15 y empezar...\\ \\ Sep 26, 2025](https://voxel51.com/events/madrid-ai-ml-and-computer-vision-meetup-september-26-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/286af5bdd8c96bc26657a4ed7234c5eb6565f793-960x540.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Getting Started with FiftyOne for Manufacturing Use Cases - Sept 30, 2025\\ \\ Sep 30, 2025](https://voxel51.com/events/getting-started-with-fiftyone-for-manufacturing-use-cases-sept-30-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/c5114c78b43a502f91389b79f56e15f0f09aecd7-960x540.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Women in AI - October 2, 2025\\ \\ Hear talks from experts on the latest topics in AI, ML, and computer vision on October 2.\\ \\ Oct 2, 2025](https://voxel51.com/events/women-in-ai-october-2-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/c25840d7b086070df8290f52a85856ba1fd0874d-2881x1620.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Visual AI in Manufacturing: How Multimodal Data Powers Adaptive Process Control\\ \\ Join this webinar to learn how multimodal visual AI drives defect detection and adaptive control in manufacturing.\\ \\ Oct 8, 2025](https://voxel51.com/events/visual-ai-in-manufacturing-how-multimodal-data-powers-adaptive-process-control) [![](https://cdn.sanity.io/images/h6toihm1/production/81c77dd40160a833db0cf8aeebc0596d56fc4fcb-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ AI, ML and Computer Vision Meetup en Español - October 23, 2025\\ \\ Join the Meetup to hear talks in Spanish from experts on cutting-edge topics across AI, ML, and computer vision.\\ \\ Oct 23, 2025](https://voxel51.com/events/ai-ml-and-computer-vision-meetup-en-espanol-october-23-2025) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-164-lllmstxt|> ## Data Annotation Whitepaper [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) Whitepaper # Your Data, Your Advantage **Your AI advantage starts with your data.** As evidenced by Meta's acquisition of Scale AI, outsourcing annotation now comes with serious strategic risks. This whitepaper explores why retaining control over your proprietary datasets is essential—and how model–powered auto-labeling lets you scale securely, without handing data to third parties. Inside, you’ll get benchmarks, cost comparisons, and practical workflows for bringing annotation in-house. Learn how leading teams are cutting labeling costs by up to 100,000× while improving security, flexibility, and compliance. If data is your edge, this is how you protect it. ![](https://cdn.sanity.io/images/h6toihm1/production/0056db2d46f0d1f7c6394d430f2109e3b9eaeb46-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=1600) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-165-lllmstxt|> ## Composed Image Retrieval [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Composed Image Retrieval at CVPR 2025 Jun 2, 2025 • 23 min read Article content In this article [Why it matters](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#72afc6d4d69d) [A primer on composed image retrieval](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#c20b3695a743) [Generative zero-shot composed image retrieval](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#f428447caf7a) [Missing target-relevant information prediction with world model for accurate zero-shot composed image retrieval](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#22e0c6e2e720) [Imagine and seek: Improving composed image retrieval with an imagined proxy](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#384398fc1e0f) [The future of retrieval is compositional](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#40f415afbbfd) In this article [Why it matters](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#72afc6d4d69d) [A primer on composed image retrieval](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#c20b3695a743) [Generative zero-shot composed image retrieval](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#f428447caf7a) [Missing target-relevant information prediction with world model for accurate zero-shot composed image retrieval](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#22e0c6e2e720) [Imagine and seek: Improving composed image retrieval with an imagined proxy](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#384398fc1e0f) [The future of retrieval is compositional](https://voxel51.com/blog/composed-image-retrieval-at-cvpr-2025#40f415afbbfd) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _The evolution of image search_ If you’ve ever struggled to find the perfect car online, thinking, “I want this exact model, but in blue,” you’ve encountered the limitations of traditional image searches. While standard search lets you look for “blue sedans” or find similar vehicles to a reference image, it doesn’t easily combine these approaches. Enter **Compositional Image Retrieval (CIR)**, one of the coolest (and useful) areas of Visual AI research showcased at CVPR 2025. ## Why it matters Beyond academic interest, CIR has transformative potential for e-commerce, creative applications, and everyday search experiences. Imagine finding products matching your preferences by saying “like this but more formal” or “this couch in a different fabric.” It bridges the gap between how humans naturally communicate their visual preferences and how search systems operate. In this blog post, I’ll highlight some of the interesting CIR research presented at CVPR 2025, showcasing how researchers are tackling these challenges and pushing the boundaries of what’s possible in visual search technology. - [Generative Zero-Shot Composed Image Retrieval](https://hal.cse.msu.edu/papers/cig-generative-zero-shot-composed-image-retrieval/) - [Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval](https://arxiv.org/abs/2503.17109) - [Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy](https://arxiv.org/abs/2411.16752) ## A primer on composed image retrieval Compositional Image Retrieval **operates at the intersection of vision and language**, allowing users to **search using a multimodal query**: a reference image combined with a text modification. The reference image provides the visual foundation, while the text specifies desired changes to particular attributes. In formal terms, CIR involves **three key elements**: - A **reference image** that establishes the visual starting point - A **modification text** that specifies the desired changes (e.g., “make it red”, “with a sunroof,” “in the sport trim”) - A **gallery of candidate images** from which the system retrieves results The task requires the system to understand what’s in the reference image and how the textual instruction should transform it. For example, if you show a sedan and ask for “the same model but as a convertible in red,” the system must: 1. **Parse** the visual content of the sedan (make, model, features, design elements, etc.) 2. **Identify** which attributes should change (body type, colour) and which should remain (make, model, other features) 3. **Apply** these specific modifications conceptually 4. **Retrieve** images that match this mental transformation Under the hood, CIR systems typically **learn an embedding function that maps the combination of reference image and modification text to the same vector space as potential target images**. The mathematical goal is to make this composed representation similar to the true target images and dissimilar to irrelevant ones. What makes CIR distinct from other retrieval approaches is its compositional nature. Instead of finding exact matches to a query, **it performs a semantic transformation first, then searches based on the result of that transformation**. This bridges the gap between how humans naturally communicate our visual preferences (“like this, but different in that specific way”) and how retrieval systems operate. ### Beyond traditional search In traditional retrieval, the system directly encodes text (“blue SUV”) or an image into a representation matched against a database. The query itself represents what you want to find. In CIR, **the system must first understand the reference image’s attributes, interpret how the text modification should alter these attributes, and then combine this information to create a new representation that doesn’t directly match either input**. This composed representation describes an image that may not even exist in the exact form imagined. CIR does not match the raw multimodal query directly to gallery images. It processes (or “semantically transforms”) the reference image and modification text into a unified representation that embodies the desired target state. It then **uses this transformed representation to retrieve images from the gallery that are semantically** closest to that desired state. ### Different embedding space operations Traditional retrieval systems typically operate through: - **Text-to-image matching**: Text and images are projected into a shared embedding space where semantically similar items cluster together - **Image-to-image matching**: Finding database images closest to a query image in feature space CIR requires more complex operations: - **Attribute disentanglement:** Breaking down the reference image into modifiable properties - **Selective feature modification:** Applying the text modification to only relevant dimensions of the image embedding - **Feature preservation:** Maintaining unmentioned attributes at their original values - **Compositional reasoning:** Understanding how multiple modifications might interact For instance, when requesting “the same car but with a panoramic roof and in metallic gray,” the system must understand these are independent modifications that can co-exist, rather than treating it as a single atomic change. These techniques go beyond traditional contrastive learning approaches in image-text retrieval systems like CLIP. ### Two approaches to composed image retrieval Composed Image Retrieval (CIR) research typically follows one of two distinct paths, each with its own philosophy about how these systems should learn and operate. The traditional approach, **Supervised CIR**, relies heavily on carefully annotated training data. These systems **learn from triplets consisting of a reference image, modification text, and the corresponding target image that satisfies the modification**. While this approach often yields impressive results, it comes with significant drawbacks. Creating these annotated datasets is extremely labour-intensive and expensive, inherently limiting the size and scope of what these models can learn. Supervised models excel within their training domains but may struggle with novel modifications or image types. In contrast, **Zero-shot CIR (ZS-CIR)** takes a fundamentally different approach by eliminating the need for task-specific annotated triplets. Instead, these systems **leverage the knowledge embedded in large vision-language models pre-trained on vast amounts of general image-text data**. Some ZS-CIR methods convert reference images into textual representations that can be combined with modification text. In contrast, others generate synthetic training examples or cleverly combine existing pre-trained models without additional training. The key advantage is that these systems can often generalize to new domains and modification types without requiring domain-specific annotations. What’s particularly exciting is how the gap between these approaches has begun to narrow. While supervised methods historically held the performance edge, recent advances in large-scale vision-language models have enabled some zero-shot approaches to match or even exceed traditional supervised methods. This evolution suggests a future where CIR systems can offer strong performance and the flexibility to work across diverse domains without requiring extensive manual annotation for each new application. ### Evaluation complexity The evaluation of traditional retrieval systems is relatively straightforward: given a query, is the ground truth image ranked highly in the results? CIR introduces multiple layers of complexity: - Multiple valid targets may exist for a single query - The degree of modification matters (how metallic should “metallic gray” be?) - Attribute preservation needs evaluation (did unmentioned attributes remain unchanged?) - The same modification applied to different reference images should produce consistent changes Unlike the complexity of CIR systems themselves, evaluation frameworks are relatively straightforward, focusing on retrieval effectiveness through key metrics: #### Primary metrics: - **Recall@K (R@K):** Measures how often the correct target appears in the top K results. Common reporting includes R@1, R@5, R@10, and R@50. - **Recall\_subset@K:** Addresses the “false negative” problem in CIRR by evaluating against visually similar image subsets, focusing on the model’s ability to distinguish subtle text-specified differences. - **Mean Average Precision@K (mAP@K):** Essential for CIRCO evaluation where multiple ground-truth images may satisfy a single query. #### Benchmark datasets: - **FashionIQ:** Fashion-focused with clothing attribute modifications - **CIRR:** Features image clustering to mitigate false negatives - **CIRCO:** First dataset with multiple ground-truth targets per query Current evaluation focuses exclusively on **retrieval outcomes rather than intermediate reasoning processes**. While CIR systems must handle complex operations internally (attribute understanding, transformation application, selective feature modification), **success is measured purely by whether the correct image(s) were retrieved**. The CIR evaluation challenge is addressing several inherent complexities: multiple valid targets may exist, modification degrees vary in subjective importance, unmentioned attributes should remain preserved, and similar modifications should produce consistent results across different reference images. ### Why it’s xhallenging (and interesting) The field has seen rapid growth since its emergence around 2019, with approaches ranging from supervised learning using annotated triplets to zero-shot methods leveraging large vision-language models. Several factors make this a rich research area: - **Multimodal Understanding:** Systems must integrate and align information from images and text, where each provides complementary signals. - **Selective Attribute Modification:** The model must identify which visual elements to change while preserving everything else. - **Data Complexity:** Training requires triplets (reference image, modification text, target image), which are much harder to collect than simple image-text pairs. - **Multiple Valid Interpretations:** Multiple target images might be equally valid for a single query, complicating training and evaluation. ## Generative zero-shot composed image retrieval The paper’s framework adds a **generative step** to the ZS-CIR pipeline. Creating a visual representation (the pseudo-target image) that attempts to capture the desired result of the composition provides additional visual information that helps bridge the “representation gap” between composed embeddings in the language space and target image embeddings in the image space. This CIG component is an effective **add-on** that can be integrated with existing CIR methods to boost performance. The paper introduces a novel approach to **Composed Image Retrieval (CIR)**, specifically focusing on improving **Zero-Shot CIR (ZS-CIR)**. ![](https://cdn.sanity.io/images/h6toihm1/production/9903d5214128554abe02198ad51baafc74ad4ad0-1055x550.webp?auto=format&dpr=2&fit=max&q=75&w=1055) **Important links:** - [Paper on arXiv](https://hal.cse.msu.edu/assets/pdfs/papers/2025-cvpr-cig-generative-zero-shot-composed-image-retrieval.pdf) - GitHub: Not yet released - [Project Page](https://hal.cse.msu.edu/papers/cig-generative-zero-shot-composed-image-retrieval/) - Model: Not yet released One key element discussed and utilized by the paper is **Textual Inversion**, which **maps image features into a semantic token embedding space**. Textual inversion essentially learns a “pseudo token” or word embedding that represents the visual content of a specific image. This mapping aims to create representations compatible with the text encoder of a pretrained vision-language model, such as CLIP. Once the reference image is transformed into textual pseudo-tokens, these tokens are combined with tokens from the modification text. This forms a “unified query” that consists entirely of text-like tokens. This unified query is then encoded using the text encoder of the VLP model, allowing the image’s information to be integrated into textual prompts or sentences, enabling multimodal composition. The method generates **pseudo-target images** that visually represent what the modified reference image should look like when modified by the delta caption. These generated images serve as additional visual information to enhance retrieval performance. ### Training process During the training phase of their **Composed Image Generation (CIG)** model, textual inversion is used to map the image latent embedding of a training image into the token embedding space. This pseudo-token embedding is combined with the image’s caption to compose a prompt embedding. This composed prompt serves as a textual condition for training a latent diffusion model. At a high level, the model training approach is: - **Self-supervised** training using only standard image-caption pairs, not requiring expensive CIR triplet datasets. - A pre-trained textual inversion network maps image embeddings into the token embedding space. - Composed prompt embeddings are constructed by combining pseudo-tokens with the image’s caption. - A latent diffusion model (Stable Diffusion variants) is fine-tuned to reconstruct original images using these composed prompts as conditioning. ### Inference workflow During inference, the model addresses the challenge of retrieving images based on a reference image and text modification. The key innovation is using a generative component to create a visual preview of what the target image should look like after applying the requested modifications. This pseudo-target image helps bridge the modality gap between language-space embeddings and image-space embeddings. After generation, the pseudo-target image is processed to extract complementary information that enhances the retrieval process. This approach effectively provides two perspectives on the query: the original text-based representation and a visually informed representation that better aligns with the target image space. At a high level, the model inference approach is: - Generate a pseudo-target image which visually represents how the modified image should look. - The pseudo-target image is mapped back to the token embedding space. - A second composed text embedding is created from the pseudo-target image and the delta caption. - The original and pseudo-target-based embeddings are combined with a weighting hyperparameter. - Images are retrieved by computing cosine similarity with the combined embedding. This process leverages generated visual representations to bridge the gap between query and target image spaces. ### Key insights from generative CIR research - **Generating pseudo-target images effectively bridges the representation gap** between language-space embeddings and image-space embeddings, providing visual information that better aligns with target images. - **Self-supervised training is sufficient:** Fine-tuning a diffusion model on image reconstruction using composed prompts induces the ability to generate useful pseudo-target images without requiring expensive triplet datasets. - **Textual embedding fusion is superior:** Combining the original composed embedding with the pseudo-target-derived embedding at the textual level yields better performance than image-level or token-level fusion. - **Visual detail preservation:** Unlike single-token representations, pseudo-target images maintain rich visual content from the reference while incorporating requested modifications, creating more effective retrieval queries. ## Missing target-relevant information prediction with world model for accurate zero-shot composed image retrieval A major challenge in ZS-CIR is **accurately modifying a reference image according to manipulation text, especially when the text specifies visual content that is** **_missing_** **from the reference image**. Existing methods typically map the reference image into a pseudo-token of CLIP’s language space, but struggle with this because they ignore the missing target content. That’s because the CLIP embedding is coarse-grained and loses the visual details needed for CIR tasks. ![Figure 2 from the paper](https://cdn.sanity.io/images/h6toihm1/production/b1beb1b9ea971ea0b14d4a21de181c18901a9e55-1257x439.webp?auto=format&dpr=2&fit=max&q=75&w=1257)Figure 2 from the paper **Important links:** - [Paper on arXiv](https://arxiv.org/abs/2503.17109) - [GitHub](https://github.com/Pter61/predicir) - Model: Not yet released This paper introduces PrediCIR ( **Predi** ct target image feature before retrieval for zero-shot **C** omposed **I** mage **R** etrieval), which explicitly predicts missing target visual content before performing image-to-word mapping. Rather than directly mapping existing features to pseudo-tokens, PrediCIR first predicts what visual elements are needed to fulfill the manipulation instruction, then adaptively combines this predicted content with the reference image’s existing features. It does this through three interconnected modules that work together during pre-training and inference: 1. World view generation 2. Target content predictor 3. Predictive cross-modal architecture ![Figure 1 from the paper](https://cdn.sanity.io/images/h6toihm1/production/83a1dc1182084349b69c4a05ba55517dee8c1913-507x574.webp?auto=format&dpr=2&fit=max&q=75&w=507)Figure 1 from the paper ### World view generation **World view generation** is a fundamental component and the **initial step** in the pre-training process of the PrediCIR model. The main goal of this module is to **construct source and target views along with corresponding actions** from existing image-caption pairs, _without_ requiring extra supervision. This module generates **pseudo triplets** in . #### How it works: - An **original image** from an image-caption pair is designated as the **target view**. - A **corrupted version** of that original image is created by **randomly cropping** certain visual content. This corrupted image serves as the **source view**. Random cropping is preferred over masking to align with the frozen CLIP model and preserve coherent regional context. The crop size and aspect ratios were analyzed for their influence on performance. - The **caption** associated with the original image is used as the **action**, representing the intent to transform the source view into the target view. The caption is embedded using the frozen CLIP language encoder to obtain an action embedding. #### Why this data structure? This specific structure simulates the ZS-CIR problem where a reference image (like the cropped image) needs to be modified according to a text instruction (the caption/action) to become a target image (the original image). It teaches the model to understand what content is “missing” in the source view relative to the target view, based on the action. ### Target content predictor The triplets generated by the World View Generation module are then used to **train the Target Content Predictor (TCP)** module, which functions as a **world model**. Acting as a world model predictor (similar to a JEPA framework), the TCP takes the latent features of the source view (the cropped image patches), the action embedding (derived from the caption via the CLIP language encoder), and mask tokens representing the locations of missing content in the target view as input. Guided by the action, the **TCP learns to predict the latent representation of the missing target visual content**. This prediction occurs in the latent space. The TCP’s output includes the predicted latent features for the missing content and enhanced latent features for the source content. These predicted missing features and enhanced source features from the TCP are then passed to the Predictive Cross-Modal Alignment module. ### Predictive cross-modal alignment The **Predictive Cross-Modal Alignment (PMA)** module **bridges the gap between the visual features predicted by the Target Content Predictor (TCP) and the CLIP language space** used for retrieval. It takes the features representing the predicted missing visual content and the enhanced source content from the TCP, along with the global source feature, and **combines them adaptively**. This combined representation is mapped into the word token space to create a **pseudo-word token**. The pseudo-word token represents the potential visual content of the target image, including the elements predicted by the TCP. This token is appended to a simple prompt like “a photo of S\*”. The PMA is trained using a **contrastive loss** that encourages embedding the prompt sentence to align closely with the actual global feature embedding of the target image in the CLIP vision-language space. This is the final step in the prediction-based image-to-word mapping that takes the predictor’s output and transforms it into a format (the pseudo-token) suitable for composing a query within the CLIP language space for Zero-Shot Composed Image Retrieval. ### Key lessons and takeaways for practitioners The PrediCIR paper strongly suggests that explicitly predicting missing content is a powerful paradigm for ZS-CIR. Practitioners should consider leveraging prediction-based world models, carefully generating training data that simulates missing information, adaptively fusing original and predicted features, and aligning these predictive outputs with established multimodal spaces like CLIP to improve performance and generalization ability in CIR tasks. Here are some of the key insights from the paper: 1. **Prediction is powerful:** Explicitly predicting the visual content needed to fulfill a manipulation instruction, especially when missing in the reference image, improves ZS-CIR performance. Especially for “missing content” manipulations like changing domains (e.g., to origami) or adding objects. 2. **World models for visual transformations:** Training a world model predictor using a JEPA-like framework on synthetic “world views” (source, action, target) generated from image-caption data learns visual prediction capabilities without needing extensive, hand-annotated ZS-CIR triplets. 3. **Data generation matters:** Randomly cropping images to create source views creates diverse scenarios of missing content and aligns well with the frozen CLIP architecture used. Trying to predict the _entire_ target image is less effective than predicting the missing _parts_, suggesting partial prediction is better for managing computation and avoiding overfitting. 4. **Improved pseudo-token quality:** The prediction process yields a higher-quality pseudo-token that better represents the intended target image’s potential content and fine-grained details, which are important for accurate retrieval. ## Imagine and seek: Improving composed image retrieval with an imagined proxy Traditional ZS-CIR methods often rely on projecting query images into the text feature space and combining them with query text features for retrieval. This approach can suffer from a natural gap between images and text, making it difficult to guarantee detailed alignment and often overlooking important semantic information present in the image but not explicitly captured by text features. This is particularly challenging with complex captions. Existing methods that leverage large language models (LLMs) to generate descriptions also tend to focus on the text side, neglecting the potential for direct imagination on the image side. **This paper introduces Imagined Proxy for CIR (IP-CIR)**, a **training-free method** that leverages the power of imagination from generative models to address these limitations. ![Figure 1 from the paper](https://cdn.sanity.io/images/h6toihm1/production/ac1e07dca2b7c7d3bbd06ac9afebec5eaf1591ec-1694x1370.webp?auto=format&dpr=2&fit=max&q=75&w=1600)Figure 1 from the paper **Important links:** - [Paper on arXiv](https://arxiv.org/abs/2411.16752) - GitHub: Not yet released - Model: Not yet released IP-CIR harnesses the power of **imagination from generative models** to create an **“imagined proxy” image** aligned with the query image and the relative text description. This proxy image aims to provide additional details like style, instance attributes, and spatial relationships that text-based retrieval might miss. This contrasts with methods that rely solely on text modifications or text-space projection. The method then carefully integrates this visual proxy information with the original query image and textual guidance through robust features and a balanced retrieval metric. The framework involves three main conceptual steps: 1. **Imagined retrieval proxy generation** 2. **Constructing robust proxy features** 3. **Balancing retrieval results** The innovation here is the direct use of **image generation** to create a visual “imagination” of the target image, which _should_ provide a rich, image-side feature representation to complement and enhance traditional text-based retrieval in ZS-CIR. ![Figure 2 from the paper](https://cdn.sanity.io/images/h6toihm1/production/f2a14132dcb3adf0b17241ccf66a019d1c5b297e-1314x488.webp?auto=format&dpr=2&fit=max&q=75&w=1314)Figure 2 from the paper ### Imagined retrieval proxy generation The first stage focuses on generating a visual representation of what the user is looking for — an “imagined proxy image” that combines elements from the reference photo and the text description. IP-CIR starts by analyzing the user’s query image using BLIP2, which automatically generates detailed captions describing what’s in the image. These **captions are then combined with the user’s text modifications and fed into Qwen1.5–32B**, a large language model that serves as the reasoning engine. Qwen analyzes this information and creates a detailed spatial layout for the target image. This layout includes precise descriptions of each object and its positions. It decides which elements should come from the original image versus the text description. This essentially creates a blueprint that says, “Keep the dog from the photo, but add a red hat as described in the text.” IP-CIR then uses controllable image generation techniques to create an actual proxy image following this layout. The generation process incorporates visual features from the original query image while adding new elements described in the text. The result is a concrete visual representation of the search target. ### Constructing robust proxy features Raw proxy images sometimes emphasize the wrong details or miss subtle textual requirements. This stage creates a more reliable search representation by combining multiple information sources. Qwen generates detailed captions describing what the ideal target image should contain after applying all the requested modifications. IP-CIR compares these target descriptions with the original BLIP2 captions to identify exactly what should change. This creates a “semantic perturbation” that captures the direction and nature of the requested modifications. The final search feature combines three components: - Features from the generated proxy image (the visual imagination) - Features from the original query image (preserving important context) - The semantic perturbation (ensuring text requirements are met) This multi-source approach compensates for potential weaknesses in any individual component while preserving the most important information from visual and textual inputs. ### Balancing retrieval results The final stage addresses how to effectively combine results from traditional text-based search with the new image-based proxy approach. IP-CIR now has two ways to search: using traditional text-based methods and using the new proxy-based visual approach. Each method produces similarity scores, but simply averaging them can lead to poor results where images excel in one area but fail in another. IP-CIR uses “balanced similarity” by multiplying the text-based score with the proxy-based score. This multiplication is key — it ensures that only images performing well in both approaches receive high, balanced scores. An image that completely fails the text or visual test will have a very low balanced score. The final search ranking combines the original text-based similarity with this balanced similarity using a weighted average. A lambda parameter controls how much weight to give each component, allowing the system to adapt to different types of content and datasets. ### Why this approach works This three-stage process creates a search system that truly understands complex visual-textual queries. By generating actual proxy images, the system can capture spatial relationships and visual details that pure text descriptions might miss. Combining multiple information sources in the feature construction makes it more robust against any single component’s limitations. And by carefully balancing different similarity measures, it ensures that results satisfy both visual and textual requirements rather than excelling in just one area. The result is a search system that can handle requests like “find me a living room like this one, but with a blue couch instead of brown” — understanding both the visual context of the reference image and the specific textual modifications requested. ### Key lessons and takeaways for practitioners 1. **Text-only retrieval has limitations:** Relying solely on text features for CIR overlooks important semantic information and detailed visual alignment due to the inherent gap between images and text. Text features can be coarse and easily confusable in fine-grained, complex imagination scenes. 2. **Visual imagination enhances retrieval:** Creating an “imagined proxy” provides valuable additional information, such as style, instance attributes, and spatial relationships, that are often difficult for text-based methods to capture accurately, especially with complex captions. 3. **Harnessing generative models is feasible:** Controllable generative models make it possible to create high-quality images that align with both query images and relative captions. 4. **LLMs for reasoning and layout:** LLMs can understand the complex relationship between the query image (represented by automatically generated captions like those from BLIP2) and the relative text. They can reason out detailed spatial layouts for the imagined target image, including object descriptions, bounding box coordinates, and even determine which attributes should come from the original image versus the text. They can also help derive semantic features representing the direction of the textual edit. 5. **Raw proxy features aren’t enough; robust features are key:** Simply using the raw features of a generated proxy image for retrieval can introduce noise or cause the system to overemphasize irrelevant details, such as a background not specified in the text. A more effective approach is to construct a **robust proxy feature** by combining information from the proxy image, the original query image (to compensate for lost details), and a “semantic perturbation” feature derived from inferred target captions (to capture the edit direction and mitigate focus on irrelevant details). 6. **Balancing modalities improves accuracy:** Combining the retrieval similarities obtained from image-based proxy features and traditional text-based baseline methods is critical for achieving accurate results. A balanced metric, such as the one used in IP-CIR which involves multiplying the similarities, ensures that a high final score is only achieved if an image has reasonably high similarity in _both_ modalities. This prevents retrieving images that only match one aspect (image or text) but not the other. ## The future of retrieval is compositional The three approaches showcased at CVPR 2025 represent distinct yet complementary philosophies in advancing Composed Image Retrieval. Each tackles the fundamental challenge of bridging vision and language from different angles. **Generative CIR** creates visual previews of target images, excelling at complex transformations. **PrediCIR** analytically predicts missing content, proving powerful for domain transfers and object additions. **IP-CIR** combines multiple AI systems for holistic reasoning across visual and textual modalities. All three methods recognize a fundamental truth: traditional text-image matching isn’t sufficient for compositional queries. They require more sophisticated mechanisms to bridge the gap between how humans describe visual modifications and how systems process them. The shift toward zero-shot approaches signals progress toward truly generalizable Visual AI systems that understand and manipulate visual concepts without exhaustive training on every scenario. As these methods mature and potentially converge, we’re approaching search systems that don’t just find what exists, but understand what we imagine. The implications extend far beyond academic research. From e-commerce platforms enabling intuitive product discovery to creative tools that understand artistic intent, CIR is positioning itself to transform how we interact with visual information. The question isn’t whether these systems will become mainstream, but how quickly they’ll integrate into the tools we use every day. The future of search isn’t just retrieval — it’s visual reasoning. [similarity search](https://voxel51.com/blog/tag/similarity-search) [vector search](https://voxel51.com/blog/tag/vector-search) [generative AI](https://voxel51.com/blog/tag/generative-ai) [text-to-image](https://voxel51.com/blog/tag/text-to-image) [embeddings](https://voxel51.com/blog/tag/embeddings) [zero-shot](https://voxel51.com/blog/tag/zero-shot) [benchmark](https://voxel51.com/blog/tag/benchmark) ![](https://cdn.sanity.io/images/h6toihm1/production/a41a0477c7a98264f600772e9568607d070eea59-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Sahota Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/663dd6a3e6f3e57a03425f932440b5d242133451-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Search and curate video data with FiftyOne, Twelve Labs, and Databricks Vector Search\\ \\ Product & News, Integrations\\ \\ • \\ \\ Jun 5, 2025](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs) [![](https://cdn.sanity.io/images/h6toihm1/production/7ec61c80f387b16f24b4b2fe33f804264a38b486-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Rethinking How We Evaluate Multimodal AI\\ \\ Event Recaps\\ \\ • \\ \\ Jun 12, 2025](https://voxel51.com/blog/rethinking-how-we-evaluate-multimodal-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/1f6c282282c17a88941d51776c229bde5c4849b3-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ VGGT is a Pure Neural Approach to 3D Vision\\ \\ Event Recaps\\ \\ • \\ \\ Jun 25, 2025](https://voxel51.com/blog/vggt-is-a-pure-neural-approach-to-3d-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-166-lllmstxt|> ## Ancera Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/d321cc301ad25d9c4e6890dd2d1eee225ec72bb9-2432x2432.png?auto=format&dpr=2&fit=max&q=75&w=300) [Case Studies](https://voxel51.com/customers) Ancera Ancera speeds up pathogen detection model development, regains months of engineering time May 3, 2025 Ancera streamlined manual workflows and enhanced model performance by adopting FiftyOne to identify pathogen risk amongst poultry farms. Model performance improved by over 7% Reviewed and relabeled detections across hundreds of high-resolution images 20,000 Feedback loop reduced from months to days Article content In this article [Challenge](https://voxel51.com/customers/ancera#29a0126a0d02) [Solution](https://voxel51.com/customers/ancera#37e24bcbf000) [Key Results](https://voxel51.com/customers/ancera#1064d62268e4) In this article [Challenge](https://voxel51.com/customers/ancera#29a0126a0d02) [Solution](https://voxel51.com/customers/ancera#37e24bcbf000) [Key Results](https://voxel51.com/customers/ancera#1064d62268e4) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Ancera](https://www.ancera.com/), a software analytics company for the poultry industry, focuses on improving live production and processing by monitoring and responding to pathogen risk, such as Salmonella and Coccidia, across the supply chain. Using high-resolution imaging systems and computer vision models, Ancera has helped poultry integrators save millions of dollars and gain real-time visibility into which farms face microbial threats and where interventions are needed to minimize downtime. ![](https://cdn.sanity.io/images/h6toihm1/production/0fa18ba79a0570362d9a851dd8fb80876865d021-1200x900.jpg?auto=format&dpr=2&fit=max&q=75&w=1200) > "The biggest benefit of using [FiftyOne](https://voxel51.com/) Enterprise has been the speed of development. What used to take weeks or even months can now be done in days, with fewer people and better results. It’s freed up our team to focus on what they do best while accelerating our computer vision pipeline." – Kemal Eren, Lead Computer Vision Engineer at Ancera ## Challenge ### Manual workflows and limited model visibility hindered Ancera’s operational scalability. Ancera’s platform collects and processes large volumes of poultry data from hatcheries, feed mills, breeder farms, growers, and processing units to detect pathogen risks using computer vision models. As Ancera’s imaging operations scaled, manual data and model workflows were slowing development, and the ML team faced several challenges: - **Slow model analysis and visibility:** Ancera’s ML engineers lacked an intuitive way to view, filter, and root-cause suboptimal model performance, often taking weeks of effort across multiple people to scroll through large volumes of data to identify and correct mislabeled data. - **Fragmented collaboration:** Non-technical stakeholders (microbiologists) couldn’t easily access or interpret the results of their lab tests, without engineering assistance, slowing iterations and feedback. As a result, deploying new models was taking over 3-4 months instead of a few weeks. ## Solution ### Ancera adopted FiftyOne Enterprise for their complex visual data and model analysis layer to accurately classify Eimeria species. Ancera’s data includes distinct data types such as high-resolution microscopy images, DNA sequencing results, and environmental metadata like flock health and growth rates from a variety of sources. For example, fecal samples from live birds or swabs from machinery in processing plants are translated into images. The variability in sample quality and background noise makes data curation quite complex. The Ancera team integrated FiftyOne Enterprise to streamline data curation by enabling the team to visually sift through the noise, isolate relevant objects, and maintain high-quality training and test datasets across diverse imaging conditions. They incorporated FiftyOne capabilities to find model performance issues using the ability to analyze false positives/negatives, confusion matrices, and model scores. ## Key Results ### Faster feedback loops and over 7% improvement in model performance. Using FiftyOne Enterprise, Ancera accelerated the development and performance of their computer vision model for accurately classifying Eimeria species, a pathogen that causes coccidiosis, impacting animal health, productivity, and overall farm profitability. - The team gained **instant visibility** into **model** **predictions** and **label** **quality**. Using the filtering and patch view capabilities, they were able to review and relabel 20,000 detections across hundreds of high-resolution images in just a few days, a task that previously took multiple weeks. - This improvement enabled retraining the model with higher-quality data, resulting in more than **7% improvement in performance** with a **noticeable** **boost** in F1 score, precision, and recall. - **Cross-team efficiency:** Ancera’s microbiologists could analyze the results of their lab results independently using a browser link to step through the predictions in FiftyOne, freeing the ML team from manual report generation and speeding up QA cycles. > " [FiftyOne](https://voxel51.com/) has helped us shrink our feedback loop from months to days, allowing us to catch issues we didn’t know existed. It’s a big part of our strategy to speed up model development across teams." – Kemal Eren, Lead Computer Vision Engineer at Ancera ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-167-lllmstxt|> ## Getting Started with FiftyOne [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # FiftyOneGetting Started Take the next step to an easier, simpler, and faster way to build visual AI. [Try in browser](https://try.fiftyone.ai/datasets) [Book a demo](https://voxel51.com/get-started#form) ![Visual AI Plant Embeddings](https://cdn.sanity.io/images/h6toihm1/production/2f0882e49ebe24975b416df8b6b09d9980b209ae-1258x1200.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=350&q=75&w=350) ![](https://cdn.sanity.io/images/h6toihm1/production/768506901322a2845afad1f2fa46befeafbd1f35-1258x1200.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=350&q=75&w=350) ## Get started with FiftyOne [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-168-lllmstxt|> ## Page Not Found 404 Page not found [Back to Home](https://voxel51.com/) <|firecrawl-page-169-lllmstxt|> ## Fyma Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/1f113782bbe98bcb928919a47a9f2bb5636402e2-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Fyma Fyma relies on FiftyOne to visualize CCTV datasets for object detection tasks Apr 18, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Fyma](https://www.fyma.ai/) is a UK-based company that automates human sight—something sensors are unable to do—to provide you with the power to make informed, data-driven decisions regarding your real estate portfolio and how it is being used across a wide variety of spaces, from office buildings to retail. > "At Fyma we do object detection based on CCTV cameras and have around 50K images in different datasets. We use FiftyOne to manage all of the datasets, annotating with the Label Studio integration, exporting to YOLOv5 from Ultralytics, importing predictions and then using FiftyOne Brain methods to visualize everything. We were able to find almost 600 errors in our dataset the first time we used FiftyOne!" – Madis Pukkonen, Software Engineer at Fyma ![](https://cdn.sanity.io/images/h6toihm1/production/69a2241818008e7f1e44fa4b0b0e452df168c9dc-2000x1333.jpg?auto=format&dpr=2&fit=max&q=75&w=1600) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-170-lllmstxt|> ## Finegrain Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/b224afdc229760342dd440fdcc31e50b5550d452-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Finegrain FiftyOne fuels Finegrain’s ML workflows and data curation Apr 10, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) In a world where over 3 billion photos are shared daily and 90% of shoppers make decisions based on visuals, [Finegrain](https://finegrain.ai/) is making professional-quality images accessible to everyone. While smartphone cameras have evolved, creating stunning photos traditionally demands time-consuming editing and technical expertise—but Finegrain’s platform changes that, beautifying the internet by making high-quality visual content effortless to create and share. > "Finegrain is working on beautifying the Internet. Our first product is a dead simple photo editor supercharged with AI featuring breakthrough visual AI tools developed in-house and used by 500k+ creators on Hugging Face. FiftyOne helps us curate visual data more efficiently and streamline our machine learning workflows." – Denis Brulé, Co-Founder and CEO at Finegrain ![](https://cdn.sanity.io/images/h6toihm1/production/e3fc3bee3cc9e3d9686d4d843cf53369da262445-2560x1441.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) Loading related posts... [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-171-lllmstxt|> ## Paris AI Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/d601ec2435d1a3218f73efcde251ede170e59860-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/d601ec2435d1a3218f73efcde251ede170e59860-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) In-person EMEA Meetups Paris AI, ML and Computer Vision Meetup - July 16 Jul 16, 2025 5:30-8:30 PM Paris Marriott Opera Ambassador Vendôme Meeting Room 16 Bd Haussmann Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/adf29b4949fefa465ee40777a87dfe910cab2770-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Julien Simon Arcee AI Bio ![](https://cdn.sanity.io/images/h6toihm1/production/ab547ef15e03315bbef63a0c915de5896f39865d-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Gabriel Trégoat Pruna AI Bio ![](https://cdn.sanity.io/images/h6toihm1/production/bb53d4a9879495d53752948d94386a820f3a9bff-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Louis Dupont Bridgy AI Bio ![](https://cdn.sanity.io/images/h6toihm1/production/95311d0b30f6accc28ae51b4263695fbce7fa77d-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Harpreet Sahota Voxel51 Bio About this event Hear talks from experts on cutting-edge topics in AI, ML, and computer vision on July 16 Schedule Building and working with Small Language Models ![](https://cdn.sanity.io/images/h6toihm1/production/adf29b4949fefa465ee40777a87dfe910cab2770-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Julien Simon Arcee AI Bio This session focuses on practical techniques for using small open-source language models (SLMs) in enterprise settings. We'll explore modern workflows for adapting SLMs with domain-specific pre-training, instruction fine-tuning, and alignment. Along the way, we will introduce and demonstrate open-source tools such as DistillKit, Spectrum, and MergeKit, which implement advanced techniques crucial for achieving task-specific accuracy while optimizing computational costs. We'll also discuss some of the models and solutions built by Arcee AI. Join us to learn how small, efficient, and adaptable models can transform your AI applications. Accelerating sustainable inference with Pruna AI ![](https://cdn.sanity.io/images/h6toihm1/production/ab547ef15e03315bbef63a0c915de5896f39865d-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Gabriel Trégoat Pruna AI Bio This talk explores how to make AI faster and more sustainable. We’ll look at the high costs and carbon impact of fine-tuning and self deploying models, and show how optimization techniques available in the Pruna library can reduce size and latency with little to no quality loss. What I Learned About Systematic AI Improvement ![](https://cdn.sanity.io/images/h6toihm1/production/bb53d4a9879495d53752948d94386a820f3a9bff-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Louis Dupont Bridgy AI Bio Most AI teams go through the same story: fast early progress, and then suddenly things slow down. The AI isn’t broken, but new changes don’t seem to help, and it’s not even clear how to tell if things are getting better. I’ve faced this plateau myself—both in my own work and while helping other teams. In this talk, I’ll share what I’ve learned about getting unstuck: how to build genuine confidence in your AI, what “trust” really means in practice, and practical steps to move from “it kind of works” to “this is actually improving.” My goal is to give you real-world ideas you can use when you hit the same wall. Visual Agents: What it takes to build an agent that can navigate GUIs like humans ![](https://cdn.sanity.io/images/h6toihm1/production/95311d0b30f6accc28ae51b4263695fbce7fa77d-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Harpreet Sahota Voxel51 Bio We’ll examine conceptual frameworks, potential applications, and future directions of technologies that can “see” and “act” with increasing independence. The discussion will touch on both current limitations and promising horizons in this evolving field. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-172-lllmstxt|> ## Verified Auto Labeling [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # FiftyOneVerified Auto Labeling Annotation doesn’t have to slow you down — or blow your budget. Get AI into production more quickly with Verified Auto Labeling. [Join the waitlist](https://voxel51.com/annotation#waitlist) ![Annotation bounding boxes of fish](https://cdn.sanity.io/images/h6toihm1/production/cd2d43fa103cf8d556459158d01a04abdfb47ff9-1600x1600.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1600,1600&w=350) ![](https://cdn.sanity.io/images/h6toihm1/production/7468290612338c48881e781a94a18dd22a570c37-1600x1600.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1600,1600&w=350) Automated Labeling ## Reduce annotation costs with Verified Auto Labeling Avoid the hidden costs associated with other auto-labeling tools. ### Get auto-labeling with built-in QA Verified Auto Labeling uses foundation models to automatically generate labels — and adds confidence scoring to prioritize the ones that require human review. It streamlines your annotation workflow, cutting QA, and annotation costs while maintaining near-human accuracy. [Learn more](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance) ![Verified Auto Labeling](https://cdn.sanity.io/images/h6toihm1/production/a33865d602e87c957f6292cad6bebeebf25294d6-1200x1200.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1200,1200&w=600) ML Research ## Auto-labeling rivals human performance The latest paper from our ML researchers, _Auto-Labeling Data for Object Detection_, benchmarks auto-labeling against human annotation. We reveal how foundation models can deliver labels at near-human accuracy, while reducing annotation costs by up to 100,000×. [Read the research](https://voxel51.com/whitepapers/auto-labeling-data-for-object-detection) ![](https://cdn.sanity.io/images/h6toihm1/production/587ebdf76015a294f53ed163598b17f566e39d6f-1280x640.png?auto=format&dpr=2&fit=max&q=75&w=640) ![](https://cdn.sanity.io/images/h6toihm1/production/a7a1085305de96e5ec83621e54a23dfed5cb725c-2560x640.png?auto=format&dpr=2&fit=max&q=75&w=1280) Annotation Calculator ## Estimate your savings with our auto-labeling calculator [Access underlying calculations](https://voxel51.com/blog/how-we-built-annotation-savings-estimator) ![](https://cdn.sanity.io/images/h6toihm1/production/391d2e18dd82dab450a36f2a41c03e134dbcc11a-3840x1560.png?auto=format&dpr=2&fit=max&q=75&w=1280) Number of images to label Task Type Classification Detection Segmentation Data curation for annotation ## Label smarter Optimize your annotation budget by ensuring each labeled data point maximizes model performance. ### Label the right data, not just more data Use FiftyOne to identify and prioritize the most valuable data points for annotation. [Explore embeddings for annotation](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities) ![Computer vision embeddings](https://cdn.sanity.io/images/h6toihm1/production/eda6b1f41ad09ea94312fe8902f8837fa89cb74a-1200x1200.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1200,1200&w=600) ### Identify annotation mistakes Surface annotation errors by highlighting samples with high false positives, unexpected predictions, detection errors, or inconsistencies in labels. [Learn more](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023) ![](https://cdn.sanity.io/images/h6toihm1/production/21fa0491c438b82f8954f6eb9cfea22a79162a63-1300x1300.png?auto=format&dpr=2&fit=max&q=75&rect=0,0,1300,1300&w=600) ## Questions? We have answers. ### What is auto-labeling? ### How does auto-labeling reduce annotation costs? ### How does Verified Auto Labeling improve dataset quality compared to traditional auto-labeling tools? ### How do I choose the best auto-labeling tool for my computer vision project? ### What are best practices for optimizing human review in data annotation workflows? ## Get on the waitlist for Verified Auto Labeling [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-173-lllmstxt|> ## Segmentation Maps vs Bounding Boxes [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Learn](https://voxel51.com/blog/category/learn) Why Are Image Segmentation Maps Superior to Bounding Boxes? Feb 26, 2025 • 12 min read Article content In this article [The story behind image segmentation](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#8d57b95b7870) [Understanding segmentation maps](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#2bccca0bf714) [Harnessing FiftyOne for segmentation map visualization](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#b8ec651ff105) [Creating high-quality segmentation maps](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#f7251da01b80) [Interpreting segmentation maps](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#d1ece3775a2c) [Real-world applications of segmentation maps](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#f302eea3adcf) [The future of segmentation maps: Transformers](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#614f9e84f693) [Key takeaways](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#2bdf1bdc0ee4) In this article [The story behind image segmentation](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#8d57b95b7870) [Understanding segmentation maps](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#2bccca0bf714) [Harnessing FiftyOne for segmentation map visualization](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#b8ec651ff105) [Creating high-quality segmentation maps](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#f7251da01b80) [Interpreting segmentation maps](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#d1ece3775a2c) [Real-world applications of segmentation maps](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#f302eea3adcf) [The future of segmentation maps: Transformers](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#614f9e84f693) [Key takeaways](https://voxel51.com/blog/why-are-image-segmentation-maps-superior-to-bounding-boxes#2bdf1bdc0ee4) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) A **segmentation map** is a visual depiction caused by dividing an image into many parts, or segments, to get a pixel-level understanding. Segmenting is crucial our daily lives, from helping children recognize sounds to helping businesses understand their customers' preferences. It helps us make sense of chaos in various contexts of sounds, markets, or pixels. The evolution of image segmentation began in the 1950s-1970s with some early milestones and, later, in 2012, with biological neural networks. It has significantly advanced medical imaging (see ISBI'12 [Challenge](https://people.idsia.ch/~juergen/2012deeplearningwinsbraincontest.html)) and formed the backbone of modern computer vision. Today, segmentation provides a detailed lens for visual analysis. Paper: U-Net - [ArXiv](https://arxiv.org/pdf/1505.04597v1.pdf) In computer vision, traditional object detection methods can struggle with nuanced boundaries. Segmentation maps address this by identifying “what” an object is and “where” it begins and ends, providing pixel-level granularity for better scene interpretation. A segmentation map is a visual representation of an image where each pixel is labeled a class. In this article, we’ll explore the power of segmentation maps in autonomous driving with the help of FiftyOne. We will explore segmentation tools (using [FiftyOne](https://docs.voxel51.com/getting_started/install.html) and the [KITTI](https://docs.voxel51.com/dataset_zoo/datasets.html#kitti) dataset) as well provide practical code snippets in-line here and in the linked [notebook](https://github.com/vanshi-jain/Semantic-Maps). Here’s what you’ll learn: - Understand the basics of segmentation maps, their types, and practical segmentation scenarios. - Use FiftyOne’s tools to visualize segmentation saps and gain insights from segmented scenes. - Learn data annotation and augmentation strategies to generate high-quality, precise segmentation maps with FiftyOne. ## The story behind image segmentation Image segmentation has progressed steadily from its original deterministic methods. Earlier segmentation methods included classical techniques such as _Support Vector Machines_ (SVMs), k-means clustering, edge detection, and graph-based approaches. These foundational methods paved the way for today's practices. The year 2015 marked a turning point with the rise of deep learning-based techniques like _Fully Convolutional Networks_ ( [FCNs](https://arxiv.org/abs/1411.4038)). The success of these models relies on their ability to learn intricate feature representations through feature maps. **Feature maps** are essentially sets of arrays generated by layers within these models. Each array highlights specific features detected in the input image, such as edges, textures, or shapes. These feature maps collectively form a rich **feature representation**, which is a transformed and more abstract encoding of the original image, optimized for the segmentation task. Later, the introduction of **_U-Net_** architecture revolutionized the field with its encoder-decoder architecture and efficient upsampling. This made segmentation more robust and suitable for real-time applications. U-Net’s architecture is particularly effective in segmentation because its encoder path progressively extracts feature maps at different scales. These feature maps are then upsampled and combined in the decoder path to generate precise segmentation masks. Fast-forward to 2023, and the **Segment Anything Model( [SAM](https://arxiv.org/abs/2304.02643v1))** has emerged as a pinnacle segmentation model. SAM uses vision transformers to enable zero-shot segmentation across diverse datasets, and it sets a new standard for segmentation capabilities. By learning highly generalizable feature maps, SAM can effectively segment objects in diverse and unseen images. Let's now briefly discuss segmentation maps before exploring SAM on the Virtual KITTI dataset. ## Understanding segmentation maps Segmentation maps label each pixel in an image based on objects or categories, capturing more precise boundaries and offering clarity beyond simpler object detection techniques. - **Precise boundaries:** Segmentation maps accurately represent contours and edges, making them essential for surgical planning or pathfinding in autonomous systems. - **Detailed overlap analysis:** In cluttered scenes, segmentation maps offer a detailed analysis of individual objects in satellite imagery. - **Enhanced interpretability:** By providing a pixel-level view of objects, segmentation maps make interpretation more transparent and crucial in sensitive fields such as healthcare and defense. ### Types of image segmentation: Choosing the right segmentation technique Segmentation techniques vary depending on the complexity and the choice depends on the details required for a task. Here are the key approaches: - **Semantic segmentation** provides a pixel-level view of objects, segmentation maps make interpretation more transparent and crucial in sensitive fields such as healthcare and defense. PapersWithCode [Semantic Segmentation](https://paperswithcode.com/task/semantic-segmentation) - **Instance segmentation** differentiates between individual instances of the same class (e.g., multiple sheep in a scene). It is necessary when object-level granularity is crucial. PapersWithCode [Instance Segmentation](https://paperswithcode.com/task/instance-segmentation) - **Panoptic segmentation** combines semantic and instance segmentation strengths, and provides a holistic view of the scene. Every pixel is labeled with unique IDs for countable objects. This works well in complex scenes with a mix of class-level and instance-level understanding (e.g., identifying each person as a unique entity with the same class as “person”). PapersWithCode [Panoptic Segmentation](https://paperswithcode.com/task/panoptic-segmentation) To create and visualize such segmentation maps, we need to prepare a high-quality annotated dataset and train and evaluate a model for our custom use case. Here, we are developing a SAM model fine-tuned on the KITTI dataset using FiftyOne. ## Harnessing FiftyOne for segmentation map visualization FiftyOne is a tool that enables you to visualize datasets and interpret model performance faster and more effectively (for example, curating, visualizing, and analyzing segmentation maps). The key value proposition is that better data ultimately leads to better performing models. For image segmentation specifically, FiftyOne offers a few key capabilities: - **Streamlined workflows:** FiftyOne provides an intuitive UI for labeling, dataset exploration, and model evaluation. - **Powerful visualization tools:** Overlay segmentation maps on images and compute detailed analytics. - **Targeted data insights:** Filter views and identify outliers for focused evaluation. ### Dataset loading and exploration In our example, we can start by loading the KITTI dataset from the [FiftyOne Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html#kitti). The KITTI dataset provides left-camera images and 2D object detections, ideal for autonomous driving tasks. kitti\_data = foz.load\_zoo\_dataset("kitti", split="train", ) We can isolate specific classes (e.g., “van”), creating subsets of data for targeted analysis. van\_frames = kitti\_data.filter\_labels('ground\_truth', (F('label') == "Van") ) van\_frames Then we can isolate images with at least ten segmented objects detected above 80% confidence. view = sub\_data.filter\_labels(         "segmentations",         (F("confidence") > 0.8) ).match(         F("segmentations.detections").length() > 10 ) This filtering is only made possible if we apply a segmentation model and view the predictions. Let's do that now. ### Running predictions with SAM Note that applying SAM to generate segmentation masks requires the **transformers** and **segment-anything** libraries. In the following code, we'll select 500 random images on which to run predictions, apply the model, and save the segmentation results in a _label\_field_ for further analysis. \# Segment inside boxes sub\_data = kitti\_data.take(500) sub\_data.apply\_model(model,         label\_field="segmentations",         prompt\_field="ground\_truth",) Unlike other models, SAM allows location-specific segmentation. This approach avoids clutter by running segmentation only within specified bounding box areas. Semantic Segmentation classifies each pixel in an image into a category. Here, all cars are labeled as “cars” and people as “pedestrians.” The following **Patches View** allows close inspection of bounding box areas in the image. Instance Segmentation goes beyond classification to differentiate between individual objects. Here, every car in a scene is labeled as a unique entity. And then we can create an equivalent Patches View of the same instance segmentation task. By harnessing FiftyOne and SAM, image segmentation becomes more targeted and efficient, and yields actionable insights (such as evaluating model performance vs. the ground truth). ## Creating high-quality segmentation maps Producing high-quality segmentation maps requires well-annotated data, diverse perspectives, and balanced classes. Here’s how to ensure your segmentation maps are accurate, robust, and ready for real-world applications. ### Efficient data annotation High-quality annotations are the foundation of any segmentation model. Even though it is expensive to annotate every object, it is pretty important as poor labeling can mislead the model's predictions and skew the results. #### Best practices - Use tools like [CVAT](https://docs.voxel51.com/tutorials/cvat_annotation.html) that support pixel-level annotation. - Collaborate with domain experts to capture task-specific nuances, such as distinguishing similar classes. - Conduct quality checks with peers to spot inconsistencies in annotations, like missing a leg of a chair in the corner (trust me, it’s hard to annotate a chair). Source: [Voxel51](https://docs.voxel51.com/tutorials/cvat_annotation.html) ### Data augmentation Real-world images often show variability due to lighting, perspective, and spatial resolution changes. Introducing diversity through data augmentation enhances model robustness and generalization across varied conditions. Key techniques include: - Applying geometric transformations (rotations, scaling, etc.) to simulate different perspectives of the same input image. - Altering photometric features (brightness, saturation, etc.) to account for lighting variations for a robust model. - Focusing on specific regions (cropping) or standardized sizes (resizing) Tools like FiftyOne's integration with [Albumentations](http://tools%20like%20fiftyone's%20integration%20with%20albumentations%20can%20simplify%20these%20processes./) can simplify these processes. Source: [Voxel51](https://docs.voxel51.com/tutorials/data_augmentation.html) ### Balancing the dataset A class imbalance can introduce bias in the model's decision curve. For example, training solely on daytime urban images may harm performance in other settings. It's best to ensure that your dataset has a diverse representation with an almost equal proportion of classes. If not, then the underrepresented classes can be augmented to achieve balance. ## Interpreting segmentation maps Segmentation maps provide a significant amount of information, but their true value lies in how we act on them. Let's explore a few practical workflows. ### Overlaying segmentation maps on original images A raw segmentation mask may appear abstract to us, especially to non-technical audiences, as it is just a colorful mask of binary digits. We combine binary information with the original image and make sense of it by overlaying segmentation masks onto each object. This approach enhances clarity and makes the relationships between objects and their surroundings more visible. For example, an image of a van with a segmentation mask (in the second image) shows its position and its precise boundaries, complementing its exact size and shape. ### Quantifying object areas It doesn't stop there; these masks allow for accurate measurements of object areas as each pixel is associated with a class label. It might be helpful in various applications, such as estimating the proportion of road space occupied by vehicles in a scene to monitor traffic flow. By comparing the Manhattan distance between bounding boxes and filtering based on thresholds (e.g., selecting only those with a distance greater than 1), we can clearly distinguish occupied from free space, which is beneficial for urban planning and autonomous driving. \# Bboxes are in \[top-left-x, top-left-y, width, height\] format manhattan\_dist = F("bounding\_box")\[0\] + F("bounding\_box")\[1\] \# Only contains predictions whose bounding boxes' upper left corner \# is a Manhattan distance of at least 1 from the origin view = sub\_data.filter\_labels( "segmentations", manhattan\_dist > 1 ) ### Identifying object instances and locations Instance segmentation identifies and localizes individual objects, aiding analysis in areas such as: - **Autonomous vehicles:** Detecting pedestrians and vehicles for safe navigation. - **Agriculture:** Locating diseased crops for targeted treatment. - **Retail:** Optimizing store layouts by mapping product locations. Here is an example of the dataset's segmented masks of different “pedestrian” objects. ### Why segmentation is superior to bounding boxes And now cutting to the chase: why is segmentation superior to bounding boxes? While bounding boxes estimate object locations relatively precisely by enclosing them in a rectangle, segmentation maps offer additional pixel-level precision. Segmentation maps provide clearer shapes than bounding boxes, eliminating irrelevant background elements. Minor imperfections (e.g., a missing part of a foot in the below image) can have a cascading effect on model accuracy. Overall, segmentation maps deliver the granularity needed for detailed analyses. Another example shows segmented objects from the cyclist class with their masks overlayed. Bounding boxes can be a poor geometric fit for this sort of shape. Segmentation mapping solves this with pixel-level labeling. ## Real-world applications of segmentation maps Segmentation maps are revolutionizing how we analyze satellite data for environmental monitoring and urban planning. [Planet Labs](https://www.planet.com/products/satellite-imagery-analysis/) is leading this space with key use cases, including regular updates on building infrastructure changes, insights into new and existing road developments, and early identification of construction activities. Segmentation maps also empower advanced robotic capabilities across various industries. For example, in cooking, vision-guided segmentation helps identify food items and assists in food preparation. Imagine experiencing Michelin-star dining with robotic chefs delivering precision and artistry. [Chef Robotics](https://www.chefrobotics.ai/artificial-intelligence) is pioneering this vision with its ChefOS platform, which features pick-and-place robots and other innovative tools. ## The future of segmentation maps: Transformers The paper [Attention is All You Need](https://arxiv.org/abs/1706.03762) highlights significant progress in segmentation, with deep learning models now using transformers and attention mechanisms. **Vision Transformers (ViTs) and Segmenters** use self-attention for better global feature understanding, outperforming traditional convolutional methods. **Hybrid architectures (Swin Transformers)** combine CNNs and transformers, enhancing accuracy through hierarchical representations and finer spatial details. Other approaches like **weakly-supervised learning** reduces dependency on extensive labeled datasets. Some key methods - **Class Activation Maps (CAMs)** and **Grad-CAM++** help identify object regions with minimal supervision. The [CLIMS](https://arxiv.org/abs/2203.02668) model (2023) integrates text and images, achieving robust segmentation without pixel-level annotations. Segmentation models are increasingly tailored to specific industries: - **Healthcare:** [nnU-Net: No New Net](https://arxiv.org/abs/1809.10486) offers a self-configuring solution for medical images, setting new benchmarks in tumor detection. - **Agriculture:** [DeepWeeds](https://arxiv.org/abs/1810.05726) aids in identifying invasive plants, improving crop yield and pest control. - **Autonomous Systems:** [BEVFormer](https://arxiv.org/abs/2203.17270) enhances navigation for self-driving cars through bird’s-eye view perception and obstacle detection. - **Retail:** Improved segmentation maps enhance self-checkout systems by accurately identifying products and improving user experience. These advancements promise to make segmentation more accessible and impactful in real-world applications. ## Key takeaways In this article, we explored segmentation as a computer vision task using the KITTI dataset for autonomous navigation. Segmentation allows for a deeper analysis of images, while object detection provides a high-level overview of the objects present. We covered: - **Understanding Segmentation Maps:** Segmentation maps classify every pixel in an image, providing a comprehensive view of scenes through semantic and instance segmentation. - **Using FiftyOne for Visualization and Insights:** FiftyOne’s tools facilitate the visualization of segmentation maps, allowing developers to overlay predictions, explore dataset patterns, and gain insights. - **Generating High-Quality Segmentation Maps:** Accurate segmentation maps require careful data annotation and refinement. FiftyOne simplifies this process, enabling efficient and reliable results. ### Next steps With tools like FiftyOne, you can truly harness the power of segmentation maps in your projects! We invite you to [dive into FiftyOne](https://docs.voxel51.com/getting_started/install.html) and visualize your datasets—discover how engaging and insightful this can be. Don’t forget to check out our related [blogs](http://voxel51.com/blogs), where you can find advanced features and inspiring real-world case studies! We’re eager to hear your thoughts—how do you envision segmentation maps influencing the future of AI? Please share your insights or any challenges you face.. Keep in touch for more on using FiftyOne in advanced AI workflows. Whether you're aiming to boost model performance or learn new techniques, we’re here to support your journey! [object segmentation](https://voxel51.com/blog/tag/object-segmentation) [segmentations](https://voxel51.com/blog/tag/segmentations) [images](https://voxel51.com/blog/tag/images) [Visual AI](https://voxel51.com/blog/tag/visual-ai) Voxel Team Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/d2e24d0a14de508f8ccd36eaffd3b909c9f193b9-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Image Embeddings Transform Computer Vision Capabilities\\ \\ Learn\\ \\ • \\ \\ Nov 25, 2024](https://voxel51.com/blog/how-image-embeddings-transform-computer-vision-capabilities) [![](https://cdn.sanity.io/images/h6toihm1/production/9252e8bb5db5c4805f4a6f315b51527ee5c22072-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Guide to AI Image Segmentation\\ \\ Learn\\ \\ • \\ \\ Dec 19, 2024](https://voxel51.com/blog/a-guide-to-ai-image-segmentation) [![](https://cdn.sanity.io/images/h6toihm1/production/04eff439f21847f1068ea3f38b72f8bc7ec77f0e-2340x1308.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Image Similarity Search: Unlocking Pattern Detection in Visual Data\\ \\ Learn\\ \\ • \\ \\ Apr 16, 2025](https://voxel51.com/blog/image-similarity-search-unlocking-pattern-detection-in-visual-data) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-174-lllmstxt|> ## AI, ML, and Computer Vision Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/ecef2f07e575df5b0633b855b060d9c55dc86f01-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) ![](https://cdn.sanity.io/images/h6toihm1/production/ecef2f07e575df5b0633b855b060d9c55dc86f01-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=420) Virtual Americas Meetups AI, ML and Computer Vision Meetup en Español - August 21, 2025 Aug 21, 2025 9 AM Pacific Online. Fill in the form to register! Speakers ![](https://cdn.sanity.io/images/h6toihm1/production/9baa21472c4bbbec0978f5818bff5cdd3c5b0a3f-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Ernesto Cuartas DeepLearning.ai Bio ![](https://cdn.sanity.io/images/h6toihm1/production/ee8cb85526aeb9307655e52c76a0059029214940-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Paula Ramos, PhD Voxel51 Bio ![](https://cdn.sanity.io/images/h6toihm1/production/19b84529dd0146a685f35009b7137fa30c7d86c4-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Kevin Blanco Developer Advocate Bio ![](https://cdn.sanity.io/images/h6toihm1/production/325594468417cbd82ed7b0f2ad54ef48609a8513-480x480.png?auto=format&dpr=2&fit=max&q=75&w=42) Ivonne Mejia PMO Lead Bio About this event Hear talks from experts on the latest topics at the AI, ML and Computer Vision Meetup en Español. Schedule Quiero ser parte del mundo de AI, como lo logro? ![](https://cdn.sanity.io/images/h6toihm1/production/9baa21472c4bbbec0978f5818bff5cdd3c5b0a3f-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Ernesto Cuartas DeepLearning.ai Bio En esta charla, compartiré mi trayectoria personal hacia el mundo de la inteligencia artificial (IA), comenzando con mi formación como ingeniero electrónico y mi doctorado en neuroinformática. Destacaré cómo mi tesis laureada sobre modelos volumétricos realistas para la localización precisa de fuentes EEG abrió puertas a oportunidades en procesamiento digital y visión 3D. Con experiencia docente en la Universidad Nacional de Colombia y certificaciones en machine learning y deep learning, discutiré cómo estos hitos me llevaron a desempeñarme como desarrollador de currículo para DeepLearning.AI, ofreciendo valiosas lecciones para quienes deseen seguir un camino similar. Domina tus Datos Médicos: De la Curación al Impacto Clínico ![](https://cdn.sanity.io/images/h6toihm1/production/ee8cb85526aeb9307655e52c76a0059029214940-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Paula Ramos, PhD Voxel51 Bio Los datos de alta calidad son la base de un aprendizaje automático efectivo en el ámbito de la salud. Esta charla presenta estrategias prácticas y técnicas emergentes para gestionar datasets de imágenes médicas, desde la generación de datos sintéticos y la curación, hasta la evaluación y el despliegue. Comenzaremos con casos de estudio reales de investigadores y profesionales que están transformando sus flujos de trabajo en imágenes médicas mediante prácticas centradas en los datos. Luego pasaremos a un tutorial práctico utilizando FiftyOne, la plataforma open-source para la inspección visual de datasets y la evaluación de modelos. Los asistentes aprenderán a cargar, visualizar, curar y evaluar datasets médicos en distintos tipos de imágenes. Ya seas investigador, clínico o ingeniero de ML, esta charla te brindará herramientas e ideas prácticas para mejorar la calidad de tus datos, la fiabilidad de tus modelos y su impacto clínico. Agentes AI Multi-Fuente y Embebidos ![](https://cdn.sanity.io/images/h6toihm1/production/19b84529dd0146a685f35009b7137fa30c7d86c4-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Kevin Blanco Developer Advocate Bio Demostraré cómo construir agentes de IA contextualmente conscientes, capaz de responder y tomar acciones entre multiples sistemas privados y la implementación de RAG semántico a través de fuentes de datos dispares, embebidos en sistemas existentes, todo esto sin necesidad de una infraestructura compleja de MLOps Más allá del modelo: Metodología y buenas prácticas para liderar proyectos exitosos de IA con CPMAI ![](https://cdn.sanity.io/images/h6toihm1/production/325594468417cbd82ed7b0f2ad54ef48609a8513-480x480.png?auto=format&dpr=2&fit=max&q=75&w=96) Ivonne Mejia PMO Lead Bio El éxito de los proyectos de IA no depende solo del modelo o de los datos, sino de cómo se gestionan desde el inicio. En esta charla exploraremos la metodología CPMAI (Cognitive Project Management for AI) avalada por el Project Management Institute - PMI, un marco estructurado que permite a los equipos de IA alinear sus iniciativas con objetivos de negocio, gestionar riesgos éticos y mejorar los resultados. Compartiremos buenas prácticas que pueden ser adaptadas por profesionales técnicos para mejorar la entrega de valor en cada fase del proyecto e implementar soluciones de IA éticas y responsables. [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) AI, ML and Computer Vision Meetup en Español - August 21, 25 <|firecrawl-page-175-lllmstxt|> ## ArgosAI Case Study [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/d2f0a45f9678ebc503ec7b90fbeea293efb719ce-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) ArgosAI FiftyOne enables ArgosAI to increase the accuracy of their ML models for aviation use cases Apr 18, 2025 ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [ArgosAI](https://argosai.com/) specializes in enhancing airside safety and efficiency through advanced AI-driven systems like A-FOD® for foreign object debris detection and A-CAM® for continuous airside monitoring. Our AI platform ensures early anomaly detection, 24/7 operational awareness, and reduced physical runway inspections, leading to superior performance and reduced operational complexity. > "FiftyOne plays a vital role in our project, offering a robust framework for visualizing datasets and refining machine learning models. It significantly enhances the accuracy and reliability of our AI models, which is crucial for real-world applications. Moreover, it aligns with our commitment to continuous improvement in airside operations. > > FiftyOne's powerful tools for dataset curation and model analysis empower us to fine-tune our systems for optimal performance, further reinforcing our dedication to safety and efficiency in airside operations. An integral part of our workflow is FiftyOne's CVAT integration, streamlining the annotation process and thereby enhancing the accuracy of our machine learning models." – Serkan Mengi, Machine Learning Engineer at ArgosAI ![](https://cdn.sanity.io/images/h6toihm1/production/e4338646dae4b58ecf02ccd444fdd5ef900b5fcb-2560x1707.jpg?auto=format&dpr=2&fit=max&q=75&w=1600) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-176-lllmstxt|> ## Databricks and Voxel51 Partnership [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Databricks and Voxel51: Scaling Data-Centric Visual AI on the Data Intelligence Platform Jul 22, 2025 • 5 min read Article content In this article [AI teams are inundated with visual data, but starving for insights](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai#c892b6458116) [Unifying Databricks infrastructure with Voxel51’s visual layer](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai#cb3ed947df38) [Built for enterprise AI development](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai#784505eaec74) [What’s next: bringing auto-labeling to Databricks](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai#b96dbeee3a19) [The future of Visual AI/ML](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai#164480e2600b) In this article [AI teams are inundated with visual data, but starving for insights](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai#c892b6458116) [Unifying Databricks infrastructure with Voxel51’s visual layer](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai#cb3ed947df38) [Built for enterprise AI development](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai#784505eaec74) [What’s next: bringing auto-labeling to Databricks](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai#b96dbeee3a19) [The future of Visual AI/ML](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai#164480e2600b) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Databricks](https://www.databricks.com/) and [Voxel51](https://voxel51.com/) are excited to announce a partnership that unlocks a new era of scalable, data-centric multimodal AI. As more organizations adopt computer vision, they often struggle to manage and make sense of massive amounts of visual data. This partnership brings together Databricks’ Data Intelligence Platform with Voxel51’s powerful tools for visual data understanding and analysis. Powering visual AI systems in automotive tech, manufacturing, healthcare, agriculture, and beyond, this joint solution enables computer vision teams to have full control over their data so they can build faster and more reliably. _Join our upcoming webinar on **Sept 4, 2025, @9am PT**, showcasing how Porsche is advancing its autonomous vehicle (AV) development by leveraging the power of Visual AI from Voxel51 and the scalability of Databricks._ [Register for the webinar](https://events.databricks.com/FY260904-WB-EngineeringRD/registration?scid=701Vp00000U6EaCIAV&utm_medium=Partner&utm_source=n/a) ## **AI teams are inundated with visual data, but starving for insights** Visual data is everywhere. From autonomous vehicles, retail shelves, industrial automation, satellite imagery, and medical diagnostics, organizations developing AI heavily rely on it. And yet, despite having more visual data than ever, most AI teams struggle to make meaningful use of it. The common impulse is to “just add more data.” But without visibility into what that data contains, teams end up with bloated, noisy datasets that rarely move the performance needle. Even with a lakehouse full of visual data, it’s still difficult to answer: - What data is redundant, mislabeled, or low-quality? - Where are my model’s blind spots and failure cases? - What samples do I actually need to augment or relabel? Without tools to explore, retrieve, and understand visual data at scale, teams waste time and compute, chasing performance improvements that never materialize. The backing of an open, scalable Data Intelligence Platform, coupled with an intuitive visual intelligence and data engine, is critical to handling the speed, scale, and complexity of today’s AI systems. ## **Unifying Databricks infrastructure with Voxel51’s visual layer** This partnership brings together the core components of a modern visual AI workflow: - **Databricks** provides a scalable data infrastructure backbone with: - [Lakehouse storage](https://www.databricks.com/product/lakehouse-storage) for reliable, versioned data storage that ensures consistency. - [Unity Catalog](https://www.databricks.com/product/unity-catalog) for centralized governance, access control, and lineage across datasets and AI assets. - [Databricks Vector Search](https://www.databricks.com/product/machine-learning/vector-search) for high-speed similarity queries. - **Voxel51** provides an interactive visual intelligence layer and data engine to annotate, explore, curate, and analyze visual datasets and models. Teams building AI applications can now pair Voxel51’s rich visualization and analysis capabilities with [Databricks Vector Search](https://www.databricks.com/product/machine-learning/vector-search) engine for seamless embedding generation, indexing, and discovery—all governed through Databricks Unity Catalog. This provides the best of both worlds: the power of Databricks’ cloud-scale infrastructure and enterprise-grade governance, with the intuitive control of FiftyOne’s visual interface. ![](https://cdn.sanity.io/images/h6toihm1/production/c6f68bf5c7dcec4cd93a0fd3d273896506f48039-1600x562.png?auto=format&dpr=2&fit=max&q=75&w=1600) Together, the integrated solution enables: - Unified access to multimodal data (image, video, 3D, audio) to visualize, build, and augment high-quality datasets directly from [Databricks Unity Catalog](https://www.databricks.com/product/unity-catalog) (a centralized governance layer for all data and AI assets) and Databricks Volumes (the native storage layer for data and files). - Semantic and hybrid search (e.g., image or text queries) powered by Databricks Vector Search and Voxel51 - Embeddings generation using foundation models in Databricks, visualized interactively in FiftyOne - Integrated [model evaluation workflows](https://voxel51.com/evaluation) in FiftyOne to surface failure modes and data gaps - [Automated labeling](https://voxel51.com/annotation) QA pipelines to verify, tag, and improve annotations in context What makes this solution stand out in production is its tight integration and scalability. Teams can create a Vector Search index backed by a Delta table of embeddings, and Databricks handles everything from storage and access control to REST-based querying. The result is a fast, governed, and seamless way to explore and manage visual datasets, whether it’s thousands or billions of visual data samples. ## **Built for enterprise AI development** Every enterprise AI use case is different, and so are the data and infrastructure requirements. This strategic partnership provides the essential components that vision teams need to move from raw data to production-quality datasets and models. And it all happens on infrastructure that scales. The joint solution brings together Databricks’ trusted data and AI platform and Voxel51’s visual intelligence layer to help teams build, refine, and deploy computer vision systems at scale—faster, more securely, and with full control over data quality. - **_Faster iteration:_** Visualize, curate, and evaluate massive multimodal datasets directly from Unity Catalog and speed up model development - **_Smarter search:_** Run high-speed similarity search over millions of images and videos using natural language or visual examples. - **_Better models:_** Identify edge cases, label issues, and failure modes to improve training data. - **Secure integration:** Build secure multimodal AI natively on Databricks with Delta Lake and Unity Catalog for governance. - **_Flexible and extensible_**:Supports cloud, hybrid, and airgapped environments with plugin extensibility so teams have the flexibility to adapt for specific use case scenarios. ## **What’s next: bringing auto-labeling to Databricks** As we work together, our next milestone is to bring [Voxel51’s Verified Auto Labeling technology](https://voxel51.com/annotation) directly into Databricks, enabling scalable, production-grade annotation workflows powered by foundation models and Databricks compute. By allowing ML teams to run large-scale, configurable auto-labeling pipelines using Databricks, organizations can use foundational AI models to enable faster, cheaper, and more efficient dataset annotation and curation at production scale. ## **The future of Visual AI/ML** As foundation models and scalable infrastructure continue to evolve, the bottleneck in real-world AI is shifting from algorithms to the quality, curation, and deep understanding of data. The future of visual AI lies in data-centric, multimodal systems where competitive advantage comes not from bigger models, but from better, more explainable data. The Databricks and Voxel51 partnership gives AI teams the full-stack solution they need to tackle that: from fast, secure data infrastructure to intuitive, visual tools that put human understanding back into the loop. If you’re working on multimodal or vision-based AI, we’d love to show you what this looks like in practice. [Request a demo](https://voxel51.com/sales) to see how Databricks and Voxel51 can help you build smarter, faster, and more reliable AI systems with your visual data. Join our upcoming webinar on **Sept 4, 2025, @9am PT** to learn how Porsche is advancing its autonomous vehicle (AV) development by leveraging the power of Voxel51 and Databricks. [Register here](https://events.databricks.com/FY260904-WB-EngineeringRD/registration?scid=701Vp00000U6EaCIAV&utm_medium=Partner&utm_source=n/a). [vector search](https://voxel51.com/blog/tag/vector-search) [integrations](https://voxel51.com/blog/tag/integrations) [dataset curation](https://voxel51.com/blog/tag/dataset-curation) ![](https://cdn.sanity.io/images/h6toihm1/production/8d61ff90b31d151405f9e21a33c2802509f34651-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Brian Moore Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/6aafb2b5fa699824c252fabfe2607eaeb820616a-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Enabling the AV Datasets of the Future with NVIDIA NuRec and FiftyOne\\ \\ Product & News\\ \\ • \\ \\ Aug 11, 2025](https://voxel51.com/blog/enabling-av-datasets-nvidia-nurec-and-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/663dd6a3e6f3e57a03425f932440b5d242133451-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Search and curate video data with FiftyOne, Twelve Labs, and Databricks Vector Search\\ \\ Product & News, Integrations\\ \\ • \\ \\ Jun 5, 2025](https://voxel51.com/blog/search-curate-video-fiftyone-databricks-twelvelabs) [![](https://cdn.sanity.io/images/h6toihm1/production/35e10bff3f7e49854806cbbd163022932bf07548-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How Voxel51 is Powering Physical AI with Databricks\\ \\ Product & News\\ \\ • \\ \\ Aug 20, 2025](https://voxel51.com/blog/powering-physical-ai-with-voxel51-and-databricks) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-177-lllmstxt|> ## Computer Vision Glossary [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) # Computer vision & visual AI glossary Browse our glossary to learn about the techniques and technologies being used in the world of artificial intelligence, machine learning, computer vision, and active learning. A B C D G I L M A [Accuracy](https://voxel51.com/glossary/accuracy) [Annotation](https://voxel51.com/glossary/annotation) B [Bounding Box](https://voxel51.com/glossary/bounding-box) C [Classification in Computer Vision](https://voxel51.com/glossary/classification-in-computer-vision) [Confusion Matrix](https://voxel51.com/glossary/confusion-matrix) [Convolutional Neural Network (CNN)](https://voxel51.com/glossary/convolutional-neural-network-cnn) D [Data Curation](https://voxel51.com/glossary/data-curation) [Data Labeling](https://voxel51.com/glossary/data-labeling) [Dataset](https://voxel51.com/glossary/dataset) G [Generative Pre-Trained Transformer](https://voxel51.com/glossary/gpt) I [Image Blurriness](https://voxel51.com/glossary/image-blurriness) [Image Embeddings](https://voxel51.com/glossary/image-embeddings) L [Logistic Regression](https://voxel51.com/glossary/logistic-regression) M [Mean Squared Error](https://voxel51.com/glossary/mean-squared-error) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-178-lllmstxt|> ## FiftyOne Plugins Overview [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/e1903ef08e46b0f1a697bc38cbbd544a74ec3e10-3024x640.png?auto=format&dpr=2&fit=max&q=75&w=1512) # FiftyOne Plugins Unlock infinite ways to extend and customize your development with FiftyOne. Task Search Generative Vision-Language Zero-Shot OCR Evaluation EDA Visualization Utility Quality Control Data Augmentation Annotation Ingestion Text-to-Image VQA Experiment Tracking Modality Audio Time Series Multimodal PDF Video Image Text Type Community developed Voxel51 developed [Active Learning\\ \\ Accelerate your data labeling
 with Active Learning](https://github.com/jacobmarks/active-learning-plugin) [Annotation\\ \\ Utilities for integrating
 FiftyOne with annotation tools](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/annotation) [Anonymize\\ \\ Anonymize/blur images based on a FiftyOne Detections field](https://github.com/swheaton/fiftyone-media-anonymization-plugin) [Audio Loader\\ \\ Load in your audio datasets as spectograms](https://github.com/danielgural/audio_loader) [Audio Retrieval\\ \\ Find the images in your dataset most similar to an audio file](https://github.com/jacobmarks/audio-retrieval-plugin) [Brain\\ \\ Utilities for working
 with the FiftyOne Brain](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/brain) [Clustering\\ \\ Compute clustering on your data in a visual, intuitive way with FiftyOne and Sklearn](https://github.com/jacobmarks/clustering-plugin) [Clustering Algorithms\\ \\ Find the clusters in your data using some of the best algorithms available](https://github.com/danielgural/clustering_algorithms) [Concept Interpolation\\ \\ Find images that best interpolate between two text-based extremes](https://github.com/jacobmarks/concept-interpolation) [Concept Space Traversal\\ \\ Navigate concept space with 
CLIP, vector search, and FiftyOne](https://github.com/jacobmarks/concept-space-traversal-plugin) [Dashboard\\ \\ Create your own custom dashboards from within the App](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/dashboard) [Data Augmentation\\ \\ Apply, test, compose & save augmentation transformations to your dataset](https://github.com/jacobmarks/fiftyone-albumentations-plugin) [Delegated\\ \\ Utilities for managing your delegated operations](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/delegated) [Double Band Filter\\ \\ Filter on two numeric ranges simultaneously](https://github.com/jacobmarks/double-band-filter-plugin) [Edit Label Attributes\\ \\ Edit attributes of your labels directly in the FiftyOne App](https://github.com/ehofesmann/edit_label_attributes) [Emoji Search\\ \\ Semantically search emojis and copy to clipboard](https://github.com/jacobmarks/emoji-search-plugin) [Evaluation\\ \\ Utilities for evaluating 
models with FiftyOne](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/evaluation) [FiftyOne Tile\\ \\ Tile your high resolution images to squares for training small object detection models](https://github.com/mmoollllee/fiftyone-tile) [FiftyOne Timestamps\\ \\ Compute datetime-related fields (e.g. sunrise, dawn, evening, weekday) from sample filenames or creation dates](https://github.com/mmoollllee/fiftyone-timestamps) [Filter Values\\ \\ Find the images in your dataset Filter a field of your FiftyOne dataset by one or more values](https://github.com/ehofesmann/filter-values-plugin) [Florence-2\\ \\ Connect Microsoft's Florence-2 Vision-Language Model to your data](https://github.com/jacobmarks/fiftyone_florence2_plugin) [GPT-4 Vision\\ \\ Analyze a single image or multiple images at once with GPT-4 Vision](https://github.com/jacobmarks/gpt4-vision-plugin) [Hello World\\ \\ An example of JavaScript
 and Python components and operators in a single plugin](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/hello-world) [Hugging Face Hub\\ \\ Push FiftyOne datasets to the Hugging Face Hub, and load datasets from the Hub into FiftyOne](https://github.com/voxel51/fiftyone-huggingface-plugins/tree/main/plugins/huggingface_hub) [Hugging Face Transformers\\ \\ Apply transformer models from the Hugging Face Hub to your FiftyOne datasets](https://github.com/voxel51/fiftyone-huggingface-plugins/tree/main/plugins/transformers) [I/O\\ \\ A collection of 
import/export utilities](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/io) [Image Captioning\\ \\ Generate & store captions for samples using state-of-the-art captioning models](https://github.com/jacobmarks/fiftyone-image-captioning-plugin) [Image Deduplication\\ \\ Find exact and approximate
 duplicates in your dataset](https://github.com/jacobmarks/image-deduplication-plugin) [Image Issues\\ \\ Find common image quality issues in your datasets](https://github.com/jacobmarks/image-quality-issues) [Image Quality Issues\\ \\ Find common image quality 
issues in your datasets](https://github.com/jacobmarks/image-quality-issues) [Image to Video\\ \\ Bring images to life with Stable Video Diffusion](https://github.com/danielgural/img_to_video_plugin) [Indexes\\ \\ Utilities working with
 FiftyOne database indexes](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/indexes) [Keyword Search\\ \\ Perform keyword search
on a specified field](https://github.com/jacobmarks/keyword-search-plugin) [Line2D\\ \\ Visualize x,y-Points
as a line chart](https://github.com/wayofsamu/line2d) [MLflow\\ \\ Track model training experiments on your FiftyOne datasets with MLflow](https://github.com/voxel51/fiftyone_mlflow_plugin) [Model Comparison\\ \\ Compare two object
detection models](https://github.com/allenleetc/model-comparison) [Multi Annotator Toolkit\\ \\ Find and analyze annotation issues in datasets with multiple annotators per image](https://github.com/Madave94/multi-annotator-toolkit) [Multimodal RAG\\ \\ Create and test multimodal RAG pipelines with LlamaIndex, Milvus, and FiftyOne](https://github.com/jacobmarks/fiftyone-multimodal-rag-plugin) [OCR\\ \\ Run optical character 
recognition with PyTesseract](https://github.com/jacobmarks/pytesseract-ocr-plugin) [Operator Example\\ \\ Examples showing how to use the operator type system to build custom FiftyOne operations](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/operator-examples) [Optimal Confidence Threshold\\ \\ Find the optimal confidence threshold for your detection models automatically](https://github.com/danielgural/optimal_confidence_threshold) [Outlier Detection\\ \\ Find those troublesome outliers in your dataset automatically](https://github.com/danielgural/outlier_detection) [PDF Loader\\ \\ Load your PDF documents into FiftyOne as per-page images](https://github.com/brimoor/pdf-loader) [Panel Example\\ \\ Examples demonstrating common patterns for building Python panels](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/panel-examples) [Plotly Map\\ \\ Plotly-based Map Panel with adjustable marker cosmetics](https://github.com/allenleetc/plotly-map-panel) [Plugin Utilities\\ \\ Utilities for managing and
 building FiftyOne plugins](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/plugins) [Pytesseract\\ \\ Run optical character recognition with PyTesseract](https://github.com/jacobmarks/pytesseract-ocr-plugin) [Reverse Image Search\\ \\ Find the images in your dataset 
most similar to an image from file system or the internet](https://github.com/jacobmarks/reverse-image-search-plugin) [Runs\\ \\ Utilities for managing custom runs on your dataset](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/runs) [SDK Utilities\\ \\ Call your favorite SDK utilities from the App](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/utils) [Segments.ai\\ \\ Integrate FiftyOne with the Segments.ai annotation tool](https://github.com/segments-ai/segments-voxel51-plugin) [Semantic Document Search\\ \\ Perform semantic search 
on text in your documents](https://github.com/jacobmarks/semantic-document-search-plugin) [Semantic Video Search\\ \\ Use Twelve Labs to search your videos with natural language](https://github.com/danielgural/semantic_video_search) [Text-to-Image\\ \\ Create your own AI Art Gallery with Text-to-image models and FiftyOne](https://github.com/jacobmarks/text-to-image) [Twilio Automation\\ \\ Automate data ingestion
 with Twilio](https://github.com/jacobmarks/twilio-automation-plugin) [VQA\\ \\ Ask (and answer) open-ended visual questions about your images](https://github.com/jacobmarks/vqa-plugin) [VoxelGPT\\ \\ An AI assistant that can query
 visual datasets, search the 
FiftyOne docs, and answer general computer vision questions](https://github.com/voxel51/voxelgpt) [YouTube Player Panel\\ \\ Play YouTube videos
in the FiftyOne app](https://github.com/jacobmarks/fiftyone-youtube-panel-plugin) [Zero-Shot Prediction\\ \\ Run zero-shot (open vocabulary) prediction on your data](https://github.com/jacobmarks/zero-shot-prediction-plugin) [Zoo\\ \\ Download datasets & run inference with models from the FiftyOne Zoo, without leaving the App](https://github.com/voxel51/fiftyone-plugins/tree/main/plugins/zoo) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-179-lllmstxt|> ## Voxel51 Press Releases [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/14713df0d4dec67cd3bb5e9c292c607820df061d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=540&q=75&w=960) Featured Databricks and Voxel51: Scaling Data-Centric Visual AI on the Data Intelligence Platform [Read article](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/c7a1751b81cb92767b622014dfbc4019beb7bce5-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured Voxel51 Accelerates Autonomous Vehicle Development with NVIDIA Omniverse Integration [Read article](https://www.prnewswire.com/news-releases/voxel51-accelerates-autonomous-vehicle-development-with-nvidia-omniverse-integration-302091979.html) ![](https://cdn.sanity.io/images/h6toihm1/production/c727ebde95a4d6135d0bf2740a99aa8931220c33-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured Voxel51 Raises $30M Series B Funding to Make Visual AI a Reality [Read article](https://voxel51.com/blog/voxel51-raises-30m-series-b-funding-to-make-visual-ai-a-reality) ![](https://cdn.sanity.io/images/h6toihm1/production/083d5b43658db73b3c96641645aef9229377f47f-2162x1216.webp?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=540&q=75&w=960) Featured Voxel51 & Google Collaborate to Make Downloading & Visualizing Open Images A Breeze [Read article](https://voxel51.com/blog/fiftyone-open-images-collaboration) ![](https://cdn.sanity.io/images/h6toihm1/production/14713df0d4dec67cd3bb5e9c292c607820df061d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=540&q=75&w=960) Featured Databricks and Voxel51: Scaling Data-Centric Visual AI on the Data Intelligence Platform [Read article](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/c7a1751b81cb92767b622014dfbc4019beb7bce5-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured Voxel51 Accelerates Autonomous Vehicle Development with NVIDIA Omniverse Integration [Read article](https://www.prnewswire.com/news-releases/voxel51-accelerates-autonomous-vehicle-development-with-nvidia-omniverse-integration-302091979.html) ![](https://cdn.sanity.io/images/h6toihm1/production/c727ebde95a4d6135d0bf2740a99aa8931220c33-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured Voxel51 Raises $30M Series B Funding to Make Visual AI a Reality [Read article](https://voxel51.com/blog/voxel51-raises-30m-series-b-funding-to-make-visual-ai-a-reality) ![](https://cdn.sanity.io/images/h6toihm1/production/083d5b43658db73b3c96641645aef9229377f47f-2162x1216.webp?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=540&q=75&w=960) Featured Voxel51 & Google Collaborate to Make Downloading & Visualizing Open Images A Breeze [Read article](https://voxel51.com/blog/fiftyone-open-images-collaboration) All press [![](https://cdn.sanity.io/images/h6toihm1/production/c7a1751b81cb92767b622014dfbc4019beb7bce5-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ NVIDIA Opens Portals to World of Robotics With New Omniverse Libraries, Cosmos Physical AI Models and AI Computing Infrastructure\\ \\ Aug 12, 2025](https://nvidianews.nvidia.com/news/nvidia-opens-portals-to-world-of-robotics-with-new-omniverse-libraries-cosmos-physical-ai-models-and-ai-computing-infrastructure) [![](https://cdn.sanity.io/images/h6toihm1/production/8b39b4ff857d388f096362caec22e8124d9c54a3-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ China’s AI Mirage: How “Open Source” Hides What Matters Most\\ \\ Aug 4, 2025](https://www.unite.ai/chinas-ai-mirage-how-open-source-hides-what-matters-most/) [![](https://cdn.sanity.io/images/h6toihm1/production/b0fe74c8178b3203491cedc39c06c7a399195e7f-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Companies Should Embrace 'Small AI' for ROI\\ \\ Jul 30, 2025](https://fortune.com/2025/07/30/what-is-artificial-intelligence-return-on-investment-small-ai/) [![](https://cdn.sanity.io/images/h6toihm1/production/6cf24a31341b73d93a98508e1eab908a23b4288f-800x418.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Autonomous Driving, Visual AI, and the Road Ahead with Porsche and Voxel51\\ \\ Jul 30, 2025](https://open.spotify.com/episode/01z4OB6WSJ6Pti0dRPSAWX) [![](https://cdn.sanity.io/images/h6toihm1/production/14713df0d4dec67cd3bb5e9c292c607820df061d-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Databricks and Voxel51: Scaling Data-Centric Visual AI on the Data Intelligence Platform\\ \\ Jul 22, 2025](https://voxel51.com/blog/databricks-and-voxel51-partnership-scaling-data-centric-visual-ai) [![](https://cdn.sanity.io/images/h6toihm1/production/626c4fbd38ccca99f4d11594b217558624530043-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Jason Corso on Manufacturing Tomorrow Podcast\\ \\ Jul 14, 2025](https://podcast.osu.edu/mfgtmw/2025/07/14/jason-corso-voxel51/) [![](https://cdn.sanity.io/images/h6toihm1/production/5867679ced4ba5824f3972314c0f9548c8b5cdf3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ AI in Automotive Podcast - Dr Jason Corso, Co-founder & Chief Scientist - Voxel51\\ \\ Jul 11, 2025](https://www.ai-in-automotive.com/aiia/504/jasoncorso) [![](https://cdn.sanity.io/images/h6toihm1/production/cdf4cf020ec30891e9a1eec76a52a89036f322d2-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Smart Tech Solutions To Help Retailers Minimize Shrink\\ \\ Jul 9, 2025](https://www.forbes.com/councils/forbestechcouncil/2025/07/09/smart-tech-solutions-to-help-retailers-minimize-shrink/) [![](https://cdn.sanity.io/images/h6toihm1/production/576ce6fa1aae2291a602d54c66403324feec90bb-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Zero-Shot Auto-Labeling: The End of Annotation for Computer Vision with Jason Corso\\ \\ Jun 24, 2025](https://twimlai.com/podcast/twimlai/zero-shot-auto-labeling-the-end-of-annotation-for-computer-vision/) [![](https://cdn.sanity.io/images/h6toihm1/production/cdf4cf020ec30891e9a1eec76a52a89036f322d2-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Tesla Model Y Juniper FSD Test Drive Fails 3 Times, Forbes\\ \\ Jun 23, 2025](https://www.forbes.com/sites/brookecrothers/2025/06/22/my-tesla-model-y-juniper-fsd-test-drive-fails-spectacularly/) [![](https://cdn.sanity.io/images/h6toihm1/production/6ef6221c55258d5132bdb6deaa8c5494cbbd7cda-3840x2161.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&rect=6,15,3825,2140&w=480)\\ \\ Zero-shot auto-labeling saves 100,000x on annotation costs\\ \\ Jun 4, 2025](https://voxel51.com/blog/zero-shot-auto-labeling-rivals-human-performance) [![](https://cdn.sanity.io/images/h6toihm1/production/3a033eb2f111561b2fb00f3afc1e5820293ba6f0-1200x675.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ New Data and Model Workflows from Voxel51 Accelerate Visual AI Development for Enterprises\\ \\ Mar 25, 2025](https://voxel51.com/blog/new-data-and-model-workflows-from-voxel51-accelerate-visual-ai-development-for-enterprises) Load more [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-180-lllmstxt|> ## On-Demand Webinars [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/849de3e6ac8096687759c612a76c564e100463b1-3024x640.png?auto=format&dpr=2&fit=max&q=75&w=1512) # Voxel51 On-Demand Webinars Format Virtual Region Americas Category Webinars & Workshops Meetups Industry Healthcare Agriculture [![](https://cdn.sanity.io/images/h6toihm1/production/1b5ee3eebfb4e43dcbb8b4a0eb78d19855cb00fd-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Understanding Visual Agents - August 7, 2025\\ \\ Join the Meetup to hear talks from experts on understanding visual agents.\\ \\ Aug 7, 2025](https://voxel51.com/events/understanding-visual-agents-august-7-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/7a2015ac7fa2e7fc9be86672715c936918544b69-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Women in AI - July 24\\ \\ Hear talks from experts on cutting-edge topics in AI, ML, and computer vision on July 24.\\ \\ Jul 24, 2025](https://voxel51.com/events/women-in-ai-july-24) [![](https://cdn.sanity.io/images/h6toihm1/production/56af9060688bf339888e57b7a508690dacb75b11-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Getting Started with FiftyOne for Healthcare Use Cases - July 23, 2025\\ \\ Visual AI is revolutionizing healthcare by enabling more accurate diagnoses, streamlining medical workflows, and uncover...\\ \\ Jul 23, 2025](https://voxel51.com/events/getting-started-with-fiftyone-for-healthcare-use-cases-july-23-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/0b65abff4c08a6ad58713d835a3418f3a7e4a38d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ AI, ML and Computer Vision Meetup - July 17, 2025\\ \\ Join the Meetup to hear talks from experts on cutting-edge topics across AI, ML, and computer vision.\\ \\ Jul 17, 2025](https://voxel51.com/events/ai-ml-and-computer-vision-meetup-july-17-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/13e655d277629be8c1b9e0c056902773db6b9621-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Best of CVPR – July 11, 2025\\ \\ Welcome to the Best of CVPR series, your virtual pass to some of the groundbreaking research, insights, and innovations ...\\ \\ Jul 11, 2025](https://voxel51.com/events/best-of-cvpr-july-11-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/325653df18e5644209b2cf7743bdc67b1c51c8a4-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Best of CVPR – July 10, 2025\\ \\ Welcome to the Best of CVPR series, your virtual pass to some of the groundbreaking research, insights, and innovations ...\\ \\ Jul 10, 2025](https://voxel51.com/events/best-of-cvpr-july-10-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/a40cbd0a71e7aa91b458ffe06d5eb6d1fc8bb43b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Best of CVPR – July 9, 2025\\ \\ Welcome to the Best of CVPR series, your virtual pass to some of the groundbreaking research, insights, and innovations ...\\ \\ Jul 9, 2025](https://voxel51.com/events/best-of-cvpr-july-9-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/4dbe6c37a9ce253eaf07d67643e7feaae0522d45-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual AI in Healthcare – June 27, 2025\\ \\ Hear talks from experts on cutting-edge topics at the intersection of AI, ML, computer vision and healthcare.\\ \\ Jun 27, 2025](https://voxel51.com/events/visual-ai-in-healthcare-june-27-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/2cf62ff9c416cc2d719ca3b763968e6d8a54cb32-1952x1112.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Visual AI in Healthcare – June 26, 2025\\ \\ Hear talks from experts on cutting-edge topics at the intersection of AI, ML, computer vision and healthcare.\\ \\ Jun 26, 2025](https://voxel51.com/events/visual-ai-in-healthcare-june-26-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/70c926c57595d018caf8c55437c48f54cdc1db81-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual AI in Healthcare – June 25, 2025\\ \\ Hear talks from experts on cutting-edge topics at the intersection of AI, ML, computer vision and healthcare.\\ \\ Jun 25, 2025](https://voxel51.com/events/visual-ai-in-healthcare-june-25-2025) [![Auto Labeling Workshop: Smarter annotation at scale](https://cdn.sanity.io/images/h6toihm1/production/5212603b698f079790dcca848109776fbc871d0a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Verified Auto Labeling: Smarter Annotation at Scale – June 24, 2025\\ \\ Want to take your computer vision workflows to the next level? Then this workshop is for you! Join us for 90 minutes as ...\\ \\ Jun 24, 2025](https://voxel51.com/events/verified-auto-labeling-smarter-annotation-at-scale-june-24-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/e1d2d134c3c81b6e4e39c24e70635d5f2b2a9c41-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ AI, ML and Computer Vision Meetup en Español – June 20, 2025\\ \\ Join the Meetup to hear talks in Spanish from experts on cutting-edge topics across AI, ML, and computer vision.\\ \\ Jun 20, 2025](https://voxel51.com/events/ai-ml-and-computer-vision-meetup-en-espanol-june-20-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/619dc7486f70359ac3237c38c591b68808c05386-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ AI, ML and Computer Vision Meetup – June 19, 2025\\ \\ Join the Meetup to hear talks from experts on cutting-edge topics across AI, ML, and computer vision.\\ \\ Jun 19, 2025](https://voxel51.com/events/ai-ml-and-computer-vision-meetup-june-19-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/f15572465b6cd96530e85d1fc3dd5f1e7f2675e7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Getting Started with FiftyOne Workshop – June 18, 2025\\ \\ Join us for the free 90-minute, hands-on workshop to learn how to get started with open source FiftyOne.\\ \\ Jun 18, 2025](https://voxel51.com/events/getting-started-with-fiftyone-workshop-june-18-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/f0b44cf7cf4b959162ac789d04d9c31fa1e879e4-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Mosaic AI + FiftyOne: Scaling Physical AI for Mobility and Autonomous Use Cases – June 17, 2025\\ \\ In this technical session with machine learning engineer Dan Gural, he’ll show you how the Mosaic AI integration works i...\\ \\ Jun 17, 2025](https://voxel51.com/events/mosaic-ai-fiftyone-scaling-physical-ai-for-mobility-and-autonomous-use-cases-june-17-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/1df3365d49d3baf42d81c3a983d77bd43bbd6154-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Building Visual AI in the Enterprise Workshop – June 4, 2025\\ \\ Want to take your computer vision workflows to the next level? Then this workshop is for you! Join us for 90 minutes as ...\\ \\ Jun 4, 2025](https://voxel51.com/events/building-visual-ai-in-the-enterprise-workshop-june-4-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/f8a88691e9a9fad6e00d99c1cb4e08ab40a3424c-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Best of WACV – May 30, 2025\\ \\ Welcome to the Best of WACV series, your virtual pass to some of the groundbreaking research, insights, and innovations ...\\ \\ May 30, 2025](https://voxel51.com/events/best-of-wacv-may-30-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/10b6f9117aec3c0df287aa0ec2e5543ae7c77837-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Best of WACV – May 29, 2025\\ \\ Welcome to the Best of WACV series, your virtual pass to some of the groundbreaking research, insights, and innovations ...\\ \\ May 29, 2025](https://voxel51.com/events/best-of-wacv-may-29-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/b0fc180e0237b8cdef9d31d4df8e77ab335b6903-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ AI, ML and Computer Vision Meetup – May 22, 2025\\ \\ May 22, 2025](https://voxel51.com/events/ai-ml-and-computer-vision-meetup-may-22-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/cf98729af40ac4bd03b3a2bc3740b568b38ab8db-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Advanced Computer Vision Data Curation and Model Evaluation Workshop – May 21, 2025\\ \\ May 21, 2025](https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-workshop-may-21-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/930c1f96d2fd976d9e160abfd2e72d72a2627814-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Image Generation: Diffusion Models & U-Net Workshop – May 20, 2025\\ \\ May 20, 2025](https://voxel51.com/events/image-generation-diffusion-models-u-net-workshop-may-20-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/9722bbc96eb8ae7b4a1b406da7e003210c83d9e9-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Object Detection & Instance Segmentation: YOLO in Practice Workshop – May 13, 2025\\ \\ May 13, 2025](https://voxel51.com/events/object-detection-instance-segmentation-yolo-in-practice-workshop-may-13-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/5eb759151b8a947bbf71512a5ad3d02b9bab10e3-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Image Embeddings: Zero-shot Classification with CLIP Workshop – May 6, 2025\\ \\ May 6, 2025](https://voxel51.com/events/image-embeddings-zero-shot-classification-with-clip-workshop-may-6-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/b67f49f06d96d5d5d45f61d94de48ee470a8da2e-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Model Optimization: Data Augmentation & Regularization Workshop – April 29, 2025\\ \\ Apr 29, 2025](https://voxel51.com/events/model-optimization-data-augmentation-regularization-workshop-april-29-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/0b3091351d4e58c2b9e1f12d0ff2a0a2d49a4780-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ AI, Machine Learning & Computer Vision Meetup – April 24, 2025\\ \\ Apr 24, 2025](https://voxel51.com/events/ai-machine-learning-computer-vision-meetup-april-24-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/cf98729af40ac4bd03b3a2bc3740b568b38ab8db-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Advanced Computer Vision Data Curation and Model Evaluation Workshop – April 23, 2025\\ \\ Apr 23, 2025](https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-workshop-april-23-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/8f9dd8de0ff1c79457f02a63b15d0f783819214f-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Convolutional Neural Networks: Advanced Upsampling & U-Net for Semantic Segmentation Workshop – April 22, 2025\\ \\ Apr 22, 2025](https://voxel51.com/events/convolutional-neural-networks-advanced-upsampling-u-net-for-semantic-segmentation-workshop-april-22-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/05bdf560d751c4c988761ba67a1e3bfbd00f0c08-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Interpretability in Computer Vision: CAM & Grad-CAM Workshop – April 15, 2025\\ \\ Apr 15, 2025](https://voxel51.com/events/interpretability-in-computer-vision-cam-grad-cam-workshop-april-15-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/addf6701fee69208f697beef31449a14c766ee2a-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Multi-label Classification with Binary Cross Entropy: Amazon Satellite Images Workshop – April 8, 2025\\ \\ Apr 8, 2025](https://voxel51.com/events/multi-label-classification-with-binary-cross-entropy-amazon-satellite-images-workshop-april-8-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/8c95ed68864812b7e340d704c8405ad3f4a27100-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Training Techniques for Convolutional Networks Workshop – April 1, 2025\\ \\ Apr 1, 2025](https://voxel51.com/events/training-techniques-for-convolutional-networks-workshop-april-1-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/7036a21059f542e3905da18327ad1bccdb9c3ca9-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Visual AI for Agriculture – March 26, 2025\\ \\ Mar 26, 2025](https://voxel51.com/events/visual-ai-in-agriculture-march-26) [![](https://cdn.sanity.io/images/h6toihm1/production/f62041aabd52cf6dd5c83ed0736e3107af0880de-653x314.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Convolutional Neural Networks – LeNet5 Workshop – March 25, 2025\\ \\ Mar 25, 2025](https://voxel51.com/events/convolutional-neural-networks-lenet5-workshop-march-25-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/5811fb1137d7fc7a122ad555769f67f812a9edd8-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ AI, Machine Learning & Computer Vision Meetup – March 20, 2025\\ \\ Mar 20, 2025](https://voxel51.com/events/ai-machine-learning-computer-vision-meetup-march-20-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/cf98729af40ac4bd03b3a2bc3740b568b38ab8db-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Advanced Computer Vision Data Curation and Model Evaluation Workshop – March 19, 2025\\ \\ Mar 19, 2025](https://voxel51.com/events/advanced-computer-vision-data-curation-and-model-evaluation-workshop-march-19-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/48972f0c30204adb9d9d04bd3dd3db79c1d849e4-1306x628.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Training & Evaluation of Classification Models Workshop – March 18, 2025\\ \\ Mar 18, 2025](https://voxel51.com/events/training-evaluation-of-classification-models-workshop-march-18-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/05dd0018a7fe2dc1512f4500cd10b75756131369-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Neural Networks Fundamentals: Multilayer Perceptrons for Regression Workshop – March 11, 2025\\ \\ Mar 11, 2025](https://voxel51.com/events/neural-networks-fundamentals-multilayer-perceptrons-for-regression-workshop-march-11-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/e90f2527954eebc089a7b8ac98ec431b506cdd31-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Foundations of Computer Vision Workshop – March 4, 2025\\ \\ Mar 4, 2025](https://voxel51.com/events/foundations-of-computer-vision-workshop-march-4-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/500efce2909bc916201b78f773a36b4cb6c889a4-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ AI, Machine Learning & Computer Vision Meetup – Feb 20, 2025\\ \\ Feb 20, 2025](https://voxel51.com/events/ai-machine-learning-computer-vision-meetup-feb-20-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/98908c53f94a0a46159f5c4feb1fc313eeafa26b-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Best of NeurIPS – Feb 6, 2025\\ \\ Feb 6, 2025](https://voxel51.com/events/best-of-neurips-feb-6-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/a59cc71173d2ad6d2c6021da295628daf0a8b27b-1024x576.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Best of NeurIPS – Feb 4\\ \\ Feb 4, 2025](https://voxel51.com/events/best-of-neurips-feb-4-2025) [![](https://cdn.sanity.io/images/h6toihm1/production/79f3e6076700bb28a224195b44611f3a2de19e0c-1024x576.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual AI for Geospatial – Jan 29, 2025\\ \\ Jan 29, 2025](https://voxel51.com/events/visual-ai-for-geospatial-jan-29-2025) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-181-lllmstxt|> ## Data-Centric Computer Vision [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/0056db2d46f0d1f7c6394d430f2109e3b9eaeb46-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured Your Data, Your Advantage [Read article](https://voxel51.com/whitepapers/your-data-your-advantage) ![](https://cdn.sanity.io/images/h6toihm1/production/552aa52a004e7137aba09ab689db8bccbc12809d-3840x2160.jpg?auto=format&dpr=2&fit=max&q=75&w=960) Featured Auto-Labeling Data for Object Detection [Read article](https://voxel51.com/whitepapers/auto-labeling-data-for-object-detection) ![](https://cdn.sanity.io/images/h6toihm1/production/50a686dd34d9caecdff5e2c62af0814b83710d44-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured The Best Data-Centric Computer Vision Tools for the Enterprise [Read article](https://voxel51.com/whitepapers/the-best-data-centric-computer-vision-tools-for-the-enterprise) ![](https://cdn.sanity.io/images/h6toihm1/production/0056db2d46f0d1f7c6394d430f2109e3b9eaeb46-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured Your Data, Your Advantage [Read article](https://voxel51.com/whitepapers/your-data-your-advantage) ![](https://cdn.sanity.io/images/h6toihm1/production/552aa52a004e7137aba09ab689db8bccbc12809d-3840x2160.jpg?auto=format&dpr=2&fit=max&q=75&w=960) Featured Auto-Labeling Data for Object Detection [Read article](https://voxel51.com/whitepapers/auto-labeling-data-for-object-detection) ![](https://cdn.sanity.io/images/h6toihm1/production/50a686dd34d9caecdff5e2c62af0814b83710d44-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=960) Featured The Best Data-Centric Computer Vision Tools for the Enterprise [Read article](https://voxel51.com/whitepapers/the-best-data-centric-computer-vision-tools-for-the-enterprise) All whitepapers [![](https://cdn.sanity.io/images/h6toihm1/production/2b9cf0dc8ac97d8cfb20b412b7eec7435249ea1f-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual AI for Defect Detection in Manufacturing\\ \\ Why visual inspection AI fails on the factory floor—and how to fix it. Learn how top manufacturers build reliable, scalable defect detection systems.](https://voxel51.com/whitepapers/visual-ai-for-defect-detection-in-manufacturing) [![](https://cdn.sanity.io/images/h6toihm1/production/60424cea21d1d2754c2e14e6e80f55bc6bac1447-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Vision AI Models Fails\\ \\ Why vision AI models fail — and how to prevent it. Learn data-centric strategies to catch labeling errors, bias, and drift before they hit production.](https://voxel51.com/whitepapers/why-vision-ai-models-fails) [![](https://cdn.sanity.io/images/h6toihm1/production/50a686dd34d9caecdff5e2c62af0814b83710d44-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Best Data-Centric Computer Vision Tools for the Enterprise\\ \\ Explore the top data-centric CV platforms for scalable, secure, and cost-efficient model development. Compare tools, tradeoffs, and best practices.](https://voxel51.com/whitepapers/the-best-data-centric-computer-vision-tools-for-the-enterprise) [![](https://cdn.sanity.io/images/h6toihm1/production/0056db2d46f0d1f7c6394d430f2109e3b9eaeb46-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Your Data, Your Advantage\\ \\ Reclaim control of your AI data. Learn how in-house auto-labeling boosts security, cuts costs, and protects your competitive edge.](https://voxel51.com/whitepapers/your-data-your-advantage) [![](https://cdn.sanity.io/images/h6toihm1/production/552aa52a004e7137aba09ab689db8bccbc12809d-3840x2160.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Auto-Labeling Data for Object Detection\\ \\ How far can zero-shot auto labeling take us? To find out, we rigorously benchmarked Verified Auto Labeling across popular computer vision datasets.](https://voxel51.com/whitepapers/auto-labeling-data-for-object-detection) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) Computer Vision & Visual AI Whitepapers - Voxel51 <|firecrawl-page-182-lllmstxt|> ## FiftyOne and Open Images [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Press](https://voxel51.com/blog/category/press) Voxel51 & Google Collaborate to Make Downloading & Visualizing Open Images A Breeze May 13, 2021 • 3 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) VOXEL51’S OPEN SOURCE TOOL “FIFTYONE” ENABLES RAPID DATASET ANALYSIS TO IMPROVE QUALITY, ACCURACY AND DIVERSITY OF COMPUTER VISION DATASETS & MODEL PERFORMANCE **Ann Arbor, MI** – [Voxel51](https://voxel51.com/) today announced a collaboration with Google to support Google’s Open Images Dataset, one of the largest visual datasets in the world used by AI researchers and the machine learning community for common object detection and other computer vision tasks. Through the collaboration, Open Images users will gain access to Voxel51’s free, open source machine learning developer tool, “FiftyOne,” to more efficiently and effectively load, visualize and evaluate datasets from Open Images’ annotated collection of over 9 million images. Machine learning’s massive rise in adoption is powered by the widespread availability of large, manually labeled datasets, but what truly makes this data useful in complex, real-world applications, is the ability to easily and rapidly analyze and experiment with this data at scale. Starting today, when users visit the [Open Images download page](https://storage.googleapis.com/openimages/web/download.html), they will be directed to FiftyOne, where they can easily download all or part of Open Images, visualize its rich annotations and evaluate models with the official open Images protocol. FiftyOne enables users to rapidly analyze and improve the quality, accuracy and diversity of computer vision datasets in order to improve model performance. Users can also select specific dataset subsets, types of annotations and classes of objects to download. Google reached out to Voxel51 about the dataset analysis capabilities of FiftyOne following an error analysis study that Voxel51 performed on Google’s Open Images Dataset. In a [blog post](https://towardsdatascience.com/i-performed-error-analysis-on-open-images-and-now-i-have-trust-issues-89080e03ba09), Voxel51 detailed how they utilized FiftyOne’s powerful visualization and model analysis features to observe patterns in dataset errors that frequently stifle model performance. In this case, even though the object detection ground truth in Open Images [is 98% correct](https://arxiv.org/abs/1811.00982), nearly a third of the “false positive errors” among modern models were due to the small amount of ground-truth imperfections rather than model errors. “FiftyOne enables researchers to analyze and improve the quality of their datasets rapidly, replacing the weeks of manual labor that would otherwise be required without this technology,” said [Jordi Pont-Tuset](http://jponttuset.cat/), Research Scientist at Google. “High-quality data is critical to the success of machine learning systems. Without the right tools to analyze and curate datasets, machine learning development can be inefficient and ineffective.” “In order to push the frontier of what’s possible in machine learning, model output must be closely analyzed in conjunction with the datasets in a detailed manner,” said Jason Corso, co-founder and CEO of Voxel51 and director of the Stevens Institute for Artificial Intelligence. “The classical aggregate analysis methods most commonly used in machine learning are only part of what is necessary to develop high performing datasets and models. We created FiftyOne to fill a gap in developer tools for machine learning and we’re proud to be empowering others to build better datasets and to train better models through this collaboration with Google.” Since launching in August 2020, FiftyOne has transformed the machine learning and computer vision development lifecycle. FiftyOne provides a flexible and open core software ecosystem to build machine learning workflows, with an emphasis on the important role that high quality data plays at every stage of the workflow. Among its state-of-the-art features are a novel query language to easily search and filter images and videos based on their content, a user-friendly web-based application for visualizing data, access to powerful tools like embeddings that can uncover hidden patterns and automatically label data, and the flexibility to work locally, remotely, in the cloud, or in Jupyter/Colab notebooks. FiftyOne is available at [fiftyone.ai](https://voxel51.com/docs/fiftyone/). **About Voxel51** Headquartered in Ann Arbor, Michigan, and founded in 2016 by Dr. Jason Corso and Dr. Brian Moore, Voxel51 is an AI software company that is democratizing access to software 2.0 by providing the open core software building blocks that enable computer vision and machine learning engineers to rapidly engineer data-powered workflows. Learn more at [voxel51.com](https://voxel51.com/). ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/62b79d1d13bfc9926b5f560049adc33fd5d2cbf3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 Launches Computer Vision Industry’s First Open-Source Rapid Dataset Experimentation Tool\\ \\ Press\\ \\ • \\ \\ Aug 12, 2020](https://voxel51.com/blog/fiftyone-open-source-launch) [![](https://cdn.sanity.io/images/h6toihm1/production/b2822f56ac527e57cfa5a5e6dbdd2fea96ae2817-1934x1110.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51’s Coronavirus Physical Distancing Index Tracks Reaction to Social Distancing Around the World\\ \\ Press\\ \\ • \\ \\ Apr 1, 2020](https://voxel51.com/blog/voxel51-physical-distancing-index) [Voxel51 Launches Image-To-Video Tool Enabling Models Trained on Images to Automatically Process Video\\ \\ Press\\ \\ • \\ \\ Oct 3, 2019](https://voxel51.com/blog/voxel51-launches-image-to-video-tool) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-183-lllmstxt|> ## FiftyOne Open-Source Launch [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Press](https://voxel51.com/blog/category/press) Voxel51 Launches Computer Vision Industry’s First Open-Source Rapid Dataset Experimentation Tool Aug 12, 2020 • 2 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) #### “FIFTYONE” HELPS DATA SCIENTISTS TACKLE DATA QUALITY LIMITATIONS THAT IMPACT PREDICTIVE PERFORMANCE OF IMAGE-BASED MACHINE LEARNING MODELS **Ann Arbor, MI** – [Voxel51](https://voxel51.com/) today announced the launch of [FiftyOne](https://voxel51.com/fiftyone/), the computer vision industry’s first open-source rapid dataset experimentation tool that addresses the most fundamental pain point for machine learning scientists — performance-limiting datasets. “Nothing hinders the success of machine learning systems more than poor-quality data. Yet the process of continuous data quality management is incredibly challenging and time consuming,” said Jason Corso, Voxel51 co-founder and CEO. “We created this tool, which brings over 15 years of academic research and experience in creating computer vision and machine learning systems to offer engineers a better toolbox and a more efficient way to improve the quality, accuracy and diversity of image datasets in order to mitigate the consequences of bad data and to improve the predictive performance of production models.” Available now at [voxel51.com](https://voxel51.com/), the new transformative tool enables computer vision and machine learning scientists to easily bring their insight and intuition into their training data and to shrink the time-consuming and labor-intensive process of constant iteration and evaluation. FiftyOne’s interactive dashboard and Python library allows users to easily visualize, explore and analyze their data to curate superior datasets with the ability to scale to the size and level of accuracy required in order to be useful in complex, real-world applications. “There’s no other solution that’s as easy to use to rapidly experiment with datasets and to identify limitations that stifle model performance,” said Brian Moore, Voxel51 co-founder and CTO. “We’re transforming the art of dataset curation into a science and we are eager to share the results of our work with developers and machine learning scientists across industry and academia.” FiftyOne makes it easy to search, filter and sort images by their predicted and ground truth classifications as well as their related hardness, mistakenness, and uniqueness. FiftyOne supports the most commonly used image dataset types such as Berkeley DeepDrive, COCO, CVAT, KITTI, TFRecords and VOC. A collection of open source datasets and support for TensorFlow and PyTorch ML frameworks are also built into FiftyOne along with helper functions and tutorials. Users are able to easily view subsets of dataset samples; remove duplicate and near-duplicate images; cleanup labeling mistakes; recommend samples for annotation; rank samples by representativeness; and discover the most unique samples, which is essential to improving model performance. FiftyOne is flexible and scalable, allowing natural integration into existing workflows. **About Voxel51** Headquartered in Ann Arbor, Michigan, and founded in 2016 by University of Michigan professor Dr. Jason Corso and Dr. Brian Moore, Voxel51 is an AI software company that enables computer vision data scientists to rapidly curate and experiment with their datasets in order to build higher performing machine learning systems. For more information visit [voxel51.com](https://voxel51.com/). ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0a21c6ca2af5253f72f6b88f9d8f6dafa5fb4ab3-2560x1390.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51 & Google Collaborate to Make Downloading & Visualizing Open Images A Breeze\\ \\ Press\\ \\ • \\ \\ May 13, 2021](https://voxel51.com/blog/fiftyone-open-images-collaboration) [![](https://cdn.sanity.io/images/h6toihm1/production/b2822f56ac527e57cfa5a5e6dbdd2fea96ae2817-1934x1110.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51’s Coronavirus Physical Distancing Index Tracks Reaction to Social Distancing Around the World\\ \\ Press\\ \\ • \\ \\ Apr 1, 2020](https://voxel51.com/blog/voxel51-physical-distancing-index) [Voxel51 Launches Image-To-Video Tool Enabling Models Trained on Images to Automatically Process Video\\ \\ Press\\ \\ • \\ \\ Oct 3, 2019](https://voxel51.com/blog/voxel51-launches-image-to-video-tool) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-184-lllmstxt|> ## Physical Distancing Index [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Press](https://voxel51.com/blog/category/press) Voxel51’s Coronavirus Physical Distancing Index Tracks Reaction to Social Distancing Around the World Apr 1, 2020 • 2 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) #### INTERACTIVE AI-POWERED TOOL ALLOWS USERS TO WATCH LIVE STREET CAMS AND EXPLORE HOW COVID-19 HAS CHANGED SOCIAL BEHAVIOR IN REAL TIME **Ann Arbor, MI** – Video data scientists at AI startup Voxel51, have developed The Voxel51 [Physical Distancing Index (PDI)](https://voxel51.com/press/voxel51-physical-distancing-index/) to help track how COVID-19 and preventative measures to contain its spread have impacted human activity around the globe in real time. Launched today, the free AI-powered interactive tool enables users to explore a day-by-day timeline of social activity in some of the world’s most popular public spaces. The PDI score helps people understand and compare how the coronavirus is changing social behaviors over time and enables municipalities to visualize how they’re doing from a public health perspective. Voxel51’s video analysis technology currently gathers and analyzes historical and real time video streams from public street cameras in London, New York City, Fort Lauderdale, New Jersey, Prague and Dublin, Ireland. The company’s cutting-edge computer vision and deep learning models are able to detect the number of pedestrians, vehicles, and other human-centric objects in each live stream. A PDI score or the average amount of human activity over the previous 24 hours is automatically calculated every 15 minutes to track the change over time. PDI scores are represented in each time-stamped data point on the chart along with relevant COVID-19 news headlines from the day, and a snapshot image from the live video scene. Users can also compare scores across locations to see the differences in PDI relative to each city over time. The PDI dashboard is updated in real-time and will continue to integrate live street cams from additional cities around the world. The PDI is non-invasive and does not extract any identifying information about the individuals in the video, or use other data sources like mobile phone signals, which allow only approximate location estimates. Instead, Voxel51 focuses on specific centralized locations of interest in each city and watches the changes in public social activity in real time. “The coronavirus has had an intense impact on daily life across the globe, and we wanted to find a way to use our technology as a tool for public awareness,” said Jason Corso, Voxel51 co-founder and CEO. “We were able to apply our computer vision methods to compare how people are responding to physical distancing measures around the world and to study how exposure rates correlate with news on the coronavirus and physical distancing mandates over time.” “The data in the graphs speaks for itself. Coronavirus and the necessary preventative measures across the globe have had an intense impact on daily life,” said Brian Moore, Voxel51 co-founder and CTO. “All of the cities we’re tracking have seen a sharp decline in public activity during the month of March.” **About Voxel51** Headquartered in Ann Arbor, Michigan, and founded in 2016 by University of Michigan professor Dr. Jason Corso and Dr. Brian Moore, Voxel51 is a video-first AI company that provides the leading video understanding platform to transform raw video and images into actionable intelligence. Voxel51’s AI technology brings state-of-the-art deep learning models to users, enabling them to make real-time, data-driven decisions across the computer vision system lifecycles. For more information visit [voxel51.com](https://voxel51.com/). ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0a21c6ca2af5253f72f6b88f9d8f6dafa5fb4ab3-2560x1390.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51 & Google Collaborate to Make Downloading & Visualizing Open Images A Breeze\\ \\ Press\\ \\ • \\ \\ May 13, 2021](https://voxel51.com/blog/fiftyone-open-images-collaboration) [![](https://cdn.sanity.io/images/h6toihm1/production/62b79d1d13bfc9926b5f560049adc33fd5d2cbf3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 Launches Computer Vision Industry’s First Open-Source Rapid Dataset Experimentation Tool\\ \\ Press\\ \\ • \\ \\ Aug 12, 2020](https://voxel51.com/blog/fiftyone-open-source-launch) [Voxel51 Launches Image-To-Video Tool Enabling Models Trained on Images to Automatically Process Video\\ \\ Press\\ \\ • \\ \\ Oct 3, 2019](https://voxel51.com/blog/voxel51-launches-image-to-video-tool) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-185-lllmstxt|> ## Image-To-Video Tool Launch [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Press](https://voxel51.com/blog/category/press) Voxel51 Launches Image-To-Video Tool Enabling Models Trained on Images to Automatically Process Video Oct 3, 2019 • 2 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) #### VOXEL51 SELECTED AS AI TOP PICK AT TECHCRUNCH DISRUPT 2019 **San Francisco, CA** – Today from TechCrunch Disrupt San Francisco, Voxel51, creators of the pioneering video understanding platform that transforms raw video into actionable intelligence for security, automotive and smart city applications, announced the release of its Image-To-Video tool. The first-of-its-kind tool enables computer vision teams to drop their proprietary image-based analytics models into the tool in order to rapidly process and analyze video at scale. Uniquely built to process video, Voxel51’s computer vision and spatio-temporal deep learning models automatically identify and classify objects, actions, patterns and behaviors in video scenes with incredible accuracy. The scalable platform enables customers with large video datasets to tag, search and integrate detected content into their workflows and human-decision making processes with great speed and accuracy. Pilot programs utilizing Voxel51’s Image-To-Video Tool are currently underway. “Until now, many computer vision researchers have focused on images rather than video because video is so much more difficult and time consuming to work with,” said Jason Corso, Voxel51 co-founder and CEO. “With the launch of our Image-To-Video tool on our platform, developers can take their existing computer vision models that were trained on images and automatically deploy them to process large volumes of video with minimal effort. This is going to save companies a great deal of time and money.” Voxel51 was recently selected by TechCrunch for its “TC Top Picks program” and has been recognized as one of five outstanding early-stage startups in the AI/Machine Learning category. Voxel51 co-founders Jason Corso and Brian Moore will present at TechCrunch Disrupt 2019 on Wednesday, Oct. 2 at 10 a.m., and reporters are invited to stop by the Voxel51 booth C17 in Startup Alley at TechCrunch Disrupt for a live demo. **About Voxel51** Headquartered in Ann Arbor, Michigan, and founded in 2016 by University of Michigan professor Dr. Jason Corso and Dr. Brian Moore, Voxel51 is a video-first AI company that provides the leading video understanding platform to transform raw video and images into actionable intelligence. Voxel51’s AI technology enables users to make real-time, data-driven decisions that improve public safety, security, and smart city applications. For more information visit [voxel51.com](https://voxel51.com/). ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0a21c6ca2af5253f72f6b88f9d8f6dafa5fb4ab3-2560x1390.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51 & Google Collaborate to Make Downloading & Visualizing Open Images A Breeze\\ \\ Press\\ \\ • \\ \\ May 13, 2021](https://voxel51.com/blog/fiftyone-open-images-collaboration) [![](https://cdn.sanity.io/images/h6toihm1/production/62b79d1d13bfc9926b5f560049adc33fd5d2cbf3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 Launches Computer Vision Industry’s First Open-Source Rapid Dataset Experimentation Tool\\ \\ Press\\ \\ • \\ \\ Aug 12, 2020](https://voxel51.com/blog/fiftyone-open-source-launch) [![](https://cdn.sanity.io/images/h6toihm1/production/b2822f56ac527e57cfa5a5e6dbdd2fea96ae2817-1934x1110.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51’s Coronavirus Physical Distancing Index Tracks Reaction to Social Distancing Around the World\\ \\ Press\\ \\ • \\ \\ Apr 1, 2020](https://voxel51.com/blog/voxel51-physical-distancing-index) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-186-lllmstxt|> ## Voxel51 Funding Announcement [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Press](https://voxel51.com/blog/category/press) Voxel51 Raises $2 Million To Advance Video Understanding Aug 8, 2019 • 2 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) #### AI FOR VIDEO PIONEER LAUNCHES VIDEO ANALYTICS PLATFORM FOR AUTOMOTIVE, SMART CITY AND SECURITY APPLICATION **Ann Arbor, MI** – Voxel51, the AI startup revolutionizing video understanding with its video-first solution that unlocks valuable intelligence from videos to make real-time data-driven decisions, today announced that it has closed a $2 million seed round from [eLab Ventures](http://elabvc.com/). The company also today announced that it has launched its video understanding platform for automotive, smart city and security applications. Most AI solutions used to analyze video are designed for processing static images or individual video frames, making it difficult to classify behaviors and actions such as walking or running. Uniquely built to process video, Voxel51’s computer vision and spatio-temporal deep learning models automatically identify and classify objects, actions, patterns and behaviors in video scenes with incredible accuracy. The scalable platform enables customers with large video datasets to tag, search and integrate detected content into their workflows and human-decision making processes. “While advancements in AI and computer vision have enabled object detection in images with pinpoint accuracy, video comprehension is still in its infancy,” said Jason Corso, Voxel51 co-founder and CEO. “We are already delivering human-level results for video, and this funding allows us to push the envelope even further. We will continue to execute on our vision to solve the complexities of video understanding and to improve the speed, accuracy and efficiency of large scale video processing and analysis.” Since the company was founded in December, 2016, it has completed a number of successful pilots in the automotive, smart city and security spaces. Voxel51 has also recently concluded a $1.25 million grant awarded by the U.S. Commerce Department’s National Institute of Standards and Technology (NIST) to develop video analytics for public safety, bringing its total investment to $3.25 million. **About Voxel51** Headquartered in Ann Arbor, Michigan, and founded in 2016 by University of Michigan professor Dr. Jason Corso and Dr. Brian Moore, Voxel51 is a video-first AI company that provides the leading video understanding platform to transform raw video and images into actionable intelligence. Voxel51’s AI technology enables users to make real-time, data-driven decisions that improve public safety, security, and smart city applications. For more information visit [voxel51.com](https://voxel51.com/). ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0a21c6ca2af5253f72f6b88f9d8f6dafa5fb4ab3-2560x1390.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51 & Google Collaborate to Make Downloading & Visualizing Open Images A Breeze\\ \\ Press\\ \\ • \\ \\ May 13, 2021](https://voxel51.com/blog/fiftyone-open-images-collaboration) [![](https://cdn.sanity.io/images/h6toihm1/production/62b79d1d13bfc9926b5f560049adc33fd5d2cbf3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 Launches Computer Vision Industry’s First Open-Source Rapid Dataset Experimentation Tool\\ \\ Press\\ \\ • \\ \\ Aug 12, 2020](https://voxel51.com/blog/fiftyone-open-source-launch) [![](https://cdn.sanity.io/images/h6toihm1/production/b2822f56ac527e57cfa5a5e6dbdd2fea96ae2817-1934x1110.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51’s Coronavirus Physical Distancing Index Tracks Reaction to Social Distancing Around the World\\ \\ Press\\ \\ • \\ \\ Apr 1, 2020](https://voxel51.com/blog/voxel51-physical-distancing-index) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-187-lllmstxt|> ## Introducing VoxelGPT [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Press](https://voxel51.com/blog/category/press) Introducing VoxelGPT: AI-Powered Computer Vision Insights Delivered Through Chat Jun 8, 2023 • 4 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) #### _Game-changing natural language plugin makes complex AI tasks accessible to everyone_ ANN ARBOR, Mich., June 7, 2023 /PRNewswire/ — [Voxel51](https://voxel51.com/), a leading innovator in data-centric computer vision and machine learning software, today unveiled an extraordinary breakthrough in computer vision: [VoxelGPT](https://voxel51.com/voxelgpt/). An extension of [FiftyOne](https://voxel51.com/fiftyone/), the world’s most prolific open source computer vision toolkit with more than one million downloads, VoxelGPT enables computer vision engineers, researchers, and the organizations they work for to curate high quality datasets, build high performing models, and move AI projects from proof-of-concept to viable and in production faster than ever before. VoxelGPT, along with the full FiftyOne experience, is available to try instantly in your browser at [gpt.fiftyone.ai](https://gpt.fiftyone.ai/). VoxelGPT brings together the power of GPT-3.5 with FiftyOne’s flexible computer vision query language. It turns natural language queries into actual Python code that can filter, sort, semantically slice, and reveal insights about the images and videos in your dataset, all without having to write a single line of code. While no code and low code solutions are nothing new, they can often force a one-size-fits-all approach to AI architecture and limit your ability to integrate with your favorite tools and libraries. VoxelGPT combines the simplicity of no code solutions with the flexibility of FiftyOne’s advanced querying and visualization, enabling you to do computer vision your way, but faster. “Bringing safe, smart, and high performing AI systems to production requires scientists and engineers to understand the many complexities of their datasets and models,” said Jason Corso, Co-Founder & CEO of Voxel51. “FiftyOne’s semantic slicing query language enables these scientists and engineers to rapidly and effectively find corner cases and failure modes in their system. With VoxelGPT, not only can they do so with natural language queries, but even users who do not know the FiftyOne Python query language can now tap into this analysis capability.” “At ADT Commercial, we’re actively developing socially responsible AI technology to create safe and secure commercial environments. We work with large-scale datasets, often numbering in the millions of samples, coming from multiple camera feeds in real time,” said Philippe Sawaya, Director of Artificial Intelligence, ADT Commercial. “Although data sits at the heart of AI, sifting through all that data to reveal critical insights is no trivial task. We’re thrilled to put VoxelGPT into the hands of our AI teams to make their lives easier and increase the value our AI solutions bring to customers.” VoxelGPT makes it simple to use natural language to create complex queries, saving time and money, while streamlining the way you work. Key capabilities include: - **Search computer vision datasets**: Ask VoxelGPT anything and it will search your datasets for you and return the results. For example: retrieve me 10 random samples, display the most unique images with a false positive prediction, and just show the images with at least two people. - **Ask computer vision, machine learning, and data science questions**: VoxelGPT is not only a Python programmer-at-your-fingertips, it’s also an educational resource to help you understand basic concepts and how to overcome data quality issues. For example: what is the difference between precision and recall?, how can I detect faces in my images?, and what are some ways I can reduce redundancy in my dataset? - **Search documentation, API specifications, and tutorials**: VoxelGPT has access to the entire collection of [FiftyOne documentation](https://docs.voxel51.com/), which can be used to quickly answer FiftyOne-related questions. For example: how do I load my custom dataset into FiftyOne?, how can I export my dataset in COCO format?, and how can I generate a 2D image for a point cloud? 100% open source under the Apache 2.0 license, VoxelGPT is free and easy to use. Try it instantly in your browser at [gpt.fiftyone.ai](https://gpt.fiftyone.ai/). To use VoxelGPT in the FiftyOne App, simply install it as a plugin. Or, install it locally to use alongside the open source FiftyOne library. [Learn more about VoxelGPT](https://voxel51.com/voxelgpt/), or [get started with VoxelGPT on GitHub](https://github.com/voxel51/voxelgpt). **About Voxel51** Voxel51 is bringing transparency and clarity to the world’s data. Our open source and commercial software enables developers, scientists, and organizations to build high-quality datasets and computer vision models that power some of today’s most remarkable machine learning and artificial intelligence. Tens of thousands of engineers and scientists have integrated open source FiftyOne into their ML workflows. Enterprise customers spanning verticals like automotive, robotics, security, retail, and healthcare rely on FiftyOne Teams to securely collaborate on their datasets and models. We’re building a fully-remote team of exceptional and diverse people who want to bring data-centric AI to the world. To learn more, visit [voxel51.com](https://voxel51.com/). _Originally published on [PRNewswire](https://www.prnewswire.com/news-releases/introducing-voxelgpt-ai-powered-computer-vision-insights-delivered-through-chat-301844496.html) on June 7, 2023_ ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0a21c6ca2af5253f72f6b88f9d8f6dafa5fb4ab3-2560x1390.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51 & Google Collaborate to Make Downloading & Visualizing Open Images A Breeze\\ \\ Press\\ \\ • \\ \\ May 13, 2021](https://voxel51.com/blog/fiftyone-open-images-collaboration) [![](https://cdn.sanity.io/images/h6toihm1/production/62b79d1d13bfc9926b5f560049adc33fd5d2cbf3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 Launches Computer Vision Industry’s First Open-Source Rapid Dataset Experimentation Tool\\ \\ Press\\ \\ • \\ \\ Aug 12, 2020](https://voxel51.com/blog/fiftyone-open-source-launch) [![](https://cdn.sanity.io/images/h6toihm1/production/b2822f56ac527e57cfa5a5e6dbdd2fea96ae2817-1934x1110.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51’s Coronavirus Physical Distancing Index Tracks Reaction to Social Distancing Around the World\\ \\ Press\\ \\ • \\ \\ Apr 1, 2020](https://voxel51.com/blog/voxel51-physical-distancing-index) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-188-lllmstxt|> ## FiftyOne Open Source Launch [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Press](https://voxel51.com/blog/category/press) Voxel51 Launches FiftyOne Open Source 1.0, Accelerating the Creation of Production-Ready Visual AI Applications Oct 1, 2024 • 4 min read Article content In this article In this article ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _Popular open source software with over 2M installs marks significant milestone enabling the successful development of visual AI projects_ **ANN ARBOR, Mich., October 1, 2024** – [Voxel51](http://www.voxel51.com/), the leading solution for building visual AI applications, today announced the milestone [release of version 1.0](https://voxel51.com/blog/announcing-fiftyone-1-0/) of open source [FiftyOne](https://voxel51.com/fiftyone/). FiftyOne Open Source provides the foundation for Voxel51’s commercial solution, FiftyOne Teams, together bringing unparalleled ease and efficiency to how visual data is used to develop robust and reliable visual AI applications. Visual data, which already makes up over [60% of all data traffic](https://www.prnewswire.com/news-releases/sandvines-2023-global-internet-phenomena-report-shows-24-jump-in-video-traffic-with-netflix-volume-overtaking-youtube-301723445.html), is critical to building innovative AI applications that understand and interact with the real world. However, AI builders find themselves without tools capable of handling the ever-expanding volumes and diverse modalities of visual data and putting it to use in developing their AI models and applications. Existing fragmented and inflexible tools result in complex, unwieldy, and often manual processes that lead to delays and [failures in over 80%](https://hbr.org/2023/11/keep-your-ai-projects-on-track) of AI projects. ### Integrated Solution for Visual AI Builders FiftyOne was developed by Voxel51’s team of computer vision and machine learning experts to reinvent how visual data is used in developing AI models and applications. With this milestone open source release, FiftyOne delivers on the vision of an integrated solution that can extract valuable intelligence from diverse visual data and automate the iterative processes of using that data to build robust AI models. Exhibiting exceptional ease-of-use, flexibility, and extensibility essential to fit how AI builders work, the milestone provides comprehensive capabilities to develop and deliver visual AI applications successfully. This release includes: - Support for additional varieties and forms of visual data – from [3D meshes](https://voxel51.com/blog/announcing-fiftyone-0-24-with-3d-meshes-and-custom-workspaces/) and scenes to point clouds and geometries – throughout exploration, curation, and model development. - A powerful [customizable and extensible framework](https://voxel51.com/blog/announcing-fiftyone-0-25/) that enables building interactive data applications, trigger operations, and custom dashboards using Python that aid in fine-tuning models and datasets, while supporting developers in the way they build their visual AI applications. - Adding native [vector search integration with Elasticsearch](https://voxel51.com/vector-search/) to expand the ecosystem of tools that can be used seamlessly with FiftyOne. - Open sourcing of the core machine learning techniques for data and models in [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html) that help uncover hidden structure and relationships in visual data by algorithmically assessing uniqueness, similarity, representativeness, and more. “It’s impossible to develop reliable, trustworthy visual AI when you’re struggling to curate and understand millions of data samples and how to use them to build AI models that will succeed in real-world applications,” said Voxel51 co-founder and CEO Brian Moore. “This milestone release delivers the solid foundation for innovation in visual AI made possible when every AI builder has the ability to bring quality data and analysis to every step of their development process. FiftyOne’s ease of use, extensibility, and transparency are democratizing best practices for visual AI development, and we’re excited to see our thriving community continue to deliver new value that advances the ecosystem.” “FiftyOne provides us with a centralized place where we can easily understand data to uncover and resolve problems with our annotations and models. It has been so easy to build with FiftyOne given the numerous integrations and flexibility to dig into the data. No wonder it is quickly becoming the standard in the computer vision community,” said Chris Hall, Data Scientist, Vivint Smart Home. ### Harnessing the Power and Promise of Open Source AI The release of open source machine learning algorithms in FiftyOne Brain demonstrates [Voxel51’s continuing commitment to open source AI](https://voxel51.com/blog/power-of-open-source-ai-how-fiftyone-drives-innovation/). The AI ecosystem is at a pivotal crossroads, marked by numerous questions and opinions about the true meaning of open source AI — an uncertainty that could shape the future of this rapidly evolving field. Unlike approaches and vendors that rely on closed, black-box, or only partially open offerings, Voxel51 believes that open source and transparency are critical in all aspects of AI, from models to data to systems. “Open source makes possible the transparency, collaboration, and rigor needed to deliver continuing innovation in AI,” said Voxel51 co-founder and Chief Science Officer Jason Corso. “We are excited to continue investing in open source so that the entire AI ecosystem can collaborate to deliver better and more robust innovation that is reliable and trustworthy.” Tens of thousands of AI builders and their teams already rely on FiftyOne for a wide variety of use cases to deliver more accurate and robust models, improving team productivity by up to 50% and model accuracy by up to 30%. FiftyOne Teams, Voxel51’s commercial offering based on open source FiftyOne, is helping leading enterprises and organizations including LG Electronics, Berkshire Grey, and Precision Planting to make visual AI a reality. To learn more: - [Visit us online](https://voxel51.com/) to discover FiftyOne - [Try FiftyOne](https://voxel51.com/try-fiftyone/) in your browser - [Book a demo](https://voxel51.com/book-a-demo/) - [Join the community](https://voxel51.com/community-resources/) - Follow us on [LinkedIn](https://c212.net/c/link/?t=0&l=en&o=4168578-1&h=67554328&u=https%3A%2F%2Fwww.linkedin.com%2Fcompany%2Fvoxel51%2F&a=LinkedIn), [X](https://c212.net/c/link/?t=0&l=en&o=4168578-1&h=3107383252&u=https%3A%2F%2Ftwitter.com%2Fvoxel51&a=X), [Slack](https://c212.net/c/link/?t=0&l=en&o=4168578-1&h=1671336270&u=https%3A%2F%2Fslack.voxel51.com%2F&a=Slack) and [GitHub](https://c212.net/c/link/?t=0&l=en&o=4168578-1&h=3413436312&u=https%3A%2F%2Fgithub.com%2Fvoxel51%2Ffiftyone&a=GitHub) ### About Voxel51 Voxel51 is helping organizations make visual AI a reality. Our open source and commercial software enables teams to build high-quality datasets and computer vision models that power leading machine learning and artificial intelligence applications. Tens of thousands of engineers and scientists have integrated open source FiftyOne into their workflows, and enterprise customers spanning from automotive, robotics, security, and retail to healthcare rely on FiftyOne Teams to securely collaborate on datasets and models. We’re growing and hiring the industry’s best talent to build a diverse and exceptional team of people who want to bring visual AI to life. To learn more, visit [voxel51.com](https://voxel51.com/). ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0a21c6ca2af5253f72f6b88f9d8f6dafa5fb4ab3-2560x1390.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51 & Google Collaborate to Make Downloading & Visualizing Open Images A Breeze\\ \\ Press\\ \\ • \\ \\ May 13, 2021](https://voxel51.com/blog/fiftyone-open-images-collaboration) [![](https://cdn.sanity.io/images/h6toihm1/production/62b79d1d13bfc9926b5f560049adc33fd5d2cbf3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Voxel51 Launches Computer Vision Industry’s First Open-Source Rapid Dataset Experimentation Tool\\ \\ Press\\ \\ • \\ \\ Aug 12, 2020](https://voxel51.com/blog/fiftyone-open-source-launch) [![](https://cdn.sanity.io/images/h6toihm1/production/b2822f56ac527e57cfa5a5e6dbdd2fea96ae2817-1934x1110.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Voxel51’s Coronavirus Physical Distancing Index Tracks Reaction to Social Distancing Around the World\\ \\ Press\\ \\ • \\ \\ Apr 1, 2020](https://voxel51.com/blog/voxel51-physical-distancing-index) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-189-lllmstxt|> ## Visual Kinship Recognition [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) Visual Kinship Recognition with the Families in the Wild Computer Vision Dataset Dec 7, 2022 • 6 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#7186aa78d70a) [About the Families in the Wild dataset](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#29470a8f601c) [What is visual kinship recognition?](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#3353e106193b) [Dataset quick facts](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#beca231d9936) [Acknowledgements](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#76b469cd9e42) [First things first, install FiftyOne](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#ac05e685a43b) [Next, import the dataset](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#9fcf86db964e) [How are faces cropped?](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#80b0a033c44b) [Levels and label types](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#b9e605cb518e) [What’s next?](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#8bf1287eab21) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#7186aa78d70a) [About the Families in the Wild dataset](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#29470a8f601c) [What is visual kinship recognition?](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#3353e106193b) [Dataset quick facts](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#beca231d9936) [Acknowledgements](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#76b469cd9e42) [First things first, install FiftyOne](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#ac05e685a43b) [Next, import the dataset](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#9fcf86db964e) [How are faces cropped?](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#80b0a033c44b) [Levels and label types](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#b9e605cb518e) [What’s next?](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset#8bf1287eab21) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the latest installment of our ongoing blog series where we highlight a dataset from the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html)! FiftyOne provides a Dataset Zoo that contains a collection of common datasets that you can download and load into FiftyOne via a few simple commands. In this post, we explore the Families in the Wild dataset and some of its use cases. You can watch an abbreviated version of this blog here: https://www.youtube.com/watch?v=t2gLHM\_Dy0c ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop The FiftyOne Dataset Zoo comprises more than 30 datasets, with new datasets being added all the time! They cover a variety of use cases including: - Video - Images - Location - Point-cloud - Action-recognition - Classification - Detection - Segmentation - Relationships - and more! ## **About the Families in the Wild dataset** [Families in the Wild](https://web.northeastern.edu/smilelab/fiw/) (FIW) is a public benchmark for recognizing families via facial images. Developed and maintained by the [Synergetic Media Learning (SMILE](https://web.northeastern.edu/smilelab/)) Lab at Northeastern University, it is the largest and most comprehensive database available for kinship recognition. The dataset contains more than 26k images of 5k faces collected from almost 1k families. A unique Family ID (FID) is assigned per family, ranging from F0001-F1018 ( _**Note:** Some families in the dataset were merged or removed since it was first released in 2016)._ The Families in the Wild dataset includes the families of publicly recognizable people like Jimmy Fallon, Bob Dylan, Michael Jackson, Margaret Thatcher, and Bruce Lee. The images span a variety of styles and eras, from old black and white photos to modern color images. ## **What is visual kinship recognition?** This discipline within computer vision aims to: - Identify who is related and who is not related - Identify different family relationship types like mother, father, children, and sibling - Create fine-grain classification so family trees can be constructed that include uncles, aunts, cousins, grandparents, etc. Visual kinship recognition can be applied to use cases like: - **Search identification:** Identifying fugitives or criminal suspects - **Modern-day refugee crisis**: Reuniting lost family members who have been separated due to war, famine, or natural catastrophes - **Missing children:** Finding relatives of missing or exploited children - **Genealogy services and research:** Using instead of or in conjunction with DNA-based approaches As you can imagine, visual kinship recognition is a challenging problem to solve because of intra and inter-class variation, individuals that look the same but are not related, insufficient data distribution, and a lack of labeled data. The goal of the FIW project was to build a large-scale kinship dataset to address some of these challenges! ## **Dataset quick facts** - **Dataset Source:** [Hosted and maintained](https://web.northeastern.edu/smilelab/fiw/) by the SMILE Lab at Northeastern University - **License:** MIT - **Last Update:** March 18, 2022 - **GitHub:** [README and development kit](https://github.com/visionjo/fiw) - **Dataset Name in FiftyOne:** fiw - **Dataset Size:** 173.00 MB - **Tags:** image, kinship, verification, classification, search-and-retrieval, facial-recognition - **Supported Splits:** test, val, train - **ZooDataset class:** [FIWDataset](https://www.voxel51.com/docs/fiftyone/api/fiftyone.zoo.datasets.base.html#fiftyone.zoo.datasets.base.FIWDataset) **Note:** For convenience, FiftyOne provides `get_pairwise_labels()` and `get_identifier_filepaths_map()` utilities for FIW. ## **Acknowledgements** What more details about the dataset? For statistics, task evaluations, benchmarks, and more, please check out the following journal article by the FIW authors: Robinson, JP, M. Shao, and Y. Fu. [“Survey on the Analysis and Modeling of Visual Kinship: A Decade in the Making.”](https://arxiv.org/pdf/2111.00598.pdf) IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), 2021 There is also an in-depth [presentation on YouTube](https://www.youtube.com/watch?v=bFMTvd_Ov1I) by Joseph Robinson about the dataset and benchmark. Ok, let’s get started exploring! ## **First things first, install FiftyOne** If you don’t already have FiftyOne installed on your laptop, it takes just a few minutes! For example on MacOs: - Verify your version of Python - Create and activate a virtual environment - Upgrade your Setuptools - Install FiftyOne ![](https://cdn.sanity.io/images/h6toihm1/production/7cc1ee53acbfa0ea5d4564e0463204a9abab3c76-1728x1080.gif?auto=format&dpr=2&fit=max&q=75&w=1600) Learn more about how to [get up and running with FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html) in the Docs. ## **Next, import the dataset** Now that you are up and running, importing the dataset and launching the FiftyOne App takes just a few more lines of code: import fiftyone as fo import fiftyone.zoo as foz dataset = foz.load\_zoo\_dataset("fiw", split="test") session = fo.launch\_app(dataset) ![](https://cdn.sanity.io/images/h6toihm1/production/d65b847dc98d7484cce7c7dcc311d8fa52f4521f-1880x1093.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## **How are faces cropped?** Faces were cropped from imagery using the five-point face detector [MTCNN](https://github.com/ipazc/mtcnn) from various phototypes (i.e., mostly family photos, along with several profile pics of individuals (facial shots)). The number of members per family varies from 3-to-26, with the number of faces per subject ranging from 1 to >10. ## **Levels and label types** Various levels and types of labels are associated with samples in this dataset. Family-level labels contain a list of members, each assigned a member ID (MID) unique to that respective family. For example: F0870.MID2 refers to member 2 of family 870. ![](https://cdn.sanity.io/images/h6toihm1/production/1cf3133db0d3be1d2d4ea347c8d6f7bd49904f21-1718x805.png?auto=format&dpr=2&fit=max&q=75&w=1600) Each member has annotations specifying gender and relationship to all other members in that respective family. The relationships in FIW are: ![](https://cdn.sanity.io/images/h6toihm1/production/454d065e00de9781cf356f8bd141c3296af20c50-265x238.png?auto=format&dpr=2&fit=max&q=75&w=265) Viewed in the FiftyOne App: ![](https://cdn.sanity.io/images/h6toihm1/production/ddb89c36419724800d018f357b2fdee924808d32-290x731.png?auto=format&dpr=2&fit=max&q=75&w=290) Within FiftyOne, each sample corresponds to a single face image and contains primitive labels of the Family ID, Member ID, etc. The relationship labels are stored as multi-label [Classifications](https://voxel51.com/docs/fiftyone/api/fiftyone.core.labels.html?highlight=classifications#fiftyone.core.labels.Classifications), where each classification represents one relationship that the member has with another member in the family. The number of relationships will differ from one person to the next, but all faces of the same person will have the same relationship labels. ![](https://cdn.sanity.io/images/h6toihm1/production/78fbcbebd1ff265f75b7237a525370e48dfae658-1295x751.png?auto=format&dpr=2&fit=max&q=75&w=1295) Additionally, the labels for the [Kinship Verification task](https://competitions.codalab.org/competitions/21843) are also loaded into this dataset through FiftyOne. As with relationships, the kinship labels are stored as classifications. Whereas a relationship label only specifies a role for the associated person, such as “parent”, a kinship label specifies both sides, with fd representing a Father-Daughter kinship or md for Mother-Daughter. ![](https://cdn.sanity.io/images/h6toihm1/production/5e37d4e9b65be502210da6e575af7ea6fe5a22a3-1886x950.png?auto=format&dpr=2&fit=max&q=75&w=1600) In order to make it easier to browse the dataset in the FiftyOne App, each sample also contains a `face_id` field containing a unique integer for each face of a member, always starting at 0. This allows you to filter the `face_id` in the App to show only a single image of each person. ![](https://cdn.sanity.io/images/h6toihm1/production/75b6f74139a6f44ca45ceb94783e2d6aedc8b0b0-1890x1080.png?auto=format&dpr=2&fit=max&q=75&w=1600) For your reference, the relationship labels are stored on disk in a matrix that provides the relationship of each member with other members of the family as well as names and genders. The i-th row represents the i-th family member’s relationships to the other members. In other words, the entry in the i-th row and the j-th column corresponds to the i-th family member’s relationship to the j-th member of the family. For example, `FID0001.csv` contains: ![](https://cdn.sanity.io/images/h6toihm1/production/6c9243f6dd88f7cd64d46f8a0cade2e565b95640-371x104.png?auto=format&dpr=2&fit=max&q=75&w=371) This family has three members, as listed under the MID column (far-left). Each MID reads across its row. We can see that MID1 is related to MID2 by 4 -> 1 (Parent -> Child), which of course can be viewed as the inverse, i.e., MID2 -> MID1 is 1 -> 4. It can also be seen that MID1 and MID3 are spouses of one another, i.e., 5 -> 5. **Note:** The spouse label will likely be removed in a future version of this dataset as it serves no value to the problem of kinship. ## **What’s next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [Families in the Wild](https://voxel51.com/blog/tag/families-in-the-wild) [FIW](https://voxel51.com/blog/tag/fiw) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/33d08c7b16ab5bfa0e4c5a4f936be4592b8e0a90-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the Berkeley Deep Drive Autonomous Vehicle Dataset\\ \\ Datasets\\ \\ • \\ \\ Jan 11, 2023](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/129c3574861e6c307549e14106753af13ecfa2bf-1308x1044.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How to Download ActivityNet and Evaluate Video Understanding Models\\ \\ Datasets\\ \\ • \\ \\ Feb 8, 2022](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models) [![](https://cdn.sanity.io/images/h6toihm1/production/7d8209de12d18f950c73a8d2e335f4a83bc6ce71-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the UCF101 Dataset: A Large-Scale, YouTube-Based Action Recognition Dataset\\ \\ Datasets\\ \\ • \\ \\ Mar 1, 2023](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-190-lllmstxt|> ## FiftyOne 0.18 Webinar Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Webinar Recap: What’s New in FiftyOne 0.18 for Computer Vision Dec 6, 2022 • 13 min read Article content In this article [Donating $200 to World Literacy Foundation](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#2b526a1f2207) [Presentation Highlights](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#08e093b3444c) [Why FiftyOne and Voxel51?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#63112670e5a4) [What is FiftyOne?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#7581fe5119ba) [What’s new in the latest version (0.18) of FiftyOne?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#429c36e30663) [1 — App Performance Improvements](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#958bd6faa7b1) [2 — App Sidebar Modes](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#2cf772801a37) [3 — Custom Sidebar Groups](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#4bee1ee862a2) [4 — Custom Label Attributes](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#1d11248106dc) [5 — Storing Field Metadata](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#f03a3230402a) [6 — Light Mode](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#b4bc0fd726ba) [Q&A from the Webinar](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#35289b68ceb0) [What’s Next](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#6d59ad07c0b7) In this article [Donating $200 to World Literacy Foundation](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#2b526a1f2207) [Presentation Highlights](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#08e093b3444c) [Why FiftyOne and Voxel51?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#63112670e5a4) [What is FiftyOne?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#7581fe5119ba) [What’s new in the latest version (0.18) of FiftyOne?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#429c36e30663) [1 — App Performance Improvements](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#958bd6faa7b1) [2 — App Sidebar Modes](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#2cf772801a37) [3 — Custom Sidebar Groups](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#4bee1ee862a2) [4 — Custom Label Attributes](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#1d11248106dc) [5 — Storing Field Metadata](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#f03a3230402a) [6 — Light Mode](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#b4bc0fd726ba) [Q&A from the Webinar](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#35289b68ceb0) [What’s Next](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision#6d59ad07c0b7) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We recently released [FiftyOne 0.18](https://voxel51.com/docs/fiftyone/release-notes.html#fiftyone-0-18-0) with a lot of exciting new features to make it easier than ever to build high quality computer vision datasets and machine learning models! Voxel51 Co-Founder and CTO [Brian Moore](https://www.linkedin.com/in/brimoor/) walked us through all the new features in a live webinar, with plenty of live demos and code examples to show it in action. You can watch the playback on [YouTube](https://www.youtube.com/watch?v=DukuxznuRuk), take a look at the [slides](https://docs.google.com/presentation/d/1diRx5Ni7a6Tum4iaLQS0tliH5VsJObVj8IeZdvEFKLc/edit?usp=sharing), see the [full transcript](https://www.rev.com/transcript-editor/shared/kvOPkhg4T5bxNbDUnvQd0i7bz9K4PkGgks7CLwMtG2HuqdWRZwwqmloqCflpX72Ej5CQQmlfpC_yP_uCL9VX00AY5mU?loadFrom=SharedLink), and read the recap below for the highlights. Enjoy! https://www.youtube.com/watch?v=DukuxznuRuk ## Donating $200 to World Literacy Foundation In lieu of swag, we gave attendees the opportunity to vote for their favorite charity and help guide our monthly donation to charitable causes. The charity that received the highest number of votes for the third time in a row was the [World Literacy Foundation](https://worldliteracyfoundation.org/). We are pleased to be making another donation of $200 (totaling now $600 across three events!) to this wonderful organization on behalf of the FiftyOne community. ![](https://cdn.sanity.io/images/h6toihm1/production/6da3098b8f6eeb34e3674a5984affbc8a663de8e-300x300.png?auto=format&dpr=2&fit=max&q=75&w=300) ## Presentation Highlights ## Why FiftyOne and Voxel51? So much data; so little time. Brian says “visual data contains a lot of information. But if you don’t actually pull up those samples, look at them, and understand the data you have, the annotations that you’re generating, and the model predictions that you’re getting, you won’t really understand the performance of your model.” While you might think the large majority of time data scientists and ML engineers spend developing a computer vision model is centered around the model itself, it’s not; instead, it’s spent on wrangling data. And, as Brian adds, “as datasets continue to increase in size, that proportion is only going to increase.” Data quality matters. It’s no secret that low quality data leads to problems such as model bias, physical danger, and reduced model performance. Brian notes “if you train on incorrect data, your model will learn incorrect things.” So the goal is to efficiently improve the quality of your datasets. The key to doing so is to get hands-on with your data to achieve high quality datasets on which to train your models. That’s why [FiftyOne](https://github.com/voxel51/fiftyone) and our company [Voxel51](https://voxel51.com/) exist! We’re focused on improving the transparency and clarity of your data, helping you understand what’s in your data, and how you can improve the quality of your datasets. ## What is FiftyOne? The open source FiftyOne project is designed to be the computer vision tool that brings together your data-centric tasks and your model-centric tasks. Data-centric tasks include the way you get your data annotated, and the way you find and store your data; model-centric tasks include the way you train your models, the frameworks you use, and the tools you use to track your experiments. FiftyOne joins those two worlds and helps you iterate on datasets and models together. ![](https://cdn.sanity.io/images/h6toihm1/production/9bceb99b460104346521abd05fcecc3fea612369-1774x732.png?auto=format&dpr=2&fit=max&q=75&w=1600) FiftyOne is meant for all things computer vision and supports tasks such as classification, detection, segmentation, polylines, and key points on images, video, and 3D data. It has a very powerful API that’s designed to help you organize your data and query it. A recent blog post came out comparing the way that you can filter your data in FiftyOne to the way that you can with tabular data in pandas. Brian shares: “So one analogy to think of is that [FiftyOne can be the pandas for computer vision](https://medium.com/voxel51/why-fiftyone-is-the-pandas-of-computer-vision-87618c3f1c3).” ![](https://cdn.sanity.io/images/h6toihm1/production/fdc71b7238ca4c84b7c796906466854d08851a9a-1186x374.png?auto=format&dpr=2&fit=max&q=75&w=1186) ## What’s new in the latest version (0.18) of FiftyOne? We recently released FiftyOne 0.18 and there were a number of new goodies that were added based on community input, including: - 1 — Significant performance improvements to the FiftyOne App - 2 — New sidebar modes to further optimize user experience for large datasets - 3 — The ability to programmatically modify your dataset’s sidebar groups - 4 — The ability to declare and filter by custom label attributes - 5 — Support for storing and viewing field metadata in the App - 6 — A new light mode option Brian did a show and tell of each, and why they matter. In the rest of this post, we’ll briefly explore the awesomeness of each one. ## 1 — App Performance Improvements It is of utmost importance for us to make FiftyOne as fast as possible and for as large of datasets as possible. With that in mind, FiftyOne 0.18 brings significant improvements in performance to the FiftyOne App. In the presentation, Brian compares a search across 180,000 objects in FiftyOne 0.17 vs the same search happening in 0.18 and sees a 10x improvement in the App! Underlying these performance wins are the new sidebar modes that substantially decrease the amount of load on the database enabling the system to load far faster. We cover the new sidebar modes in the next section. But even without making use of the new sidebar modes (aka using the classic “All” mode, which was previously the only mode), performance overall is still significantly faster, so there are performance enhancements across the board. Brian shows this in a .gif in the webinar playback [starting at 7:47](https://www.youtube.com/watch?v=DukuxznuRuk&t=467s)–8:35. ## 2 — App Sidebar Modes With FiftyOne 0.18, the App sidebar now has three different modes of operation — All, Fast, and Best. These modes control how many statistics are computed in the sidebar when you update your dataset or update the current dataset view. For smaller datasets, the default is All. However, if you have a very large dataset, FiftyOne will automatically default to Fast mode, where all of the statistics will not be computed by default, but rather they will compute when you expand the filter tray for one of those attributes. Why is this awesome? Choosing the right mode for the dataset makes the App experience much faster because now it’s only computing what you ask it to, which means you can get ahold of the underlying data and start looking at it in the App grid much faster. The best news (no pun intended)? You don’t have to change anything about your existing datasets, they will default to Best mode when you upgrade, which chooses the appropriate mode for your dataset based on its size, the type of media, and other criteria. You can toggle between All and Fast on-the-fly in the App, and it’s also possible to configure the default sidebar mode on a per-dataset basis via code. You can see all this in action from [23:51](https://www.youtube.com/watch?v=DukuxznuRuk&t=1431s)–26:39 in the webinar replay. Also visit the [sidebar mode](https://voxel51.com/docs/fiftyone/user_guide/app.html#sidebar-mode) section of the docs to learn more. ## 3 — Custom Sidebar Groups With the FiftyOne App, you’ve always been able to customize the groups in the sidebar, including the ability to create, rename, or delete sections and the fields they contain; drag and drop groups and fields just where you’d like them; and expand or collapse groups as you see fit. All of these options enable you to fully customize how you organize your App sidebar. But now with FiftyOne 0.18, you have the ability to also customize all of this on a per-dataset basis via code in the dataset’s app config! After doing so, when you work with a particular dataset in the App, it will automatically load with the sidebar groups you defined. This means you can visualize what you want immediately, while also being able to customize further in the App, as usual. In the demo, Brian shows how to add a new sidebar group called “new” and add a field to it, as well as set the metadata fields to non-expanded by default, all through code. See the demo in the playback video from [26:39](https://www.youtube.com/watch?v=DukuxznuRuk&t=1599s)–28:22. Also learn more in the [custom sidebar groups](https://voxel51.com/docs/fiftyone/user_guide/app.html#sidebar-groups) section of the docs. ## 4 — Custom Label Attributes With FiftyOne, it has always been the case that, in code, you could add as many label fields to your dataset as you’d like. You could also add your own custom attributes to the labels. As an example, in a scenario with a ground truth field containing object detections, you could add a custom attribute declaring whether that object is a crowd annotation or not. Now this capability extends beyond code, enabling you to see and filter on custom label attributes in the App! First, Brian shows code samples and describes a variety of methods to help you achieve your desired end-state from [13:18](https://www.youtube.com/watch?v=DukuxznuRuk&t=798s)–16:24 in the presentation. Then later in the demo, [starting from 18:46](https://www.youtube.com/watch?v=DukuxznuRuk&t=1126s)–23:51 in the video replay, Brian loads up an example dataset with images of animals. The dataset has some ground truth object detections in it, with a custom attribute called “mood”. With the new features in 0.18 it’s now possible to add these dynamic attributes to the dataset schema, and then view and filter by these dynamic custom attributes in the App. To show this with a popular open dataset, Brian then loads up COCO and shows that there are additional attributes already on the dataset, including an “iscrowd” attribute, and then shows how to use a variety of methods to find the dynamic attributes, add those to the dataset, and then filter and view samples by that attribute in the App. Read more about [custom label attributes](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#dynamic-attributes) in the docs. ## 5 — Storing Field Metadata As of FiftyOne 0.18, you can now store metadata, including a description and other info such as URLs or an arbitrary mapping of keys and values on the fields of your dataset. In addition, if you save those details, you can retrieve them later through code and see them in the App via a new tooltip that appears when you hover over a field or an attribute name in the sidebar. In the demo, Brian defines the description and info fields on a few fields in a dataset, then looks at the App to show the tooltip that exposes the metadata on hover. See how to do this from [28:22](https://www.youtube.com/watch?v=DukuxznuRuk&t=1702s)–29:18 in the webinar playback and here in the docs on [storing field metadata](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#storing-field-metadata). ## 6 — Light Mode With this release, there’s a new toggle in the upper right of the FiftyOne App to toggle between light mode and dark mode options. Try it out to determine your favorite mode, or anytime you’re feeling up for a different vibe while working in the App. ![](https://cdn.sanity.io/images/h6toihm1/production/967170f376258343478cc1df94aa946b8ab512fe-1600x1350.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## Q&A from the Webinar There was a lively Q&A all throughout the presentation and demos! Here’s a recap: **Are you planning to have more tutorials on using FiftyOne, especially video tutorials?** You can find some [tutorials](https://voxel51.com/docs/fiftyone/tutorials/index.html) in the docs, but yes, the plan is to generate more video content about how to work with image datasets, video datasets, and 3D datasets with FiftyOne. So stay tuned! We’ll be posting to [our YouTube channel](https://www.youtube.com/@voxel5148) when they’re ready. **Where is the data stored when you’re working with FiftyOne? (Specifically, can the data storage still be on our own servers?)** Re: the specific question — yes, you can store the data locally on your own servers. With FiftyOne Teams, you can also store (and access) data in private clouds. Re: where data is stored more broadly — open source FiftyOne is designed for a single user and local deployment. So when you import the library the first time, a mongoDB database is spun up in the background. And any metadata or samples you create, all that information is stored in the database. And then, you’re also storing pointers to the media — a path to the media, like a video, on disk. So that media lives separately from FiftyOne. You can work with FiftyOne in all kinds of different environments. [Here’s a section of the docs](https://voxel51.com/docs/fiftyone/environments/index.html) that includes comprehensive details for different local environments where you can use open source FiftyOne (as well as where the data is stored for each environment and how to launch the App in each type of environment): - Local machine: Data is stored on the same computer that will be used to launch the App - Remote machine: Data is stored on disk on a separate machine (typically a remote server) from the one that will be used to launch the App - Notebooks: You are working from a Jupyter Notebook or a Google Colab Notebook If you’re looking for native cloud storage support, check out [FiftyOne Teams](https://voxel51.com/fiftyone-teams/). It’s a fully backwards compatible version of FiftyOne that supports native cloud datasets, multiuser collaboration features, and much more! [Learn more](https://voxel51.com/fiftyone-teams/) or reach out in the [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) and we’d be happy to answer any questions you may have. **Are there video tensor loading capabilities in minibatches with FiftyOne?** Datasets in FiftyOne work with data loaders, such as the PyTorch DataLoader. You can definitely work in batches. When working with large datasets, for instance, sending tasks to annotation services, it is good practice to send in batches. **What would a video action recognition pipeline look like in FiftyOne?** So suppose you have a pipeline where you’re training and you want to apply inference on, whether it’s inference on images or inference on a fancier model, like a video action recognition model, FiftyOne does have a built-in capability for that. One way to do this in FiftyOne is to use a method called [apply\_model()](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html?highlight=apply_model#fiftyone.core.collections.SampleCollection.apply_model) which you can learn about in the docs. We’ll be adding documentation coming up soon about how you can wrap your own PyTorch model as a model instance in FiftyOne, and then just pass it to this method, along with some configuration around batch size, the number of workers for your data loader, and so on. Stay tuned for that. Another, more standard way to do that would just be to write your own loop. FiftyOne is a way to store the model predictions and it works in concert with your existing training loops. Looking at our tutorial about [Training with Detectron2](https://voxel51.com/docs/fiftyone/tutorials/detectron2.html) as an example, what you’ll see is that you run inference as normal and you train as normal. What you just need to do is add to your inference loop some code that extracts the model’s predictions and stores them in the appropriate FiftyOne types and adds them to your dataset. Inference happens on the machines where you are doing inference, and then basically use them to import FiftyOne and load the right dataset and add your data to that dataset. **Can you show how to use FiftyOne with ActivityNet?** Yes! FiftyOne provides a [Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/) that contains a collection of common datasets that you can download and load into FiftyOne via a few simple commands. You can find all of the [datasets available in the Zoo here](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html), including ActivityNet. In fact, the ActivityNet authors are recommending to download and work with their dataset through FiftyOne, as you can see on [their website](http://activity-net.org/download.html). In the FiftyOne documentation we have a comprehensive tutorial on [how to download, visualize, and evaluate on ActivityNet](https://voxel51.com/docs/fiftyone/integrations/activitynet.html), including all kinds of options for loading a specific split of the dataset, or only certain classes or only videos of a certain duration and more. You can get a brief tour of the tutorial in the webinar replay from [35:10](https://www.youtube.com/watch?v=DukuxznuRuk&t=2110s)–36:45 in this video. **Can FiftyOne provide any automations for frame by frame annotations?** Yes, in FiftyOne we have integrations with annotation tools, currently CVAT, Label Studio, and Labelbox. Video annotations are supported, as is frame by frame video annotation. You can send a video dataset up for annotation using the integration and pull back in frame by frame annotations. Often for video annotations, key frames are annotated vs every frame. FiftyOne will automatically interpolate frame annotations, while also indicating which ones were human-annotated key frames annotated. For more information, visit the [annotation integrations](https://voxel51.com/docs/fiftyone/integrations/) in the docs. **Are the changes in app config also exported when we export the dataset?** Yes! There are a lot of ways to [export datasets in FiftyOne](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html). But, if you are looking to export all of the dataset metadata you’ve customized in FiftyOne, like field metadata and app config settings, that’s all included when you export in [FiftyOne dataset format](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html#fiftyonedataset). So if you export in this format and then pull it back in, yes, it will be included. **Is video embedding clustering possible with FiftyOne? Maybe even object flow visualizations?** Yes! The part of the tool where you work with embeddings is called the [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html). In the docs, we provide examples of [visualizing embeddings](https://voxel51.com/docs/fiftyone/user_guide/brain.html#brain-embeddings-visualization). If you want to work with visualizing video embeddings, you will need to use the “compute your own embeddings and provide them in array form” option in FiftyOne. It’s possible to wrap your own model in FiftyOne’s interface. And there is an interface video embeddings model. But the choice of how you want to generate an embedding for a video is up to you because it depends on what you’re trying to optimize for. For example, if you want to embed every 10th frame and then average the embeddings or concatenate embeddings, you can do that. Or if you want to use an activity recognition model and pull some video embeddings from it, you can do that, too. Once you’ve computed some embeddings, you can pass them to the relevant methods downstream in FiftyOne and you can have your clustering workflows for video models. **Can a dataset be corrected on the go?** Yes! Everything about FiftyOne is mutable. For example, there’s support in the App for editing in the form of tagging. Once you’ve loaded in a dataset, if you find some samples that you don’t like, for example, you can tag them, and then find or exclude those samples through the App. Those changes made in the App are also persisted to the underlying dataset, so that you can edit your datasets programmatically too. You can also add or remove new fields at any time. There’s also a workflow for editing annotation data, simply check out one of FiftyOne’s [annotation integrations](https://voxel51.com/docs/fiftyone/integrations/index.html) to learn more about it. In addition, many parts of FiftyOne are pluggable, including the tabs here: ![](https://cdn.sanity.io/images/h6toihm1/production/ed591130d23a4bd8be355d66270bc4f59a774b98-1738x458.png?auto=format&dpr=2&fit=max&q=75&w=1600) There’s a plugin API whereby you could write your own plugin to render your own custom visualization of your dataset. Checkout the [docs on custom plugins here](https://github.com/voxel51/fiftyone/blob/develop/app/packages/plugins/README.md). Not only can you plug in this area of the code, you could also plug in your own visualizer. So it’s possible, and would be amazing, for someone to write a plugin that would embed the annotation tool into the App, because then you could directly edit in FiftyOne if that was of interest to you. **Is FiftyOne also an all-in-one CV tool with segmentation features, etc?** FiftyOne has [semantic segmentation capabilities](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#semantic-segmentation), but it is not meant to be an all-in-one CV tool. It is meant to be an essential part of the CV workflow to help data science teams improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. ## What’s Next - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone)! - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help! [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne 0.18](https://voxel51.com/blog/tag/fiftyone-0-18) [webinar recap](https://voxel51.com/blog/tag/webinar-recap) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/17422a76c76f14096dce21e43da51945f498a811-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne & FiftyOne Teams\\ \\ Event Recaps\\ \\ • \\ \\ Oct 8, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams) [![](https://cdn.sanity.io/images/h6toihm1/production/02ddf540125c56fc7e43080d8c64b02965a840d1-1200x681.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: Pandas-Style Queries for Computer Vision Data\\ \\ Event Recaps\\ \\ • \\ \\ Dec 20, 2022](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data) [![](https://cdn.sanity.io/images/h6toihm1/production/26fba75b02878a8803a43bf05ea3dd9153e4e491-1199x675.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne 0.19 for Computer Vision\\ \\ Event Recaps\\ \\ • \\ \\ Mar 2, 2023](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-191-lllmstxt|> ## Forsight's Dataset Management Solution [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/af1426b965796134b254c4c1ff02be13740eebc5-912x913.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=300&q=75&w=300) [Case Studies](https://voxel51.com/customers) Forsight Forsight Finds a Centralized Dataset Management Solution in FiftyOne Teams May 1, 2025 - Company: Forsight - Industry: AL Software for safety & security on jobsites - Challenge: Scaling dataset management as the R&D team grows - Solution: Dataset management through FiftyOne Teams Daily managed visual data 1.5 TB Engineering time saved 1/2 FTE Article content In this article [Success Story at a Glance](https://voxel51.com/customers/forsight#178a31b85f6c) [How Forsight Uses AI](https://voxel51.com/customers/forsight#592195690e4d) [Scaling Dataset Management as Business Grows](https://voxel51.com/customers/forsight#20e938ec7c62) [FiftyOne Teams Makes Dataset Management Easy](https://voxel51.com/customers/forsight#e38738c3fd7b) [Better data, better visibility, better models](https://voxel51.com/customers/forsight#3571ad79587a) [Cost effective dataset management](https://voxel51.com/customers/forsight#f648084326ef) [Enabling quick iteration](https://voxel51.com/customers/forsight#4be443e14e9d) [Conclusion](https://voxel51.com/customers/forsight#83960d7111ab) In this article [Success Story at a Glance](https://voxel51.com/customers/forsight#178a31b85f6c) [How Forsight Uses AI](https://voxel51.com/customers/forsight#592195690e4d) [Scaling Dataset Management as Business Grows](https://voxel51.com/customers/forsight#20e938ec7c62) [FiftyOne Teams Makes Dataset Management Easy](https://voxel51.com/customers/forsight#e38738c3fd7b) [Better data, better visibility, better models](https://voxel51.com/customers/forsight#3571ad79587a) [Cost effective dataset management](https://voxel51.com/customers/forsight#f648084326ef) [Enabling quick iteration](https://voxel51.com/customers/forsight#4be443e14e9d) [Conclusion](https://voxel51.com/customers/forsight#83960d7111ab) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) # Success Story at a Glance ### Challenge - As their R&D team grew, manual dataset management was no longer going to work - [Forsight](https://forsight.ai/) needed an easier, more transparent way to share image and video datasets among team members ### Solution - Forsight chose [FiftyOne](https://voxel51.com/) Teams as their central computer vision dataset management system where all the data is aggregated and consumed by all of their machine learning pipelines ### Results - 1.5TB of visual data is managed daily through FiftyOne Teams - Half an engineer’s time is saved - Datasets and model performance are improved - Costs are straightforward, not complex ## How Forsight Uses AI [Forsight](https://forsight.ai/) is on a mission to help keep workers and the jobsites where they work safe and secure. Forsight does this by using cutting edge AI vision technology to build solutions for dynamic environments that address challenges regarding safety, security, management and more. “Missing safety equipment or not adhering to safety regulations are common but preventable occurrences on jobsites in the US today. We founded Forsight around the idea that if we could save just one person’s life with AI and software, then our efforts would be well worth it,” said Ivan Ralašić, CTO and co-founder of Forsight. Although Forsight started with an initial focus on safety in the construction industry, the company now also works with industrial plant owners and operators, manufacturers, mining operations, and other customers to make their sites safer and day to day tasks easier through the use of real-time, vision-powered technology. Forsight develops AI software that uses CCTV cameras to detect and predict safety incidents, security threats, and management issues in real time. Their solution monitors jobsites for safety issues, like missing personal protective equipment, no-go zones for safety or security reasons, as well as vehicles, license plates, and more. It does this all in real time, which means processing camera feeds ten times a second, and many cameras in parallel. ## Scaling Dataset Management as Business Grows Forsight’s AI technology is now deployed at hundreds of customer sites and as Forsight’s business grew, so did their R&D team. Ivan noted, “initially, dataset management was a manual process, and the ingestion of data to the training instances was cumbersome with all the syncing. Luckily, it was only me training the models at the time, and I somehow managed to keep track of everything. As we added more team members, we needed a more transparent and easier way to share datasets among team members.” Ivan wanted to scale and streamline dataset management as the R&D team grew, as well as avoid the pitfalls of manual dataset versioning. Forsight’s primary requirements for a dataset management solution were that it had to fit into their existing cloud infrastructure and be cost effective. The Forsight engineering team tested a few different products and then found open source FiftyOne on their way to finding FiftyOne Teams. >> [FiftyOne](https://github.com/voxel51/fiftyone) is the open source tool for building high-quality datasets and computer vision models. >> [FiftyOne Teams](https://voxel51.com/) inherits all the goodness of open source FiftyOne and adds collaborative features built specifically for teams, including cloud-backed media, dataset permissions, versioning, sharing, and more. “The open source FiftyOne version stood out from other solutions because it had much more powerful features, and the FiftyOne Brain was a big bonus. FiftyOne worked really great, but as a team we needed additional features like central permission management and access control, and that’s where FiftyOne Teams comes into the picture,” explained Ivan. ## FiftyOne Teams Makes Dataset Management Easy Forsight is using [FiftyOne Teams](https://voxel51.com/) as their central dataset management system where all the data is aggregated and consumed by all of their machine learning pipelines. “It’s a great thing to have centralized dataset management in the form of FiftyOne Teams, which really makes our lives easier when curating datasets and training new models. All the team members have the same view on the datasets, which ensures that everyone understands the data used to train the models. This really saves a few hours of back and forth between team members,” said Ivan. Forsight’s R&D team trains different models to perform tasks such as detection of personal protective equipment, intrusion detection, construction vehicle detection, fire detection, semantic segmentation, and more. Their lightweight CV and ML algorithms run on-premises on edge devices that process the CCTV camera streams in real time. For typical deployments, the edge devices are powered by NVIDIA’s embedded Jetson technology, and for larger deployments, Forsight uses NVIDIA GPUs that can process 25+ camera streams in real time in parallel. Video feeds and inferred metadata are then sent into Forsight’s cloud provider, which is AWS. They’re using Amazon EC2 instances for training their ML models, and Amazon S3 buckets for media storage. Ivan added, “We wanted to have high throughput between our S3 buckets and the EC2 instances, which we are able to get from FiftyOne Teams.” In addition to integrating with AWS, Forsight also integrated FiftyOne Teams with CVAT for annotation, PyTorch for model training loops, and ClearML for experiment tracking. Overall, Forsight’s R&D team works with about 1.5 TB of data (mostly images, but some video datasets) on a daily basis, and all that data goes through FiftyOne Teams. ## Better data, better visibility, better models FiftyOne exists to give developers and scientists comprehensive visibility into their datasets, so they can c​urate better data and build better models. Because FiftyOne Teams is built on top of open source FiftyOne, it does all that and adds collaboration and access features that teams need. Forsight uses FiftyOne to automate dataset management, such as auto-tagging suspicious samples defined by rules set by mistakeness and embeddings. They also use FiftyOne to auto-balance their datasets and help reduce bias in their models. Adding the collaboration features of FiftyOne Teams extends dataset visibility team-wide. “The quality of our datasets is a key component of the performance of our models. FiftyOne helps us improve our datasets by identifying mislabeled data, adversarial samples, and unusually sized samples. FiftyOne Teams helps us track our datasets across multiple model runs and evaluate our models. We can see how our models are improving as we track model predictions on certain samples over multiple runs,” said Daniel Reiff, Machine Learning Engineer at Forsight. “We have seen performance improvements in our models directly due to using FiftyOne Teams for dataset management. FiftyOne Teams has greatly improved the visibility of our datasets across our entire R&D team, it has made it extremely easy for multiple team members to access and collaborate on datasets,“ explained Ivan. ## Cost effective dataset management The pricing model for other dataset management solutions can be complex, but with FiftyOne Teams, pricing is straightforward and includes unlimited data. “Some companies charge per gigabyte transferred, while others charge per image stored, per model trained, and per hour — it’s so complex. The costs are maybe manageable in the beginning to attract you to the product, but as you scale and have more and more data, then you run into problems. With FiftyOne Teams, pricing is straightforward, includes unlimited data, and that’s what we like,” explained Ivan. FiftyOne Teams also streamlines the process of working with data, freeing up valuable engineering time. Ivan estimates that the time saved with FiftyOne Teams is equivalent to half an engineer, which means the team can spend more time building new features and less time wrangling data. ## Enabling quick iteration The Forsight R&D team iterates quickly on their models, and uses FiftyOne Teams to help. “There are a ton of features in FiftyOne Teams that I use to iterate quickly on our models. For example, there are a lot of features of the [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html) that I find incredibly useful, including mistakenness and uniqueness that help me quickly identify suspicious samples to send to our labelers, get things turned around, and retrain the model.” Daniel gave a shout out to another one of his favorite features — [heatmaps](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#heatmaps). “My favorite FiftyOne Teams feature is heatmaps! At Forsight we love using Grad-CAM heatmaps to interpret what our models have learned. Intuitively, we use these heatmaps to build more adversarial samples into our datasets so we can rapidly improve our models. Being able to view these heatmaps in FiftyOne Teams is a great advantage because we can tag samples where the model has not learned the critical regions and build sets of adversarial samples,” said Daniel. Check out [this article](https://towardsdatascience.com/understand-your-algorithm-with-grad-cam-d3b62fce353) to learn more about how Forsight uses Grad-CAM heatmaps. ## Conclusion Forsight uses cutting edge AI vision technology to build solutions for jobsites that address challenges regarding safety, security, management, and more. The Forsight R&D team needed a dataset management solution and tested out a few products before finding open source FiftyOne, which ultimately led them to FiftyOne Teams. FiftyOne Teams meets all their requirements including integrating with their existing cloud infrastructure and being a cost effective solution. In addition, FiftyOne Teams has led to significant results in terms of improved speed, quality, and performance. See for yourself: - [Try FiftyOne](https://voxel51.com/docs/fiftyone/), it’s easy to get up and running in a few minutes - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help - Check out [FiftyOne Teams](https://voxel51.com/) to enable multiple users to securely collaborate on the same datasets and models ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) [Submit a case study](https://share.hsforms.com/1o6aebyQbShWveAHzweg-Ng2ykyk) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-192-lllmstxt|> ## FiftyOne 0.18 Release [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Announcing FiftyOne 0.18 with App Performance Improvements, Sidebar Modes, and Custom Attributes Nov 15, 2022 • 6 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#581b2d3f174b) [tl;dr: What’s new in this release?](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#d8a61afb939d) [Join me for a live demo and AMA on December 1](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#8a412d454c4f) [App performance improvements](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#9fe0f6ec68cb) [App sidebar modes](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#f7b42037f734) [Custom sidebar groups](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#86afb0f48c43) [Custom label attributes](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#44143768d50b) [Storing and visualizing field metadata](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#a82e130c4993) [Now in light mode](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#dfe78c6ecd11) [Community contributions](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#acd83fda43f8) [FiftyOne community updates](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#772ea068e039) [What’s next?](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#cca24d578739) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#581b2d3f174b) [tl;dr: What’s new in this release?](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#d8a61afb939d) [Join me for a live demo and AMA on December 1](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#8a412d454c4f) [App performance improvements](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#9fe0f6ec68cb) [App sidebar modes](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#f7b42037f734) [Custom sidebar groups](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#86afb0f48c43) [Custom label attributes](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#44143768d50b) [Storing and visualizing field metadata](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#a82e130c4993) [Now in light mode](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#dfe78c6ecd11) [Community contributions](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#acd83fda43f8) [FiftyOne community updates](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#772ea068e039) [What’s next?](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes#cca24d578739) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Voxel51 in conjunction with the FiftyOne community is excited to announce the general availability of [FiftyOne 0.18](https://voxel51.com/docs/fiftyone/release-notes.html#fiftyone-0-18-0)! ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into the release! ## tl;dr: What’s new in this release? This release includes: - Significant performance improvements to the FiftyOne App - New sidebar modes to further optimize user experience for large datasets - The ability to declare and filter by custom label attributes - Support for storing and viewing field metadata in the App - A new light mode option Check out the [release notes](https://voxel51.com/docs/fiftyone/release-notes.html#fiftyone-0-18-0) for a full rundown of additional enhancements and bugfixes. ## Join me for a live demo and AMA on December 1 You can see all of the new features in action in a live webinar and AMA on December 1, 2022 at 10 AM Pacific Time. I’ll be demoing all of the new features, followed by an open Q&A where you can get answers to any questions you might have. [Register here.](https://us02web.zoom.us/webinar/register/7816682079588/WN_sY5_tv1WQUSxBdEtmv8_LA) Now, here’s a quick overview of some of the new features we packed into this release. ## App performance improvements FiftyOne 0.18 features major upgrades “under the hood” that have yielded significant performance improvements when working with large datasets with many samples and/or fields in the App. For example, here’s a search across 185,000 objects in FiftyOne 0.17: \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop The same search in FiftyOne 0.18 is now **10x faster**! \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop What key changes are driving these improvements? - Sidebar statistics are now lazily computed only for the currently visible fields, so only the data that you care about is loaded - Similarly-shaped data are optimized into faceted database aggregations - Improved request architecture ensuring that components load asynchronously without blocking ## App sidebar modes Another performance-oriented feature is the new configurable **sidebar mode**, including an updated default behavior that further optimizes the performance of the App for large datasets. For context, each time you load a new dataset or view in the FiftyOne App, the sidebar updates to show statistics for the current collection. For large datasets with many samples or fields, this can involve substantial computation. The App now supports three sidebar modes that you can choose between: - **“all”**: always compute counts for all visible fields and stats for all fields whose filter tray is expanded - **“fast”**: only compute counts and stats for fields whose filter tray is expanded - **“best” (default)**: automatically choose between “all” and “fast” mode based on the size of the dataset When the sidebar mode is “best”, the App will choose “fast” mode if any of the following conditions are met: - Any dataset with 10,000+ samples - Any dataset with 1,000+ samples and 15+ top-level fields in the sidebar - Any video dataset with frame-level label fields You can toggle the sidebar mode dynamically for your current session via the App’s settings menu: ![](https://cdn.sanity.io/images/h6toihm1/production/df40d94b6f385c2e5e90e89137b5c9ce9b255b52-914x675.gif?auto=format&dpr=2&fit=max&q=75&w=914) You can also permanently configure the default sidebar mode of a dataset by modifying the `sidebar_mode` property of the dataset’s App config: \# Set the default sidebar mode to "fast" dataset.app\_config.sidebar\_mode = "fast" dataset.save() # must save after edits session.refresh() Check out [the docs](https://voxel51.com/docs/fiftyone/user_guide/app.html#app-sidebar-mode) for more information about sidebar modes and how to configure them on a per-dataset basis. ## Custom sidebar groups If you’re an existing FiftyOne user, you may have discovered that you can customize the layout of the App’s sidebar by creating/renaming/deleting groups and dragging fields between groups directly in the App, with all changes you make automatically persisting between sessions. ![](https://cdn.sanity.io/images/h6toihm1/production/df40d94b6f385c2e5e90e89137b5c9ce9b255b52-914x675.gif?auto=format&dpr=2&fit=max&q=75&w=914) Well, now you can programmatically modify your dataset’s sidebar groups by editing the `sidebar_groups` property of the dataset’s App config, including the ability to configure whether or not the group is expanded by default: import fiftyone as fo import fiftyone.zoo as foz dataset = foz.load\_zoo\_dataset("quickstart") \# Get the default sidebar groups for the dataset sidebar\_groups = fo.DatasetAppConfig.default\_sidebar\_groups(dataset) \# Collapse the \`metadata\` section by default print(sidebar\_groups\[2\].name) # metadata sidebar\_groups\[2\].expanded = False \# Modify the dataset's App config dataset.app\_config.sidebar\_groups = sidebar\_groups dataset.save() session = fo.launch\_app(dataset) Check out [the docs](https://voxel51.com/docs/fiftyone/user_guide/app.html#sidebar-groups) to learn more about configuring your dataset’s sidebar groups via the App and the SDK. ## Custom label attributes Exciting news! A highly requested feature is finally available: you can now [declare custom attributes](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#dynamic-attributes) on your label fields (or, in general, any embedded field) and filter by them in the FiftyOne App. This can be achieved in a variety of ways, including: 1\. Using `add_sample_field()` to directly declare a new label attribute: \# Declare a new \`iscrowd\` attribute on the \`ground\_truth\` objects dataset.add\_sample\_field("ground\_truth.detections.iscrowd", fo.FloatField) 2\. Providing the new `dynamic=True` option when adding samples to datasets using methods like `add_samples()` and `from_dir()` to automatically declare any dynamic attributes encountered while importing data: \# Automatically declare all new dynamic attributes dataset.add\_samples(samples, dynamic=True) 3\. Using `get_dynamic_field_schema()` to detect the names and type(s) of any undeclared dynamic attributes, and then using `add_dynamic_sample_fields()` to automatically declare them: \# View any undeclared dynamic attributes print(dataset.get\_dynamic\_field\_schema()) \# Declare them dataset.add\_dynamic\_sample\_fields() Any dynamic attributes that you declare on your dataset’s schema are now exposed as filterable in the FiftyOne App! ![](https://cdn.sanity.io/images/h6toihm1/production/a4c2bee9ed053c5be2a1c161e5abf758c9a12ff8-1400x923.png?auto=format&dpr=2&fit=max&q=75&w=1400) Refer to [the docs](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#dynamic-attributes) for plenty of additional information on custom attributes, including viewing, selecting, and excluding them from your dataset’s schema. ## Storing and visualizing field metadata You can now [store metadata](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#storing-field-metadata) such as descriptions and other information like URLs on the fields of your dataset! import fiftyone as fo import fiftyone.zoo as foz dataset = foz.load\_zoo\_dataset("quickstart") dataset.add\_dynamic\_sample\_fields() field = dataset.get\_field("ground\_truth") field.description = "Ground truth annotations" field.info = {"url": "https://fiftyone.ai"} field.save() field = dataset.get\_field("ground\_truth.detections.area") field.description = "Area of the box, in pixels^2" field.info = {"url": "https://fiftyone.ai"} field.save() session = fo.launch\_app(dataset) Moreover, you can view this information in the FiftyOne App by hovering over field or attribute names in the App’s sidebar! ![](https://cdn.sanity.io/images/h6toihm1/production/41f8cda1e1d4a556b66ad49ad1db555250012cc8-915x641.gif?auto=format&dpr=2&fit=max&q=75&w=915) Check out the docs for more information about [storing](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#storing-field-metadata) and [viewing](https://voxel51.com/docs/fiftyone/user_guide/app.html#using-the-sidebar) field metadata. ## Now in light mode Last but not least, FiftyOne now supports light mode! You can switch to light mode by clicking the sun/moon icon to the right of the [Have a Team?](https://voxel51.com/teams) button in the FiftyOne App: ![](https://cdn.sanity.io/images/h6toihm1/production/5f42b8f454aa80e77a96aff6038086f140789821-1400x1113.png?auto=format&dpr=2&fit=max&q=75&w=1400) By default, your browser will remember your last theme, but you can also store a default theme [in your App config](https://voxel51.com/docs/fiftyone/user_guide/config.html#configuring-the-app). ## Community contributions Shoutout to the following community members who contributed to this release! - [Jan Steeg](https://github.com/karsil) contributed [#2114 — new YOLOv5 export option](https://github.com/voxel51/fiftyone/pull/2114) - [Daniel Langenkämper](https://github.com/dlangenk) contributed [#2122 — new COCO import option](https://github.com/voxel51/fiftyone/pull/2122) - [Conor Doyle](https://github.com/HoopsMcann) contributed [#2128 — interpolation option for image resizing](https://github.com/voxel51/fiftyone/pull/2128) - [Laura Lin](https://github.com/lauralindy) contributed [#2198 — exposing dataset info in App](https://github.com/voxel51/fiftyone/pull/2198) - [Odd Eirik Igland](https://github.com/oddeirikigland) contributed [#2145 — handling missing data with Label Studio](https://github.com/voxel51/fiftyone/pull/2145) - [Kishan Savant](https://github.com/NeoKish) contributed [#2107 — docs fixes](https://github.com/voxel51/fiftyone/pull/2107) and [#2206 — README updates](https://github.com/voxel51/fiftyone/pull/2206) - [andife](https://github.com/andife) contributed [#2177 — CVAT documentation fixes](https://github.com/voxel51/fiftyone/pull/2177) - [oguz-hanoglu](https://github.com/oguz-hanoglu) contributed [#2132 — typo fixes](https://github.com/voxel51/fiftyone/pull/2132) ## FiftyOne community updates The FiftyOne community continues to grow! - 1,100+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,100+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 1,500+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 190+ repositories - 46+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## What’s next? - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne 0.18](https://voxel51.com/blog/tag/fiftyone-0-18) [product release](https://voxel51.com/blog/tag/product-release) ![](https://cdn.sanity.io/images/h6toihm1/production/8d61ff90b31d151405f9e21a33c2802509f34651-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Brian Moore Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Loading related posts... [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-193-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Dec 02, 2022 Dec 3, 2022 • 7 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#08e841f3aeaa) [Evaluating performance by label](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#c1ca45a0d4b4) [Exporting a dataset with splits](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#dd21965816ca) [Getting and setting values in large datasets](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#5cc803c2b360) [Storing group level metadata](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#1ea1d5e68fb3) [Display bounding boxes without masks](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#d954918cf1bb) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#bde2755bb70a) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#c424eb3418ff) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#08e841f3aeaa) [Evaluating performance by label](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#c1ca45a0d4b4) [Exporting a dataset with splits](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#dd21965816ca) [Getting and setting values in large datasets](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#5cc803c2b360) [Storing group level metadata](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#1ea1d5e68fb3) [Display bounding boxes without masks](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#d954918cf1bb) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#bde2755bb70a) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022#c424eb3418ff) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## Evaluating performance by label Community Slack member Mike Kraus asked, _“I have run a number of evaluations on the same dataset, and I want to compute evaluation metrics for each label separately. What is the best way to do this?_ If you’re interested in composite metrics like the precision, recall, or f1-score, then look no further than FiftyOne’s evaluation API. The `evaluate_detections()` and `evaluate_classifications()` methods compute separate values for these metrics for each class, which can be viewed by calling `print_report()`. import fiftyone as fo import fiftyone.zoo as foz from fiftyone import ViewField as F dataset = foz.load\_zoo\_dataset("quickstart") result = dataset.evaluate\_detections( "predictions", gt\_field="ground\_truth", eval\_key="eval" ) result.print\_report() For more granular access to the detection counts, including false positives “fp” and false negatives “fn”, the evaluation API can be combined with FiftyOne’s `filter()` method to compute results by label, tag, or more flexible conditions. To count the false positives and false negatives for each class, we can pick up where we left off above: import fiftyone as fo import fiftyone.zoo as foz from fiftyone import ViewField as F dataset = foz.load\_zoo\_dataset("quickstart") result = dataset.evaluate\_detections( "predictions", gt\_field="ground\_truth", eval\_key="eval" ) counting\_results = {} for detection\_result in \["tp", "fp", "fn"\]: view = dataset.filter\_labels("predictions", F("eval") == type) counting\_results\[type\] = view.count\_values("predictions.detections.label") print(counting\_results) Learn more about FiftyOne’s [Evaluation API](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html) in the FiftyOne Docs. ## Exporting a dataset with splits Community Slack member Joy Timmermans asked, _“Is it possible to export a dataset with splits in it?”_ There are multiple ways to do this in FiftyOne. The first approach involves creating a separate view for each split you are interested in, and exporting each one of them. For instance, the following will export only the train and validation splits for a dataset: export\_dir\_train = "/path/for/export/train" export\_dir\_val = "/path/for/export/val" \# The name of the sample field containing the label that you wish to export \# Used when exporting labeled datasets (e.g., classification or detection) label\_field = "ground\_truth" # for example \# The type of dataset to export \# Any subclass of \`fiftyone.types.Dataset\` is supported dataset\_type = fo.types.COCODetectionDataset train\_view = dataset.match\_tags("train") train\_view.export( export\_dir=export\_dir\_train, dataset\_type=dataset\_type, label\_field=label\_field, ) test\_view = dataset.match\_tags("val") test\_view.export( export\_dir=export\_dir\_val, dataset\_type=dataset\_type, label\_field=label\_field, ) Alternatively, you can leverage the “split” keyword argument as input to `export()`, which will take care of the splitting for you. import fiftyone as fo import fiftyone.zoo as foz export\_dir = "/path/for/export" label\_field = "ground\_truth" # for example \# The splits to export splits = \["train", "val"\] \# All splits must use the same classes list classes = \["list", "of", "classes"\] \# The dataset or view to export \# We assume the dataset uses sample tags to encode the splits to export dataset\_or\_view = foz.load\_zoo\_model("quickstart") # portion of COCO dataset dataset\_type = fo.types.COCODetectionDataset \# Export the splits for split in splits: split\_view = dataset\_or\_view.match\_tags(split) split\_view.export( export\_dir=export\_dir, dataset\_type=fo.types.YOLOv5Dataset, label\_field=label\_field, split=split, classes=classes, ) Learn more about [exporting datasets](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html) in FiftyOne in the FiftyOne Docs. ## Getting and setting values in large datasets Community Slack member Geoffrey Keating asked, _“Is there a point at which calling the `values()` method on a very large dataset or field will be impractical?”_ This question cuts to the core of what makes FiftyOne great for dealing with massive computer vision datasets. Whereas FiftyOne does provide a `values()` method for getting the values in a field, and the related `set_values()` method for setting a field’s values, these methods are typically only practical for small to medium sized datasets. This is because these two methods loads the entire dataset from memory at once, and the size of the dataset can exceed the amount of available RAM. An alternative to use the `select_values()` and `set_field()` methods, which load in only constant memory. Of course, these methods do sacrifice speed in the name of conserving memory. To efficiently do more general operations on a field of a large dataset, `select_fields()` and `set_field()` can be combined with `iter_samples()` to iterate through all samples. For instance, if we wanted to keep a running tally of all detections across all samples, and add this counter value in a field “detection\_number” on each detection, we could do so with: counter = 0 dataset.add\_sample\_field( "ground\_truth.detections.detection\_number", fo.IntField ) for sample in dataset.select\_fields("ground\_truth.detections").iter\_samples( autosave=True, progress=True ): for detection in sample.ground\_truth.detections: detection.set\_field("detection\_number", counter) counter += 1 Learn more about `set_field()` and [Fields](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#fields) in the FiftyOne Docs. ## Storing group level metadata Community Slack member Andreas Fehlner asked, _“I really like the new group feature. At the moment there is metadata common to all of the media in a group. Is it possible to store that metadata for all of the files in one group?”_ One way to do this is to select one group slice and iterate through all samples in that slice. For each sample, you can then get all of the other samples in the same group using the `get_group()` method, and then add the metadata to all samples in the group. left\_samples = dataset.select\_group\_slices("left") slices = dataset.group\_slices for ls in left\_samples.iter\_samples(autosave=True, progress=True): group\_metadata = ls.filepath ### Will be same for entire group group\_id = ls.group.id group = dataset.get\_group(group\_id) for s in slices: dataset.group\_slice = s sample = dataset\[group\[s\].id\] ### Select from dataset by id sample\["group\_metadata\_field"\] = group\_metadata Alternatively, if you have multiple related media files that have the same aspect ratio, you can use multiple media fields on a single sample. For instance, if we want to generate low resolution thumbnails for each image in our dataset, we can tie these thumbnails to the original images in a single sample object. import fiftyone as fo import fiftyone.utils.image as foui import fiftyone.zoo as foz dataset = foz.load\_zoo\_dataset("quickstart") \# Generate some thumbnail images foui.transform\_images( dataset, size=(-1, 32), output\_field="thumbnail\_path", output\_dir="/tmp/thumbnails", ) \# Modify the dataset's App config dataset.app\_config.media\_fields = \["filepath", "thumbnail\_path"\] dataset.app\_config.grid\_media\_field = "thumbnail\_path" dataset.save() # must save after edits Then we can add the metadata for each “group” of images or media files in just one place. Learn more about [Grouped datasets](https://voxel51.com/docs/fiftyone/user_guide/groups.html) and [Multiple media fields](https://voxel51.com/docs/fiftyone/user_guide/app.html#multiple-media-fields) in the FiftyOne Docs. ## Display bounding boxes without masks Community Slack member Lukas asked, _“Is there an easy way to display only the bounding boxes and not the masks for detections that have both?”_ Yes! This is possible in the Python API, where you can create a view that explicitly omits the masks while retaining the masks on the dataset: import fiftyone as fo import fiftyone.zoo as foz dataset = foz.load\_zoo\_dataset("quickstart") session = fo.Session() session.view = dataset.set\_field("ground\_truth.detections.mask", None) Learn more about views and [DatasetView](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) in the FiftyOne Docs. ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,100+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,100+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 1,900+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 205+ repositories - 49+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## What’s next? - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [bounding boxes](https://voxel51.com/blog/tag/bounding-boxes) [exporting with splits](https://voxel51.com/blog/tag/exporting-with-splits) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [large datasets](https://voxel51.com/blog/tag/large-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/d63c2ef9fed7cf00ea9af1c6dd3f2ee4d657c598-1200x685.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Nov 4, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 5, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/c28199522446929ca5288c0559047354eb93f01a-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Oct 28, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Oct 29, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/a17b9ee7620741f8c3225d0256d174c857a575d5-1200x672.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Oct 14, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Oct 15, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-194-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Dec 16, 2022 Dec 17, 2022 • 5 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#1c9437028a5c) [Sending grouped tasks to CVAT](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#a8efeb138c08) [Linking detections on multiple samples](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#d20c51a2ce97) [Merging large datasets into grouped dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#56c6834a4119) [Annotating tagged labels from multiple fields](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#2e36c863865f) [Visualizing point clouds in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#d5d398e86333) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#950bc503d7a1) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#081e763151b2) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#1c9437028a5c) [Sending grouped tasks to CVAT](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#a8efeb138c08) [Linking detections on multiple samples](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#d20c51a2ce97) [Merging large datasets into grouped dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#56c6834a4119) [Annotating tagged labels from multiple fields](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#2e36c863865f) [Visualizing point clouds in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#d5d398e86333) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#950bc503d7a1) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022#081e763151b2) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ![](https://cdn.sanity.io/images/h6toihm1/production/f8c59b7ff0a53a527b7002a899ccee84e95dee0a-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=1200) ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **Sending grouped tasks to CVAT** Community Slack member Daniel Fortunato asked, _“I am using FiftyOne to manage a grouped dataset with three images per group. I want to send these images to CVAT for annotation. How can I import all the images in each group as a single task with multiple jobs, rather than multiple tasks with one job each?_ If each group slice has the same media type, e.g. images, then you can do this by selecting the desired group slices with `select_group_slices()`, and grouping by the group id with `group_by()`. Calling `select_group_slices()` will return a collection of images (or more generally, a view whose `media_type` attribute is the same as the group slices, rather than `media_type = “group”`. The following `group_by()` operation will rearrange the resulting DatasetView such that samples from the same group are located at neighboring indices. Once this view has been created, it can be sent to CVAT using the `annotate()` method, with the `segment_size` argument set equal to the number of images per task — here the number of selected group slices. Putting it all together, we can see what this would look like for the [Quickstart Groups](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#dataset-zoo-quickstart-groups) dataset: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("quickstart-groups") 5slices = ["left", "right"] 6segment_size = len(slices) 7 8view = dataset.select_group_slices(slices).group_by("group._id") 9view.annotate(..., segment_size=segment_size) ``` Learn more about FiftyOne’s [annotation API](https://voxel51.com/docs/fiftyone/user_guide/annotation.html) and FiftyOne’s [integration with CVAT](https://voxel51.com/docs/fiftyone/integrations/cvat.html#cvat-integration) in the FiftyOne Docs. ## **Linking detections on multiple samples** Community Slack member Marijn Lems asked, _“Is it possible to relate multiple detections on different samples in FiftyOne?”_ In computer vision it is often the case that multiple images or media files have objects that are related to each other. For instance, in a dataset containing faces, such as [Families in the Wild](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#dataset-zoo-fiw), there could be multiple images that feature the same person’s face. Alternatively, in a grouped dataset like the [KITTI Multiview dataset](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#dataset-zoo-kitti-multiview), an image and a point-cloud could feature different representations of the same object. One way to associate these detections is by creating a Detection-level metadata field that you populate with a unique identifier for each associated group of detections. Working with the [Quickstart Groups](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#dataset-zoo-quickstart-groups) dataset, for example, linking a ground truth detection in the “left” and “right” images in the first group can be done as follows: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("quickstart-groups") 5 6left_sample = dataset.first() 7left_detection = left_sample.ground_truth.detections[0] 8right_sample_id = dataset.get_group(left_sample.group.id)["right"].id 9dataset.group_slice = "right" 10 11right_sample = dataset[right_sample_id] 12right_detection = right_sample.ground_truth.detections[0] 13 14"Tie the detections together" 15uuid = 0 16left_detection["detection_uuid"] = uuid 17left_sample.save() 18 19right_detection["detection_uuid"] = uuid 20right_sample.save() ``` Learn more about [object detections](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#object-detection) in FiftyOne in the FiftyOne Docs. ## **Merging large datasets into grouped dataset** Community Slack member Oğuz Hanoğlu asked, _“I have two large image datasets which I want to merge into a grouped dataset according to their values in a certain field. What is the most efficient way to do this?”_ When dealing with large datasets, you want to perform as many operations in-database as possible to avoid the long times typical of loading large datasets into memory. One way to do this is to create a temporary clone of each source dataset, then populate a `Group` field on each sample in these “single slice” datasets, and finally merge these datasets into a single grouped dataset. To merge datasets `ds_left` and `ds_right` based on their values in field `my_field`, this would look like: ```python 1import fiftyone as fo 2 3def group_collections(d, group_field, group_key): 4 dataset = fo.Dataset() 5 dataset.add_group_field(group_field) 6 7 group_keys = set().union(*[set(c.exists(group_key).distinct(group_key)) for c in d.values()]) 8 groups = {_id: fo.Group() for _id in group_keys} 9 10 for group_slice, sample_collection in d.items(): 11 _add_slice(dataset, groups, sample_collection, group_field, group_key, group_slice) 12 13 return dataset 14 15def _add_slice(dataset, groups, sample_collection, group_field, group_key, group_slice): 16 tmp = sample_collection.exists(group_key).clone() 17 group_values = [groups[k].element(group_slice) for k in tmp.values(group_key)] 18 19 tmp.add_group_field(group_field, default=group_slice) 20 tmp._doc.media_type = sample_collection.media_type 21 tmp.set_values(group_field, group_values) 22 tmp._doc.group_media_types = {group_slice: sample_collection.media_type} 23 tmp._doc.media_type = "group" 24 25 dataset.add_collection(tmp) 26 tmp.delete() 27 28dataset = group_collections({"left": ds_left, "right": ds_right}, "group", "my_field") ``` This approach does require the `set_values()` method rather than the `set_field()` method, as the `Group` instances need to be instantiated in memory in order to guarantee that they have unique `id`’s . However, for large datasets this approach is far more efficient than an approach which iterates through samples in the datasets to be merged. Learn more about [set\_field](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.set_field) and [set\_values](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.set_values) in the FiftyOne Docs. ## **Annotating tagged labels from multiple fields** Community Slack member Daniel Bourke asked, _“I want to send a collection of the worst predictions to Label Studio for annotation, including both ground truth and prediction label fields. How do I do this, when the label schema for Label Studio only accepts one field?”_ One way that you could do this now is to combine your incorrect ground truth and predictions into a single label field in FiftyOne first, then annotate that field in [Label Studio](https://labelstud.io/) (since it’s now just one field to re-annotate). First, you can clone the “ground\_truth” and “predictions” fields into new fields. Then, you can filter on these fields to create a view containing only the samples of interest. And finally, you can use the `merge_labels()` method to merge one, temporary, cloned field into the other cloned field: ```python 1dataset.clone_sample_field("ground_truth", "combined_field") 2dataset.clone_sample_field("predictions", "preds_to_merge") 3 4# Perform your filtering of the GT and Preds to only the ones you want to annotate 5dataset.filter_labels("combined_field", ...).save(fields="combined_field") 6dataset.filter_labels("preds_to_merge", ...).save(fields="preds_to_merge") 7 8# Merge them into one field, "preds_to_merge" will be deleted after this 9dataset.merge_labels("preds_to_merge", "combined_field") ``` This consolidates your desired information into a single field so that it is suitable for sending to Label Studio. Additionally, it preserves the original fields. Learn more about [merge\_labels](https://voxel51.com/docs/fiftyone/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.merge_labels) and [FiftyOne’s Label Studio integration](https://voxel51.com/docs/fiftyone/integrations/labelstudio.html) in the FiftyOne Docs. ## **Visualizing point clouds in the FiftyOne App** Community Slack member Marijn Lems asked, _“I have a dataset consisting of point-cloud data and I’d love to visualize this data in the FiftyOne App. How can I do this?”_ In FiftyOne, you can define, populate, and perform operations on a point-cloud only dataset, but you will not be able to visualize this out-of-the-box without having an image or other media file associated with each point cloud. This is because point-clouds trigger FiftyOne’s [3d visualizer](https://voxel51.com/docs/fiftyone/user_guide/groups.html#configuring-the-3d-visualizer) plugin, and it would be quite compute-intensive to have the sample-grid running a large number of these 3d visualizers at once. One solution to this is to create a [Grouped Dataset](https://voxel51.com/docs/fiftyone/user_guide/groups.html#) with the point-cloud media as one slice, and to generate an image for each point-cloud, saving these images into another group slice. As an example, you could generate a bird’s eye view image for the point clouds in the collection `pcd_samples`: ```python 1import numpy as np 2import cv2 3def my_birds_eye_view_function(pcd_sample): 4 return np.zeros((100, 100, 3)) 5 6 7import fiftyone as fo 8ds = foz.load_zoo_dataset("quickstart-groups") 9pcd_samples = ds.select_group_slices("pcd") 10print(pcd_samples) 11 12dataset = fo.Dataset("my-group-dataset") 13dataset.add_group_field("group", default="pcd") 14 15group_pcd_samples = [] 16group_bev_samples = [] 17 18for pcd_sample in pcd_samples: 19 group = fo.Group() 20 sample = fo.Sample(filepath=pcd_sample.filepath, group=group.element("pcd")) 21 group_pcd_samples.append(sample) 22 23 uuid = sample["filepath"] 24 bev_map = my_birds_eye_view_function(pcd_sample) 25 bev_filepath = "path/to/bev/img.png" 26 cv2.imwrite(bev_filepath, bev_map) 27 bev_sample = fo.Sample(filepath=bev_filepath) 28 bev_sample["group"] = group.element("bev") 29 group_bev_samples.append(bev_sample) 30 31dataset.add_samples(group_pcd_samples) 32dataset.add_samples(group_bev_samples) 33 34### set group slice to images so can view 35dataset.group_slice = "bev" 36session = fo.launch_app(dataset) ``` Learn more about [adding samples to a dataset](https://voxel51.com/docs/fiftyone/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.add_sample) in the FiftyOne Docs. Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,200+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,300+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 208+ repositories - 51+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,200+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,300+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 208+ repositories - 51+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## What’s next? - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [CVAT](https://voxel51.com/blog/tag/cvat) [detections](https://voxel51.com/blog/tag/detections) [grouped datasets](https://voxel51.com/blog/tag/grouped-datasets) [Label Studio](https://voxel51.com/blog/tag/label-studio) [point clouds](https://voxel51.com/blog/tag/point-clouds) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/47c4462ab8c14fe4728ce5901402986125d3f99d-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Nov 18, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 19, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/d63c2ef9fed7cf00ea9af1c6dd3f2ee4d657c598-1200x685.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Nov 4, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 5, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/03107d477b7db4be03031293fa4fe15aaea806f0-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Sept 16, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 17, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-195-lllmstxt|> ## ChatGPT and Computer Vision [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) Tunnel vision in computer vision: can ChatGPT see? Dec 16, 2022 • 17 min read Article content In this article [What is ChatGPT?](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#9af4784baa0b) [Where ChatGPT excels](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#3ce4287ef334) [Where it falters](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#238aea633702) [Where to exercise extreme caution](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#e5663f56572a) [Why ChatGPT empowers CV engineers](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#9f7c53af285d) [Just for fun](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#52b297b5fc8e) [FiftyOne Computer Vision toolset](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#7a325815be91) In this article [What is ChatGPT?](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#9af4784baa0b) [Where ChatGPT excels](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#3ce4287ef334) [Where it falters](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#238aea633702) [Where to exercise extreme caution](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#e5663f56572a) [Why ChatGPT empowers CV engineers](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#9f7c53af285d) [Just for fun](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#52b297b5fc8e) [FiftyOne Computer Vision toolset](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#7a325815be91) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/99dc871d875865fc200f14931a3cfb3118ded624-1020x1007.png?auto=format&dpr=2&fit=max&q=75&w=1020) In just two weeks, ChatGPT has taken a commanding hold of the public consciousness. More than a million people have “conversed” with OpenAI’s new chatbot, asking it to write [poems](https://www.npr.org/2022/12/10/1142045405/opinion-machine-made-poetry-is-here) and [college essays](https://www.theatlantic.com/technology/archive/2022/12/chatgpt-ai-writing-college-student-essays/672371/), generate recipe ideas, [build virtual machines](https://www.engraved.blog/building-a-virtual-machine-inside/), and oh so much more. It’s been used to [write the intro](https://www.youtube.com/watch?v=0gNauGdOkro) for news articles and YouTube videos — the fact that I even need to say that ChatGPT did _not_ write this introduction, by itself, speaks volumes. Who would have thought that a chatbot would be simultaneously [hailed as the next great disruptive technology](https://hbr.org/2022/12/chatgpt-and-how-ai-disrupts-industries), and [decried as a virus](https://techcrunch.com/2022/12/09/is-chatgpt-a-virus-that-has-been-released-into-the-wild/) that “has been released into the wild with no concern for the consequences”? And yet, underneath its seemingly intelligent facade, in some ways ChatGPT is frustratingly foolish. As a machine learning engineer at a computer vision (CV) company ( [Voxel51](https://voxel51.com/)), I’ve spent the last few days pushing ChatGPT to its limits to see what it “knows” about CV. I wanted to know what this language model means for the future (and present) of the field. Read on if you too are curious! Before diving in, a few quick words of caution: first, there are an infinite number of questions one could ask ChatGPT. While I attempted a relatively thorough investigation, you may have asked an entirely orthogonal set of questions. For instance, I only asked for Python code. If you are unsatisfied with my coverage, I encourage you to go to [https://chat.openai.com/chat](https://chat.openai.com/chat) and see for yourself. Second, ChatGPT is a generative model, so there is a degree of randomness inherent in its responses. Often when I would ask the same question multiple times, the results would look slightly different. If you ask these questions, there’s a chance that the answers you receive will also look different! Finally, and perhaps most importantly, this is only one person’s opinion. This post is broken into six parts: - [What is ChatGPT?](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#chatgpt) - [Where ChatGPT excels](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#excels) - [Where it falters](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#falters) - [Where to exercise extreme caution](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#caution) - [Why ChatGPT empowers CV engineers](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#empowers) - [Just for fun](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see#fun) ## What is ChatGPT? Released on November 30th, 2022, ChatGPT is OpenAI’s latest product to set the tech world on fire. Like GPT1, GPT2, GPT3, and InstructGPT before it, ChatGPT is a generative pretrained transformer (GPT) model, a type of large language model built with the notion of “self-attention”, which allows the model to flexibly identify which portions of the input — in this case the text from a conversation — are relevant to which other portions. [Large language models](https://hai.stanford.edu/news/how-large-language-models-will-transform-science-society-and-ai) (LLMs) are trained on [vast amounts of text data](https://commoncrawl.org/), such as books and articles, in order to learn the patterns and structures of human language. This allows them to generate text that sounds natural and human-like, making them useful for tasks like language translation and generating responses to questions. For the past few years, LLMs have been rapidly growing in popularity. These models have exponentially increased in size: whereas [the first transformer model](https://arxiv.org/pdf/1706.03762.pdf), introduced in 2017, had 65 million parameters, GPT3, which was trained up until mid-2021, had [175 billion parameters](https://beta.openai.com/docs/models/codex). And with their size, so too has their expressive power dramatically increased. ChatGPT was built on top of an updated version of GPT3 called [GPT3.5](https://techcrunch.com/2022/12/01/while-anticipation-builds-for-gpt-4-openai-quietly-releases-gpt-3-5/). This enormous expressive capacity, along with the data it was trained on (presumably [similar to GPT3](https://arxiv.org/pdf/2005.14165.pdf)), is what enables ChatGPT to be so generally capable, comprehending inputs that include a variety of human and programming languages, mathematical equations, and chord progressions. As a [generative model](https://en.wikipedia.org/wiki/Generative_model), ChatGPT is specifically designed to create new text based on the patterns it has learned from the training data. What makes this new model so fascinating is that it was [fine-tuned for conversation](https://openai.com/blog/chatgpt/). Not only does the chatbot give eerily human responses, it is able to integrate user-provided feedback to better accomplish the desired task. Of course, it isn’t perfect, but this functionality opens the door for a range of interactive applications. _FYI: three sentences in this section were written by ChatGPT. I’ll let you guess which ones!_ ## Where ChatGPT excels From what I can tell, there are three main categories of computer vision applications for which ChatGPT is fairly reliable: commonplace code, dressed individual method calls, and clean concatenations of simple components. ChatGPT’s responses to queries in any of these three categories benefit from being relatively self-contained. I’ll give examples for each. ### Commonplace code This should come as no surprise, but a generative model that was trained on a large corpus, including text and code, is generally very good at generating blocks of code that occur frequently and with little variation across the internet. When a code-based solution is essentially canonical (and likely omnipresent in the training data), ChatGPT’s probabilistic predilections will, with high probability, generate the tried and true result. Two examples of this include the fast Fourier transform (FFT), which is ubiquitous in signal processing and computer vision (among other fields): ```python 1import numpy as np 2 3def fft(x): 4 """Compute the discrete Fourier Transform of the 1D array x""" 5 x = np.asarray(x, dtype=float) 6 N = x.shape[0] 7 8 if N <= 1: # base case 9 return x 10 11 # split the array into even and odd components 12 even = fft(x[0::2]) 13 odd = fft(x[1::2]) 14 15 # compute the FFT of the even and odd components 16 T = [np.exp(-2j*np.pi*k/N)*odd[k] for k in range(N//2)] 17 return np.concatenate([even + T, even - T]) ``` and pooling layers, which are essential ingredients in any convolutional neural network (CNN): ```python 1def min_pooling(inputs, pool_size): 2 # Get the dimensions of the input array 3 n_rows, n_cols, n_channels = inputs.shape 4 5 # Calculate the size of the output array 6 out_rows = n_rows // pool_size 7 out_cols = n_cols // pool_size 8 9 # Create an empty array to store the output 10 outputs = np.zeros((out_rows, out_cols, n_channels)) 11 12 # Iterate over the input array and compute the output 13 for r in range(out_rows): 14 for c in range(out_cols): 15 for k in range(n_channels): 16 # Extract a patch of the input array 17 patch = inputs[r*pool_size:(r+1)*pool_size, c*pool_size:(c+1)*pool_size, k] 18 19 # Compute the min of the patch 20 outputs[r, c, k] = np.min(patch) 21 22 return outputs ``` Of course, these implementations are not always efficient, as in the case of the `min_pooling()` function above, which uses nested for loops. With semi-black box AI models like ChatGPT, it isn’t always clear if a generated solution is intentionally naive for the sake of simplicity and instruction, or if it genuinely selected what it believed to be the optimal solution. In scenarios like this, you can use the feedback oriented nature of ChatGPT to request improvements. In the case of pooling, for instance, ChatGPT can be instructed to parallelize the code or vectorize the for loops. Knowing what improvements to ask for is absolutely essential! ### Dressed method calls In a similar vein, ChatGPT has a pretty good grasp on all of the common machine learning and computer vision libraries (up until 2021), including PyTorch, TensorFlow, Scikit-learn, PIL, Skimage, and OpenCV. The chatbot is at its best when it can dress up methods from these libraries with the appropriate (boilerplate) preprocessing steps, such as input-output handling, converting a color image to grayscale, and reshaping arrays. Some examples of this, from simplest to most complex, are sharpening an image: ```python 1from PIL import Image, ImageFilter 2 3# Open the image 4im = Image.open('image.jpg') 5 6# Apply a sharpening filter 7sharpened_im = im.filter(ImageFilter.SHARPEN) 8 9# Save the sharpened image 10sharpened_im.save('sharpened_image.jpg') ``` Applying Otsu thresholding: ```python 1from skimage import io 2from skimage.filters import threshold_otsu 3 4# Load the image 5image = io.imread('image.png') 6 7# Convert the image to grayscale 8gray_image = np.mean(image, axis=2) 9 10# Calculate the Otsu threshold 11threshold = threshold_otsu(gray_image) 12 13# Apply the threshold to the image to create a binary image 14binary_image = gray_image > threshold ``` And computing the tail-distributed stochastic neighbor embedding (t-SNE) for a collection of images: ```python 1import numpy as np 2from scipy.spatial.distance import pdist, squareform 3from sklearn.manifold import TSNE 4from scipy.spatial.distance import pdist, squareform 5from PIL import Image 6 7def tsne(images, perplexity=30.0, learning_rate=200.0, n_components=2): 8 # Convert the images to grayscale 9 grayscale_images = [np.array(image.convert('L')) for image in images] 10 11 # Stack the grayscale images into a single array 12 stacked_images = np.stack(grayscale_images) 13 14 # Flatten the stacked images into a single array 15 flattened_images = stacked_images.reshape(len(images), -1) 16 17 # Compute the pairwise distances between the flattened images 18 distances = squareform(pdist(flattened_images, 'euclidean')) 19 20 # Compute the low-dimensional embedding of the images using t-SNE 21 tsne = TSNE(perplexity=perplexity, learning_rate=learning_rate, n_components=n_components, init='random') 22 embedding = tsne.fit_transform(distances) 23 24 return embedding ``` Unfortunately, ChatGPT is not adept or even conversant in newer CV libraries like [FiftyOne](https://github.com/voxel51/fiftyone/tree/develop/fiftyone). ### Concatenating simple components While ChatGPT can definitely behave irrationally, the model has basic logical capabilities that enable it to synthesize multiple (simple) inputs into composite outputs: just ask ChatGPT a few questions of the form _“what do X and Y have in common”?_ As an example, of this I asked ChatGPT to do the following: _“Write code to resize two images to (100,200), augment the dataset with 6 more images, blend all the input images into one, then plot all the images in a grid. Write the text of this prompt on the third plot”_ It’s solution was acceptable, effectively without modification: ```python 1import matplotlib.pyplot as plt 2from skimage import io, transform 3from skimage.util import random_noise 4 5# load and resize images 6img1 = io.imread("image1.png") 7img1 = transform.resize(img1, (100, 200)) 8img2 = io.imread("image2.png") 9img2 = transform.resize(img2, (100, 200)) 10 11# augment dataset with 6 additional images 12augmented_imgs = [] 13for i in range(6): 14 augmented_imgs.append(random_noise(img1)) 15 augmented_imgs.append(random_noise(img2)) 16 17# blend images 18blended_img = img1 * 0.5 + img2 * 0.5 19for img in augmented_imgs: 20 blended_img += img * 0.5 / len(augmented_imgs) 21 22# plot all images in a grid 23fig, ax = plt.subplots(nrows=3, ncols=3) 24ax[0, 0].imshow(img1) 25ax[0, 1].imshow(img2) 26ax[0, 2].imshow(blended_img) 27for i in range(6): 28 ax[1 + i // 3, i % 3].imshow(augmented_imgs[i]) 29 30# add text to third plot 31ax[0, 2].set_title("Resized, augmented, and blended images") ``` ### Surprising successes Other tasks it was able to navigate without issue include 1. Performing segmentation on 3d point clouds (it used DBSCAN) 2. Using DeOldify to restore color in old images 3. Perform object tracking in videos ## Where it falters As with any new technology, failure modes are potentially limitless. This discussion is not intended to be exhaustive, and I won’t even touch on general concerns such as [prompt-injection attacks.](https://research.nccgroup.com/2022/12/05/exploring-prompt-injection-attacks/) After asking ChatGPT computer vision tasks for multiple days, there seem to be a few recurring issues: long-tail scenarios, mathy manipulations, and expansive code blocks. As in the previous section, I’ll give examples of each. ### Look out for the long tail In any field, there will be a variety of tasks that are staples of certain subfields, but are dwarfed by more common motifs in the sprawling corpora employed in training LLMs. ChatGPT has its fair share of trouble with these domains, and can be quite sensitive to minutiae when prompted on niche subjects. One word can mean the difference between a desired result, and an idea getting lost in the recesses of ChatGPT’s immense representational structure. An example of this is 3D computer vision, which is a small but strong subfield of computer vision concerned with three dimensional spatial data. ChatGPT has a working knowledge of point clouds, _“A point cloud is a set of points in 3D space that represent the surface of an object. Each point in the point cloud is defined by its 3D coordinates (x, y, and z) and may also include additional information such as color or normal vector. Point clouds are commonly used in computer vision and 3D modeling applications”,_ And when I ask it to _“write code to define the PointNet architecture, which acts on point clouds”_, its response includes the symmetric max pooling function, which allows for the model to act directly on unstructured point cloud data. ```python 1import tensorflow as tf 2 3# Define the input layer 4input_layer = tf.keras.layers.Input(shape=(None, 3)) 5 6# Use a shared MLP to process the input point cloud 7mlp1 = tf.keras.layers.Dense(64, activation="relu") 8mlp2 = tf.keras.layers.Dense(64, activation="relu") 9mlp3 = tf.keras.layers.Dense(64, activation="relu") 10x = mlp1(input_layer) 11x = mlp2(x) 12x = mlp3(x) 13 14# Use max pooling to get the features of the entire point cloud 15global_features = tf.keras.layers.GlobalMaxPooling1D()(x) 16 17# Use a fully connected layer to predict the output 18output_layer = tf.keras.layers.Dense(units=1, activation=None)(global_features) 19 20# Create the model 21model = tf.keras.Model(inputs=input_layer, outputs=output_layer) ``` _“This code defines a PointNet architecture that takes in a point cloud as input and uses a shared multi-layer perceptron (MLP) to process each point independently. The output of the MLP for each point is then passed through a global max pooling layer, which extracts the features of the entire point cloud. Finally, a fully connected layer is used to predict the output based on the global features of the point cloud._ _This is just one possible implementation of the PointNet architecture, and there are many different variations and modifications that can be made to the architecture depending on the specific requirements of your project.”_ When I instead prompt ChatGPT with just _“write code for the PointNet model architecture”_, omitting explicit mention of point clouds, I get the following in response: ```python 1import torch 2from torch import nn 3 4class PointNet(nn.Module): 5 def __init__(self): 6 super(PointNet, self).__init__() 7 8 # define the neural network layers 9 self.fc1 = nn.Linear(3, 64) 10 self.fc2 = nn.Linear(64, 64) 11 self.fc3 = nn.Linear(64, 64) 12 self.fc4 = nn.Linear(64, 128) 13 self.fc5 = nn.Linear(128, 1024) 14 self.fc6 = nn.Linear(1024, 512) 15 self.fc7 = nn.Linear(512, 256) 16 self.fc8 = nn.Linear(256, 128) 17 self.fc9 = nn.Linear(128, 64) 18 self.fc10 = nn.Linear(64, 10) 19 20 def forward(self, x): 21 # apply the neural network layers to the input point cloud 22 x = F.relu(self.fc1(x)) 23 x = F.relu(self.fc2(x)) 24 x = F.relu(self.fc3(x)) 25 x = F.relu(self.fc4(x)) 26 x = F.relu(self.fc5(x)) 27 x = F.relu(self.fc6(x)) 28 x = F.relu(self.fc7(x)) 29 x = F.relu(self.fc8(x)) 30 x = F.relu(self.fc9(x)) 31 x = self.fc10(x) 32 return x ``` _“This code defines a PointNet class that extends the `nn.Module` class from PyTorch. The PointNet class defines a neural network with 10 fully connected layers, which are defined in the `__init__()` method of the class. The forward method applies these layers to an input point cloud and returns the output of the model._ _To use the PointNet model, we would create an instance of the PointNet class and then call the forward method on an input point cloud. The output of the model would be a tensor of shape (1, 10), which contains the predictions made by the model on the input point cloud”._ Despite the chatbot explicitly referencing point clouds in its response, a symmetric function is nowhere to be found in the code. This is despite [the paper](https://arxiv.org/abs/1612.00593) introducing PointNet garnering almost 10,000 citations, the [official GitHub repository](https://github.com/charlesq34/pointnet) accumulating 4000+ stars, and results related to the (correct) PointNet model single handedly [monopolizing the first page of Google’s search results](https://www.google.com/search?q=pointnet&oq=pointnet&aqs=chrome.0.69i59l2j0i512l3j69i60l3.1529j0j7&sourceid=chrome&ie=UTF-8). ### Math manipulations get messy While ChatGPT is capable of solving surprisingly involved standalone math problems and even generating its own proofs, the chatbot struggles to string together multiple mathematical lines of code. For example, ChatGPT fails when asked to implement [Otsu thresholding](https://en.wikipedia.org/wiki/Otsu%27s_method) or the [Hough transform](https://en.wikipedia.org/wiki/Hough_transform) from without using computer vision libraries. Otsu threshold: ```python 1def otsus_method(img): 2 # calculate histogram of grayscale values 3 histogram = np.zeros(256) 4 for i in range(img.shape[0]): 5 for j in range(img.shape[1]): 6 histogram[img[i, j]] += 1 7 8 # normalize histogram 9 histogram = histogram / img.size 10 11 # calculate cumulative sum of normalized histogram 12 cumulative_sum = np.cumsum(histogram) 13 14 # calculate cumulative sum of squares of normalized histogram 15 cumulative_sum_of_squares = np.cumsum(histogram ** 2) 16 17 # initialize variables 18 optimal_threshold = 0 19 max_variance = 0 20 21 # iterate over all possible threshold values 22 for i in range(256): 23 # calculate background and foreground probabilities 24 p_background = cumulative_sum[i] 25 p_foreground = 1 - p_background 26 27 # calculate means 28 mean_background = np.sum(np.arange(256) * histogram[:i+1]) / p_background 29 mean_foreground = np.sum(np.arange(i+1, 256) * histogram[i+1:]) / p_foreground 30 31 # calculate between-class variance 32 variance = p_background * p_foreground * (mean_background - mean_foreground) ** 2 33 34 # update optimal threshold if necessary 35 if variance > max_variance: 36 max_variance = variance 37 optimal_threshold = i 38 39 return optimal_threshold ``` ![](https://cdn.sanity.io/images/h6toihm1/production/606c4f1b72cb411409e5af91a8d4e375d243d341-1400x258.png?auto=format&dpr=2&fit=max&q=75&w=1400) Hough transform: ```python 1def hough_transform(img, theta_resolution=1, rho_resolution=1): 2 # convert image to grayscale 3 img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) 4 5 # detect edges in the image using Canny edge detection 6 edges = cv2.Canny(img, 50, 200) 7 8 # create accumulator array with dimensions determined by theta and rho resolutions 9 theta = np.linspace(-90, 90, 180 // theta_resolution) 10 rho = np.linspace(-np.sqrt(2) * img.shape[0], np.sqrt(2) * img.shape[0], 2 * np.sqrt(2) * img.shape[0] // rho_resolution) 11 accumulator = np.zeros((len(theta), len(rho))) 12 13 # find non-zero edge pixels in the image 14 y_idxs, x_idxs = np.nonzero(edges) 15 16 # convert x and y indices to cartesian coordinates 17 points = np.stack([x_idxs, y_idxs], axis=1) 18 19 # iterate over all edge points 20 for x, y in points: 21 # iterate over all theta values 22 for i, t in enumerate(theta): 23 # calculate rho for current theta value 24 r = x * np.cos(t) + y * np.sin(t) 25 26 # find closest rho value and increment corresponding accumulator bin 27 j = np.argmin(np.abs(rho - r)) 28 accumulator[i, j] += 1 29 30 return accumulator, theta, rho ``` ![](https://cdn.sanity.io/images/h6toihm1/production/95d7d19aedd8491f75254e026312d21d5829fc93-1400x515.png?auto=format&dpr=2&fit=max&q=75&w=1400) In many of these cases, the high-level logical flow isn’t necessarily wrong — or at least isn’t too far off. Instead, the choices made on one line seem to conflict with choices made on following lines: the attempted Otsu thresholding fails when it tries to broadcast together operands of incompatible shapes, and the Hough transform falters with data types. Some of these mathematical mishaps are easily corrected, whether by converting an int to a float, or rewriting a single faulty line of code. Other times these surface-level errors may be concealing more clandestine complications. If you’re going to use ChatGPT for these types of problems, proceed with due caution. The more modular the code you can get ChatGPT to write, the better. ### Expansive solutions breed errors It may not be too big of a surprise, but the more code that is generated, the more likely it seems that there will be an error (or multiple errors). After all, one could easily argue that humans are the same. For transformer models, it is a [well-known problem](https://arxiv.org/abs/2107.03448) that coherence (and model performance) diminishes as the number of “blocks” in the output increases. Blocks containing computer vision code are no exception. When I asked ChatGPT to “ _implement optical character recognition (OCR) from scratch”_, for instance, the model failed irredeemably. This failure was more of a complete mishmash than a series of understandable mathematical mistakes. Of course, that doesn’t mean ChatGPT is unable to perform OCR. When I prompted the chatbot with the more lenient _“apply optical character recognition”_, it showed that at least it knows how to use the relevant method in `pytesseract`: ```python 1import pytesseract 2 3# Load the license plate image 4# image = cv2.imread("license_plate.png") 5 6# Convert the image to grayscale 7gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) 8 9# Apply Otsu's thresholding to binarize the image 10threshold, binary = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY+cv2.THRESH_OTSU) 11 12# Use Tesseract to perform OCR on the license plate image 13text = pytesseract.image_to_string(binary, lang="eng") 14 15# Print the recognized text 16print(text) ``` The problem often rears its ugly head when ChatGPT is constructing neural networks in PyTorch, TensorFlow, or Sklearn, where it may string together a bunch of layers in a way that does not accomplish anything of use. ## Where to exercise extreme caution The previous three categories only scratch the surface of where things can go wrong using ChatGPT for computer vision. If you’re not careful, the AI agent might feed you egregiously erroneous answers. When I asked ChatGPT to give me the _“inference time verse year released for the 10 best object detection models”_, I received these details in response: ```python 1# Define the names and years of the 10 best object detection models 2model_names = ["Faster R-CNN", "YOLOv3", "SSD", "RetinaNet", "Mask R-CNN", "R-FCN", "SSD with FPN", "YOLOv4", "Grid R-CNN", "FCOS"] 3model_years = [2015, 2018, 2016, 2017, 2017, 2016, 2017, 2020, 2019, 2019] 4 5# Define the inference times of the 10 best object detection models (in milliseconds) 6model_inference_times = [100, 30, 20, 40, 60, 50, 25, 15, 35, 45] ``` Upon cursory glance, this might appear reasonably trustworthy. The models are all fairly well-known in object detection — and “best” is somewhat subjective anyway. The years check out too. Where this all breaks down is the inference times. These times are pulled from a variety of sources, all with different conditions and assumptions. ChatGPT’s response makes it seem like they can be compared apples to apples, which is patently false. At the end of the day, the burden of discerning and recognizing truth lies with the human asking and interpreting ChatGPT’s responses. Asking a vague question about inference times is a bad start. But even more well-posed questions can give incorrect answers. In a similar vein, you should probably be cautious whenever ChatGPT spits out specific numeric values that aren’t explicitly given as an example. Sometimes these values are appropriate, as when it knew to normalize AlexNet with given mean and standard deviation: ```python 1transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225] ``` When I asked ChatGPT to perform depth perception using stereo images, it included a line of code that read: ```python 1# Calculate the depth of each pixel 2depth = 9.0 * 0.54 / (disparity + 0.00001) ``` When questioned why it chose these values, it clarified: _“This formula uses the known parameters of the stereo camera system, such as the baseline (the distance between the two cameras) and the focal length, to convert the disparity value (which is the difference in horizontal position of a pixel in the left and right images) into a depth value._ _The specific values used in this formula (9.0 and 0.54) are just examples, and they may not be appropriate for all stereo camera systems. In general, the formula for calculating depth from disparity is:_ ```python 1depth = baseline * focal_length / (disparity + epsilon) ``` Where `baseline` is the distance between the two cameras, `focal_length` is the focal length of the cameras, `disparity` is the disparity map computed from the stereo images, and `epsilon` is a small value used to avoid division by zero. The values of these parameters will depend on the specific stereo camera system that you are using”. Another area of concern is the evaluation of dataset and model quality. When tasked with writing code to _“evaluate the quality and diversity of an image dataset”_, ChatGPT responded with the following code: ```python 1import os 2from PIL import Image 3from sklearn.neighbors import NearestNeighbors 4 5# Define the path to the dataset directory 6dataset_dir = './dataset' 7 8# Load the images in the dataset 9images = [] 10for filename in os.listdir(dataset_dir): 11 # Load the image 12 image = Image.open(os.path.join(dataset_dir, filename)) 13 14 # Convert the image to grayscale and resize it to a fixed size 15 image = image.convert('L').resize((128, 128)) 16 17 # Add the image to the list of images 18 images.append(image) 19 20# Flatten the images into a single array 21flattened_images = np.stack([np.array(image) for image in images]).reshape(len(images), -1) 22 23# Compute the nearest neighbors of the flattened images 24nbrs = NearestNeighbors(n_neighbors=5).fit(flattened_images) 25distances, indices = nbrs.kneighbors(flattened_images) 26 27# Compute the mean and standard deviation of the distances to the nearest neighbors 28mean_distance = np.mean(distances[:, 1:]) 29std_distance = np.std(distances[:, 1:]) 30 31print('Mean distance to nearest neighbor:', mean_distance) 32print('Standard deviation of distances:', std_distance) ``` Along with an accompanying exposition: _“In this example, the images variable is a list of images in the dataset. The images are first loaded from the files on disk, then converted to grayscale and resized to a fixed size using the convert and resize methods from the `PIL.Image` module. The images are then flattened into a single array using the stack and reshape methods from the numpy module. The `nbrs` variable computes the nearest neighbors of the flattened images using the `NearestNeighbors` class from the `sklearn.neighbors` module. The `mean_distance` and `std_distance` variables compute the mean and standard deviation of the distances to the nearest neighbors, which can be used as metrics to evaluate the quality and diversity of the dataset. A lower mean distance and a higher standard deviation indicate a higher-quality and more diverse dataset.”_ While this approach is not _necessarily_, or _technically_ wrong, it is at the very least unsophisticated. I received a similarly uninspiring answer when I asked ChatGPT to diagnose distribution shift in an image dataset, and it used a t-test (a test which just compares the means of two groups of samples) on the raw pixel values. You can also see that ChatGPT attempted to compute the mean squared error between two sets of images, which only makes sense if the sets have the same size, but we will give the chatbot the benefit of the doubt here and move on. ```python 1# Import necessary libraries 2import numpy as np 3from scipy.stats import ttest_ind 4from sklearn.metrics import mean_squared_error 5 6# Load the original and new datasets 7original_data = ... 8new_data = ... 9 10# Calculate the mean squared error between the original and new images 11mse = mean_squared_error(original_data, new_data) 12 13# Use a t-test to compare the means of the original and new datasets 14t, p = ttest_ind(original_data, new_data) 15 16# Print the results 17print("Mean squared error:", mse) 18print("T-test results: t =", t, "p =", p) 19 20# If the mean squared error is large and the t-test p-value is small, 21# this indicates a significant distribution shift in the data ``` Diagnosing distribution shifts in machine learning is in general a difficult problem, but at the very least, a basic understanding of the [various types of possible shifts](https://d2l.ai/chapter_linear-classification/environment-and-distribution-shift.html) is standard, and there are a variety of more nuanced approaches than applying a t-test. One last example is instructive: I prompted ChatGPT to “ _compute how mistaken my model is given a set of ground truth values, predictions, and prediction confidence scores”_, the crux of its response was _“​​you can use a metric like accuracy or precision to measure the performance of your model”_, which misses the mark entirely, and doesn’t use confidence scores at all. For reference, in the open source computer vision library [FiftyOne](https://github.com/voxel51/fiftyone/tree/develop/fiftyone), a quantity of this kind can be computed for a Dataset `dataset` by invoking `dataset.compute_mistakenness()`. ## Why ChatGPT empowers CV engineers Nevertheless, in spite of all these shortcomings, ChatGPT can still be a force for good in computer vision. At least that’s one person’s opinion. You might have noticed a theme in the examples I’ve included. ChatGPT is a tool that can help computer vision engineers and practitioners do their jobs more efficiently. It is a very powerful tool that can also be quite dangerous. And like any tool, it is only as good as its user. From a programming perspective, many of the same principles that ensure good human engineering can make for good machine-generated code: aim for modular design, decompose high-level tasks into pieces if possible, and leverage existing libraries instead of reinventing the wheel. If used effectively, ChatGPT can help CV engineers write boilerplate CV code more quickly, freeing up humans to focus on the important tasks of building high-quality datasets and evaluating ML models, using toolsets from companies like Scale, V7, and Voxel51. We are living in an exciting time. But be careful. ## Just for fun On a lighter note, no article on ChatGPT would be complete without highlighting just how peculiar, interesting, and often hilarious, its responses can be. Accordingly, I’ve prepared a few of my favorite responses that are — at least tangentially — related to computer vision: **Is a picture actually worth a thousand words?** _“It is often said that a picture is worth a thousand words, meaning that a single image can convey a complex idea or story more effectively than a lengthy explanation. This is because pictures can capture the attention of the viewer and convey information quickly and efficiently._ _However, the exact value of a picture in terms of words can vary depending on the context and the audience. For example, a picture that is rich in detail and visual information may be able to convey more than a thousand words, while a simpler image may be less effective. Additionally, the ability of a picture to convey meaning also depends on the viewer’s familiarity with the subject matter and their ability to interpret the image._ _Overall, while a picture may be able to convey a large amount of information quickly and effectively, its value in terms of words can vary and may not always be as high as a thousand”._ **Write a poem about open source computer vision package FiftyOne** _“It’s easy to use and highly flexible,_ _With support for many different platforms and specs,_ _Whether you’re a researcher or a developer,_ _FiftyOne has something for every level._ _So if you’re looking for a powerful tool,_ _To help with your computer vision tasks,_ _Give FiftyOne a try,_ _And see how it can make your work more efficient and fast.”_ **What is the most significant image of all time?** _“It is difficult to determine the most significant image of all time, as the significance of an image can vary depending on the context and the perspective of the viewer. However, some images are widely considered to be significant due to their impact on history, culture, and society._ _One example of a significant image is the “Migrant Mother” photograph, taken by Dorothea Lange in 1936. The photograph depicts a mother and her children living in poverty during the Great Depression, and it became a symbol of the struggles faced by many Americans during this time._ _Another example of a significant image is the “Earthrise” photograph, taken by astronaut William Anders during the Apollo 8 mission in 1968. The photograph shows the Earth rising above the lunar surface, and it is credited with inspiring the environmental movement and raising awareness of the need to protect our planet._ _Overall, there are many images that can be considered significant due to their historical, cultural, or societal impact. The most significant image may vary depending on the perspective and context of the viewer”._ **Tell me a joke about computer vision** _“Q: Why was the computer cold?_ _A: Because it left its Windows open”._ **Make word art** _“Here is a word art design:_ \\_\\_ \\_\_ / / / / / / / / / / / / / / /\_\_ / / /\_\_ / / /\_\_\_\_ / / /\_\_\_\_ / / /\_\_\_\_\_\_ / /\_\_\_\_\_\_ / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / / _“This design incorporates the words ‘computer vision’ into a geometric pattern. The words are arranged in a way that creates a visual image of a computer screen with a vision-like pattern in the background”._ **NOTE**: this was after I had asked it a bunch of CV questions. As far as I can tell, the words “computer vision” are not present in that image. Also when I reset the ChatGPT thread, it told me it was not able to generate word art. ## FiftyOne Computer Vision toolset [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset developed by Voxel51 that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [ChatGPT](https://voxel51.com/blog/tag/chatgpt) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/770b8cfdbd7944916b1195dc11e5b173dbab8e97-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 and the State of Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 18, 2023](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ec2410f8c8844aa682fea044da0a14e3d9c5c7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ VoxelGPT: Your AI Assistant for Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jun 7, 2023](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/a735267ad7effa9f799f850ab7c8ffa241088710-1024x1024.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Why 2022 was the most exciting year in computer vision history (so far)\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 14, 2022](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-196-lllmstxt|> ## December 2022 Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup — December 2022 Dec 13, 2022 • 14 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#e612ba09b7a9) [Meetup Recap at a Glance](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#b600be759f66) [Talk #1: Wearable Vision Sensors Summary](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#6144a9c6b0db) [Talk #1 Video Replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#e2afe45f428b) [Talk #1 Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#f37c5743fdc1) [Talk #1 Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#b667f0a5194f) [Talk #2: The Future of Data Annotation: Trends, Problems & Solutions Summary](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#564cf744bcee) [Talk #2 Video Replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#be5237ac1119) [Talk #2 Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#4c779946e6bd) [Talk #2 Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#4ff87da21768) [Talk #3: Using Similarity Learning to Improve Data Quality Summary](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#a0a010dc4d30) [Talk #3 Video Replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#75beab6f9767) [Talk #3 Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#e4830a3c9df5) [Talk #3 Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#96d7c893c515) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#757af4ed9d7a) [Upcoming Computer Vision Meetup Speakers & Schedule](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#c3afe98a5d0e) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#ab201912f23a) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#e612ba09b7a9) [Meetup Recap at a Glance](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#b600be759f66) [Talk #1: Wearable Vision Sensors Summary](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#6144a9c6b0db) [Talk #1 Video Replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#e2afe45f428b) [Talk #1 Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#f37c5743fdc1) [Talk #1 Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#b667f0a5194f) [Talk #2: The Future of Data Annotation: Trends, Problems & Solutions Summary](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#564cf744bcee) [Talk #2 Video Replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#be5237ac1119) [Talk #2 Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#4c779946e6bd) [Talk #2 Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#4ff87da21768) [Talk #3: Using Similarity Learning to Improve Data Quality Summary](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#a0a010dc4d30) [Talk #3 Video Replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#75beab6f9767) [Talk #3 Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#e4830a3c9df5) [Talk #3 Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#96d7c893c515) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#757af4ed9d7a) [Upcoming Computer Vision Meetup Speakers & Schedule](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#c3afe98a5d0e) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#ab201912f23a) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/b5ed751b5c0fbc3d2cb74f0b30e6418d3319564c-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=1200) Last week Voxel51 hosted the December 2022 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/). What a wonderful event! Our amazing speakers shared insightful presentations, the virtual room was packed, and the Q&A was vibrant! In this blog post we provide the recordings, recap presentation highlights and Q&A, as well as share the upcoming Meetup schedule so that you can join us at a future event. Hope to see you soon! ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes was [Children International](https://www.children.org/). We are pleased to be making a donation of $200 to them on behalf of the computer vision community! ![](https://cdn.sanity.io/images/h6toihm1/production/d693c1ecf054c4bbc85b6c3212e12bd131e21c21-606x201.png?auto=format&dpr=2&fit=max&q=75&w=606) ## Meetup Recap at a Glance - Talk #1 — Kris Kitani // Wearable Vision Sensors [\- Presentation recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#vision-sensors-summary) [\- Video replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#vision-sensors-video) [\- Q&A recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#vision-sensors-q-a) [\- Additional resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#vision-sensors-resources) - Talk #2 — Anna Petrovicheva // The Future of Data Annotation: Trends, Problems & Solutions [\- Presentation recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#data-annotation-summary) [\- Video replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#data-annotation-video) [\- Q&A recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#data-annotation-q-a) [\- Additional resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#data-annotation-resources) - Talk #3 — Kacper Łukawski // Using Similarity Learning to Improve Data Quality [\- Presentation recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#similarity-learning-summary) [\- Video replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#similarity-learning-video) [\- Q&A recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#similarity-learning-q-a) [\- Additional resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#similarity-learning-resources) - [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#locations) - [Computer Vision Meetup Speakers — January, February, March](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#schedule) - [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022#get-involved) ## Talk \#1: Wearable Vision Sensors Summary [Kris Kitani](https://www.linkedin.com/in/kriskitani/), associate research professor of the [Robotics Institute at Carnegie Mellon University](https://www.ri.cmu.edu/), shares some of the research and projects his lab is working on centered around the concept of wearable vision sensors. Kris explains that computer vision is currently undergoing a paradigm shift from third-person generated data, such as somebody uploading pictures that they took, to first-person data where the sensors themselves are on the user. He then shows three early prototype systems based on wearable sensors. First, Kris walks through a learning based model where the goal is to accurately estimate a person’s hand pose using a smartwatch. This model’s inputs include hand and motion history images, and the output is 19 different joint positions and their angles to describe the joints on a hand. Then, using inverse kinematics to solve for what the hand might look like, the estimated hand pose can be visualized using computer graphics. Kris moves onto the next model, which uses a chest mounted camera for pose estimation. This model centers around three streams: one that estimates how the chest camera is moving, another that is looking at how the pose is changing relative to the camera (joints, angles, etc.), and a third stream that is focused on which way the head is looking in 3D. Kris shows a demo video visualizing a few scenes reconstructed from a chest mounted camera and this model. The last model Kris that describes uses a head mounted camera for pose estimation. In this model, the input is the egocentric, forward-facing camera, and a detected object. The output is an object-aware human pose estimate of the person wearing the camera. Kris shows demo videos of what the camera is seeing, alongside a computer-generated reconstruction of what that person is doing. Learn more and see it all in action in the video replay. ## Talk \#1 Video Replay https://www.youtube.com/watch?v=HZw3aeZfX-c ## Talk \#1 Q&A Recap Here’s a recap of the live Q&A from this presentation during the virtual Computer Vision Meetup: **Can you use hand pose estimation to determine if someone pulled a trigger or not?** When you’re holding an object with a trigger, there’s only so much that your hand can do, relative to other examples where the hand has total free motion. So I would assume that hand pose estimation would actually work pretty well in this situation. Incorporating auditory information would also increase recognition. **Have you tested hand pose estimation out for different sized people?** We have not done a full, robust study that would be needed to make this a commercial product. So far, we’ve just tried it on a small group of graduate students. However, I think it’s fairly robust and with enough data, yes, it should be able to be generalized to more people if done properly. **What minimum camera resolution do you need?** We haven’t run that experiment. So far we’ve been using GoPro cameras for most of these experiments, so that will give you the range of the resolutions that we’ve been working with. **To keep ethical AI in perspective, would a robust study also include people of different ethnicity (aka skin color)?** Yes, of course. One of the questions earlier kind of alluded to that as well. Different body sizes, different skin colors would definitely need to be taken into consideration for commercial development. **Can the chest mounted camera be used to know if a person is following you?** As originally scoped, I would think not, but there might be an interesting way to do that. **How can you hallucinate or predict what an arm is doing that is out of the field of view?** That’s essentially what we’re doing with the physics simulation: we try to hallucinate what the other limbs are doing even though we can’t see them. **How are these simulated environments generated?** You could use any physics simulator that you want, but the way we’re generating it is that we’re running object detection. In newer methods, we’re doing a 3D reconstruction and then inserting these synthetic objects into the scene and then we use them when we’re running the physics simulation. There’s a whole stream of computer vision work on just scene understanding and scene reconstruction, which could be leveraged, but we’re not going that deep into it at the moment. **Is the inferencing and computation done on an edge device or on a server?** All of these are running in real time, using the device and also a desktop computer. Because they run in real time, I could conceivably run them on the embedded device, but we are not running it like that right now. **Have you thought about predicting the next move of the user instead of just recognizing the motion in the current timestamp?** Yes, I have a stream of work on trajectory forecasting and activity forecasting, which I didn’t talk about today. We do have an ACCV paper and an ICCV paper on using a wearable camera to try to predict what people are doing. **The emphasis on moving from third to first-person view is fabulous, but so much of vision for the last 20–30 years has been third-person view. In your experience so far, what is most different about doing computer vision in ego (first person) versus exo (third person)?** With a wearable first-person camera, we have a much higher resolution video of people with their hands interacting with objects, which opens up a whole new area of what it means to understand people when they’re interacting with objects. Grasp taxonomies, pressure points, detailed manipulation — it’s actually kind of intersecting with a lot of robotics research. So not only is it different, it’s opening up a whole new world of research topics because of this new perspective that we have. ## Talk \#1 Additional Resources Check out this additional resource on the presentation: - [Talk transcript](https://www.rev.com/transcript-editor/shared/a2g_dWAnD7wwf_1WWgZijhy1q8hPHQ0S064GB6YOO-SCNCuHZkp8b5lRGkDLs5mCjZDkWALrOoh0C3ahWEtngE273Fk?loadFrom=SharedLink) Thank you Kris on behalf of the entire Computer Vision Meetup community for sharing your research and early prototype systems that will inspire and influence commercial developments yet to come! ## Talk \#2: The Future of Data Annotation: Trends, Problems & Solutions Summary [Anna Petrovicheva](https://www.linkedin.com/in/anna-petrovicheva-44b24673/) is the CEO at [CVAT.ai](https://www.cvat.ai/) and CTO at [OpenCV.ai](https://www.opencv.ai/) presented on the future of data annotation. The data annotation market is growing at 26% per year. This rapid growth rate reflects the importance of categorizing and labeling data in the AI-driven world. In the talk, Anna shared five major trends related to data annotation and what innovations we can expect in the upcoming years: 1. Working with datasets is an iterative process. In the past, you mostly focused on a dataset one time. Now, people typically think about it as a “first chunk of the data, and follow on chunks, they’re different,” said Anna. The first chunk you want to be big, meaningful, and diverse; it typically covers 70% of your use case. When you deploy in production, the follow on chunks cover the 30% of the business case that you didn’t predict. beforehand. 2. Going beyond 3D. Now, there is an abundance of LiDAR data available, as well as radar data from satellites, and medical imaging from MRI scans. You will see data annotation moving to support more and more types of data, including a variety of 3D data types, so that they are easy to work with. 3. Data annotator is a profession. When the AI boom started, data scientists were annotating their data. Now there are in-house teams, data agencies, and even crowdsourcing options available for data annotation. 4. Data is moving to open source tools. One trend Anna sees is that open source tools will be as prevalent or even more popular than closed source solutions. Examples: CVAT and FiftyOne are both open source and available on GitHub, and there are many more. 5. There will be more data flows and flexible integrations. The industry is moving from manual dataset management to automated, reliable data flows available anytime. Data flows vary by company, but there are starting to be common tools underpinning data flows, such as tools like CVAT and FiftyOne, that can be connected via APIs. ## Talk \#2 Video Replay https://www.youtube.com/watch?v=7V\_nCO9bkN0 ## Talk \#2 Q&A Recap **What are the most popular annotation formats?** There are a lot of them and they very much depend on the type of the data and type of the annotation that you are working with. For example, if we are talking about bounding boxes, probably the most popular format is COCO format, which encodes the coordinates of the bounding boxes on each image. There are other formats for data segmentation, for polygon annotation, or annotation in 3D. So, it depends on the data. You can check the supported formats in CVAT in our README on GitHub. **Is there version control for annotations?** Yes, there are several tools for data versioning. My favorite one is called DVC, data version control. It’s also a fellow open source tool: [https://dvc.org](https://dvc.org/) ## Talk \#2 Additional Resources Check out these additional resources: - [Presentation slides](https://www.slideshare.net/MichelleBrinich1/computer-vision-meetup-the-future-of-data-annotation) - [Talk transcript](https://www.rev.com/transcript-editor/shared/twbuap8FZvqtqIt7iSlJ5i-xW6jsH0ZiOIgfNfvxJB18yZuHe8ZItg0Tp9tgB9rUMW2avdsVRFC-XyoVRJM_eMhJOWE?loadFrom=SharedLink) A big thank you to Anna on behalf of the entire Computer Vision Meetup community for sharing the important trends in data annotation and the AI industry more broadly! ## Talk \#3: Using Similarity Learning to Improve Data Quality Summary [Kacper Łukawski](https://www.linkedin.com/in/kacperlukawski/), developer advocate at [Qdrant](https://qdrant.tech/), gave a talk about similarity learning and how it can be used to improve data quality, in particular in the context of image-based tasks. Real-world data is a living structure, it grows day by day, changes a lot, and becomes harder to work with. The process of splitting or labeling is error-prone and these errors can prove to be very costly. Similarity learning is not only a great tool for data classification, even in the case of extreme classification, but also a surprisingly good way to improve the quality of your data. Kacper takes us on a journey to understand the inner workings of similarity learning, starting with defining traditional neural networks that have been designed to solve classification problems. He uses this as a foundation to thoroughly compare and contrast similarity learning. Kacper then provides an example of using similarity learning in action with a web furniture marketplace that includes images of products, as well as their names and the category they belong to. In this example, the marketplace has decided to automate the process of assigning a product to a category in order to determine whether or not a new category needs to be created to have a better user experience, or simply to sell a little more by identifying and correcting products with wrong categories. Kacper shows various ways similarity learning can be used to automate the process and find the items that are assigned incorrectly. You can dive into all of the details in the replay. ## Talk \#3 Video Replay https://www.youtube.com/watch?v=eQYz2dxqWHk ## Talk \#3 Q&A Recap **Can you please elaborate on differences between fine tuning the last layer and performing a clustering on embedding space based on distance?** Actually, these are two steps in the same process. First, if you take a pre-trained model, let’s say trained on ImageNet, and you want to use it in a different domain, then the original embeddings won’t be able to capture the similarities between those new objects that well. So the goal of fine tuning the last layer is to adjust the original embeddings in a way that those similar points will be closer to each other than the originals would be. Comparing the distance could be done on those original embeddings, but they do not provide the best possible quality out of the box. So, in this approach, we started with fine tuning and then compared the embeddings after that. **What are the advantages of Similarity Learning by looking at the embedding space rather than the final classification (of the network) after softmax?** You could probably use the final output of that network. But since it was done to solve a classification problem, those vectors won’t capture the regularities in your image data; instead, they will be trying to model the probability distribution of your classes. In the example I shared, we are not trying to solve the classification problem that the pre-trained model was trained for, instead we want to determine the content of the image, what it represents. Those previous vector representations are of a lower dimensionality, and they have a different data representation. This is where Similarity Learning comes in. **Is similarly learning compatible with transfer learning using pretrained models?** Similarity learning is similar to transfer learning, which takes an existing model and applies it to a different task or domain. The only difference here is that transfer learning doesn’t require any training. For example, transfer learning could be used to recognize the class on an object. But in our example, we don’t want to recognize a class, we want to have a data representation so that two objects that are similar in reality will also be similar in the embeddings space using a distance function, which is why we need to undergo a training effort and use similarity learning. **How would you optimize the similarity comparison to find outliers? Comparing every point with the other is computationally demanding with large data.** That would be an issue if we want to calculate the distance metrics to all the points, which is why we typically suggest focusing on specific subsets of the data. But of course, if you want to find some outliers in the whole set, you could approach it in a different way. Let’s say you have all your embeddings already collected, and if you want to find some outliers, you could use a vector database, such as Qdrant, and simply try to calculate the distance of every single point you have in a collection to some different points. A point which is an outlier will have the closest point quite far from it. But if the distances are quite low, then it may indicate there is a cluster in your data. And if you have a vector database like Qdrant that is already optimized, thanks to HNSW graphs you will be able to find the closest points with their distances quite easily, which will no longer require calculation of the distance metrics. **How is the anchor selected for the anchor-based similarity measure?** If you want to encode the category name, then this is already an anchor; we want to capture the points that have some encodings that are not close to the category name encoding. But with diversity search, we can select it randomly, because the assumption is that our outliers should be far away from every single item from the dataset that is not the outlier. **What are the differences between Qdrant and Milvus?** In a nutshell, Qdrant and [Milvus](https://milvus.io/) solve the same problem. At Qdrant we performed some benchmarks, which are available on our website. From my point of view, using Qdrant with some additional metadata stored for every single vector is way more efficient as compared to Milvus. But this is just my experience. I’m not an expert in Milvus personally. If you want to discuss more, please feel free to reach out and we can do a more detailed comparison. ## Talk \#3 Additional Resources Check out these additional resources on this talk: - [Presentation slides](https://www.slideshare.net/MichelleBrinich1/computer-vision-meetup-using-similarity-learning-to-improve-data-quality) - [Talk transcript](https://www.rev.com/transcript-editor/shared/0y5W1bW1YQqxGGZiNjnNghHQxTPoix-ao7IaHR79Lgs4IVlmta8BpZRsZRLtaqZqyRo8OHhbZ-LDBOOE5bn8P15m5MM?loadFrom=SharedLink) - [Article: Finding errors in datasets with Similarity Search](https://qdrant.tech/articles/dataset-quality/) - [Demo Application](https://dataset-quality.qdrant.tech/) Thank you to Kacper on behalf of the entire Computer Vision Meetup community for the very thorough exploration of using similarity learning to improve data quality. ## Computer Vision Meetup Locations Computer Vision Meetup membership has grown to more than [2,000+ members](https://www.meetup.com/pro/computer-vision-meetups/) in just a few months! The goal of the meetups is to bring together a community of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. If that’s you, we invite you to join the Meetup closest to your timezone: - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) ## Upcoming Computer Vision Meetup Speakers & Schedule We recently announced an exciting lineup of speakers for December, January, and February. Become a member of the Meetup closest to you, then register for the Zoom for the Meetups of your choice. ### January 12 - Hyperparameter Scheduling for Computer Vision — [Cameron Wolfe](https://www.linkedin.com/in/cameron-r-wolfe-04744a238/) (Alegion/Rice University) - An introduction to computer vision with Hugging Face transformers — [Julien Simon](https://www.linkedin.com/in/juliensimon/) (Hugging Face) - [Zoom Link](https://us02web.zoom.us/webinar/register/9616708727021/WN_XQIZlMP2RQuCFBwBylyRqA) ### February 9 - Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO — [Paula Ramos, PhD](https://www.linkedin.com/in/paula-ramos-41097319/) (Intel) - Understanding Speech Recognition with OpenAI’s Whisper Model — [Vishal Rajput](https://www.linkedin.com/in/vishal-rajput-999164122/) (AI-Vision Engineer) - [Zoom Link](https://us02web.zoom.us/webinar/register/9216708727643/WN_P8UHtAZGQWOx_A2HM1dcPA) ### March 9 - Lighting up Images in the Deep Learning Era — [Soumik Rakshit](https://www.linkedin.com/in/soumikrakshit/), ML Engineer (Weights & Biases) - Training and Fine Tuning Vision Transformers Efficiently with Colossal AI — [Sumanth P](https://www.linkedin.com/in/sumanth-p-09b339173/) (ML Engineer) - [Zoom Link](https://us02web.zoom.us/webinar/register/8816708728020/WN_mTdNXxTSR-e7bDG5EkH1XQ) ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. — _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [Carnegie Mellon](https://voxel51.com/blog/tag/carnegie-mellon) [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [CVAT](https://voxel51.com/blog/tag/cvat) [data annotation](https://voxel51.com/blog/tag/data-annotation) [Qdrant](https://voxel51.com/blog/tag/qdrant) [similarity learning](https://voxel51.com/blog/tag/similarity-learning) [wearable vision sensors](https://voxel51.com/blog/tag/wearable-vision-sensors) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Loading related posts... [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-197-lllmstxt|> ## FiftyOne Filtering Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Filtering Tips and Tricks — Dec 09, 2022 Dec 10, 2022 • 6 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#22caebfff7f8) [A filtering primer](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#00a5d6024cf5) [Filtering on tags, labels, frames, and keypoints](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#26f3770bc257) [Defining complicated filters](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#38b05b0e8827) [Composing filters](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#44bdd23600c4) [Accessing parent-level data](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#a6c385664501) [Filtering in the FiftyOne App](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#eb9bf8b72962) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#bc865aafbc26) [What’s next?](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#cde1e03338cd) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#22caebfff7f8) [A filtering primer](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#00a5d6024cf5) [Filtering on tags, labels, frames, and keypoints](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#26f3770bc257) [Defining complicated filters](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#38b05b0e8827) [Composing filters](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#44bdd23600c4) [Accessing parent-level data](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#a6c385664501) [Filtering in the FiftyOne App](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#eb9bf8b72962) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#bc865aafbc26) [What’s next?](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022#cde1e03338cd) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/204bd847400b1534c494dae594f0797457bc4c90-1620x906.png?auto=format&dpr=2&fit=max&q=75&w=1600) Welcome to our weekly FiftyOne tips and tricks blog where we give practical pointers for using FiftyOne on topics inspired by discussions in the open source community. This week we’ll cover [filtering](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#filtering). ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **A filtering primer** [Datasets](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#using-datasets) are the core data structure in FiftyOne, allowing you to represent your raw data, labels, and associated metadata. Querying a Dataset will return a [DatasetView](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) that contains only the samples (and possibly filtered contents) that match your query’s criteria. FiftyOne provides powerful [ViewField](https://voxel51.com/docs/fiftyone/api/fiftyone.core.expressions.html#fiftyone.core.expressions.ViewField) and [ViewExpression](https://voxel51.com/docs/fiftyone/api/fiftyone.core.expressions.html#fiftyone.core.expressions.ViewExpression) classes that allow you to use native Python operators to define the match/filter expressions that retrieve only the content of interest from your dataset. By filtering your data using ViewField and ViewExpression, you can create custom views of your dataset that employ comparison, logic, arithmetic or array operations. Continue reading for some tips and tricks to help you master filtering in FiftyOne! ## **Filtering on tags, labels, frames, and keypoints** In many cases FiftyOne’s generic `match()` method is what we want. Likewise, we do sometimes want to use the generic `filter_field()`. However, there are many cases where we might want to filter directly on special fields like tags, labels, or keypoints. FiftyOne provides built-in support for these operations with the following special filter and match operations: `filter_labels()`, `filter_keypoints()`, `match_labels()`, `match_tags()`, and `match_frames()` . These methods offer even easier ways to select the views we are interested in. Here are two equivalent ways of performing the same operation — creating a view of all samples which have the “test” tag: import fiftyone as fo import fiftyone.zoo as foz from fiftyone import ViewField as F dataset = foz.load\_zoo\_dataset("quickstart") \# Method 1: filter\_fields - verbose test\_view = dataset.filter\_field("tags", F().contains("test")) \# Method 2: match\_tags - concise test\_view = dataset.match\_tags("test") Learn more about [filter\_labels()](https://voxel51.com/docs/fiftyone/api/fiftyone.core.view.html#fiftyone.core.view.DatasetView.filter_labels), [match\_labels()](https://voxel51.com/docs/fiftyone/api/fiftyone.core.view.html#fiftyone.core.view.DatasetView.match_labels), and [match\_tags()](https://voxel51.com/docs/fiftyone/api/fiftyone.core.view.html#fiftyone.core.view.DatasetView.match_tags) in the FiftyOne Docs. ## **Defining complicated filters** While sometimes, the filtering operation one wants to perform can be stated quite simply, other times it is more involved. In these cases, even if all of the filtering is performed on a single field, it can be helpful to define the filter prior to its application, using FiftyOne’s ViewField. Suppose, for instance, that we want to create a view containing only those images in which we’ve detected more than two and fewer than six objects, unless the sample contains 10 detections, in which case we also want it to be part of our view. If we wanted to do this “inline”, this would be quite cumbersome: import fiftyone as fo import fiftyone.zoo as foz from fiftyone import ViewField as F dataset = foz.load\_zoo\_dataset("quickstart") view = dataset.match( ( (F("predictions.detections").length() < 6) & (F("predictions.detections").length() > 2) ) \| (F("predictions.detections").length() == 10) ) Instead, we could save ourselves the potential headaches arising from mismatched parentheses by defining the filter in pieces and then employing it. gt\_cond = F("predictions.detections").length() > 2 lt\_cond = F("predictions.detections").length() < 6 except\_cond = F("predictions.detections").length() == 10 view = dataset.match((gt\_cond & lt\_cond) \| except\_cond) Learn more about the FiftyOne [ViewField](https://voxel51.com/docs/fiftyone/api/fiftyone.core.expressions.html#fiftyone.core.expressions.ViewField) in the FiftyOne Docs. ## **Composing filters** In FiftyOne, you can also create composite filters which include filters on multiple different fields. This can be done by creating one view by applying one of the filters, and then creating a new view by applying a second filter to this first view. Alternatively, you can accomplish the same composite filtering operation inline without explicitly defining an intermediate view. For example, the two following approaches are equivalent: import fiftyone as fo import fiftyone.zoo as foz from fiftyone import ViewField as F dataset = foz.load\_zoo\_dataset("quickstart") \# Approach 1: intermediate view simple\_view = dataset.filter\_labels( "predictions", F("confidence") > 0.9 ) complex\_view = simple\_view.match( F("predictions.detections").length() > 2 ) \# Approach 2: composing filters complex\_view = dataset.filter\_labels( "predictions", F("confidence") > 0.9).match(F("predictions.detections").length() > 2 These two approaches are equivalent because in FiftyOne, a DatasetView is defined symbolically until operations on the view are performed. This is in contrast to similar filtering operations in pandas, where chained assignment can cause issues. Learn more about [views](https://voxel51.com/docs/fiftyone/api/fiftyone.core.view.html#module-fiftyone.core.view) in the FiftyOne Docs. ## **Accessing parent-level data** The symbolic nature of filtering queries in FiftyOne also makes it possible for queries on embedded fields to take information from their parent fields into account. The `$` symbol prepended to a field name signifies that it refers to parent-level data. This functionality can be used to leverage sample-level metadata when performing filtering operations on individual detections. For example, FiftyOne stores bounding box coordinates as relative values in \[0, 1\]. Thus, if we want to select all predicted detections with medium-sized bounding boxes in terms of absolute number of pixels, we need to use the sample width and height, which are stored in the sample’s metadata. import fiftyone as fo import fiftyone.zoo as foz from fiftyone import ViewField as F dataset = foz.load\_zoo\_dataset("quickstart") dataset.compute\_metadata() \# Computes the area of each bounding box in pixels bbox\_area = ( F("$metadata.width") \\* F("bounding\_box")\[2\] \\* F("$metadata.height") \\* F("bounding\_box")\[3\] ) \# Only contains boxes whose area is between 32^2 and 96^2 pixels medium\_boxes\_view = dataset.filter\_labels( "predictions", (32\*\*2 < bbox\_area) & (bbox\_area < 96\*\*2) ) Learn more about [embedded fields](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#tips-tricks) and the `EmbeddedDocumentField` in the FiftyOne Docs. ## **Filtering in the FiftyOne App** While everything we’ve discussed so far pertains to the Python SDK, it is also possible to filter in the FiftyOne App! When you load a session, on the left hand side, you can filter by labels, sample id, evaluation results, or other fields in your dataset. ![](https://cdn.sanity.io/images/h6toihm1/production/547eb80a45b08aef59ee2004126b82d044336a4d-1019x806.gif?auto=format&dpr=2&fit=max&q=75&w=1019) For categorical data like labels, you can add any number of selections to the view, or select the toggle “exclude” to view all of the unselected options. For numerical fields, such as image “uniqueness” or prediction “confidence”, you can drag the left and right ends of the corresponding range bar to set the range of allowed values. You can also filter using the view bar at the top of the FiftyOne App. If you click in the “+ add stage” box on the left side of the view bar, you can select one of the filtering methods and enter the desired information. Using both the side bar and the view bar, you can compose multiple filtering operations to create complex views that will be displayed in the view grid of the FiftyOne App. Learn more about the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html) in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,200+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,300+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 208+ repositories - 51+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## What’s next? - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [FAQ](https://voxel51.com/blog/tag/faq) [filtering](https://voxel51.com/blog/tag/filtering) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/95b57b1e85f9838a5ad0f07ee3b457331d0dd178-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Nov 11, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 12, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/25a70e1c9c2cae7f9df18696e0e087dddf67707d-1200x679.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Sept 23, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 24, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/93682b6d528c21a04ca55782334799a0715028ea-1200x678.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 30, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 31, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-198-lllmstxt|> ## Data-Centric Machine Learning Tools [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tutorials](https://voxel51.com/blog/category/tutorials) CVAT <> FiftyOne: Data-Centric Machine Learning with Two Open Source Tools Nov 30, 2022 • 10 min read Article content In this article [‍Introduction](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#43069128098d) [Dataset Curation](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#0f4f43d14704) [Dataset Annotation](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#71692fba0392) [Dataset Improvement](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#1b4b16645621) [Next Steps](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#c68a9f50f07a) [Summary](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#0535a2d2a2b6) In this article [‍Introduction](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#43069128098d) [Dataset Curation](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#0f4f43d14704) [Dataset Annotation](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#71692fba0392) [Dataset Improvement](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#1b4b16645621) [Next Steps](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#c68a9f50f07a) [Summary](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools#0535a2d2a2b6) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) _This article was written in collaboration with members of the CVAT team and was originally published on [cvat.ai](https://www.cvat.ai/post/data-centric)._ ![](https://cdn.sanity.io/images/h6toihm1/production/42757251267e8803167af9f5cc5b6d2ae1bade62-1400x347.jpg?auto=format&dpr=2&fit=max&q=75&w=1400) **TL;DR**: The quality of Deep Learning-based algorithms strongly depends on the quality of training data employed. This is especially true in the Computer Vision domain. Poor data quality leads to worse predictions, increased training times, and the need for bigger datasets. FiftyOne and CVAT can be used together to help you produce high-quality training data for your models. Keep reading to see how! ## ‍Introduction Recently, the “Data-Centric movement” has been gaining popularity in the machine learning space. Over the last decade, improvements in machine learning primarily focused on models, while datasets remained largely fixed. As a community, we looked for better network architectures, created scalable models, and even implemented automatic architecture search. At present, however, the performance of our increasingly powerful models is limited by the datasets on which they are trained and validated.‍ In practice, datasets rarely stay fixed. They are constantly changing as more data is collected, annotated, and models are retrained. This iterative model improvement process is called Data Loop, illustrated in the image below.‍ ![](https://cdn.sanity.io/images/h6toihm1/production/e501b582019f6dcba39acf72759a4095a6e56969-1400x624.jpg?auto=format&dpr=2&fit=max&q=75&w=1400) ‍It is generally established that the more high-quality data you feed into the model, the better performance it achieves. The estimations are (eg. [\[1\]](https://towardsdatascience.com/predicting-the-performance-of-deep-learning-models-9cb50cf0b62a), [\[2\]](https://arxiv.org/abs/1712.00409)) that to reduce the training error by half, you need four times more data. But there’s a tradeoff: the more data you use in the training, the more time is needed for the training itself, as well for the annotation. And, unlike model training, because the annotation process is largely human-led, it can’t be simply sped up by more performant hardware.‍ That’s why it is important to keep datasets just the right size to be able to annotate data quickly and with high quality. The smaller the dataset, the better the annotation quality required to achieve good training results. Annotations must not contradict each other and be accurate. Since the annotations are done by people, they require validation. And that’s where tools, like FiftyOne and CVAT, can greatly help.‍ [FiftyOne](https://fiftyone.ai/) is an open-source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster.‍ [CVAT](https://www.cvat.ai/) is one of the leading open-source solutions for annotating Computer Vision datasets. It allows you to create quality annotations for images, videos and 3D point clouds and prepare ready-to-use datasets. It has an [online platform](https://app.cvat.ai/) and [can be deployed](https://github.com/opencv/cvat) on your computer or cluster. It is a scalable solution both for personal use and for big teams.‍ In this blog post we will demonstrate how you can use these tools to create high-quality annotations for a dataset, validate the annotations, and detect and fix problems. ‍Follow along with the code in this post through [this Colab notebook](https://colab.research.google.com/drive/1WsAwB5nC3YeTMCGBXNugxQShP613x3wm?usp=sharing).‍ ## Dataset Curation To demonstrate a data-centric ML workflow, we will create an object detection dataset from raw image data. We will use images from the [MS COCO dataset](https://cocodataset.org/#home). This dataset is available in the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html). You can easily download custom subsets of the dataset and load them into FiftyOne. This dataset does have object-level annotations, but we will avoid using them in order to show how to annotate a dataset from scratch.‍ import fiftyone as fo import fiftyone.zoo as foz dataset = foz.load\_zoo\_dataset( "coco-2017", split="validation", label\_types=\[\], ) \# Visualize the dataset in FiftyOne session = fo.launch\_app(dataset) While it is easy to load a large dataset into [FiftyOne for visualization](https://voxel51.com/docs/fiftyone/user_guide/app.html), it is much harder (and often a waste of time and money) to annotate an entire dataset. A better approach is to find a subset of data that would be valuable to annotate and start with that, adding samples as needed. A subset of a dataset can be useful if it contains an even distribution of visually unique samples, maximizing the informational content per image. For example, if you are training a dog detector, it would be better to use images from a range of different dog breeds rather than only using images of a single breed.‍ With FiftyOne, we can use the [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html#visual-similarity) to find a subset of the unlabeled dataset with visually unique images.‍ import fiftyone.brain as fob \# Generate embeddings model = foz.load\_zoo\_model("clip-vit-base32-torch") embeddings = dataset.compute\_embeddings(model) results = fob.compute\_similarity( dataset, embeddings=embeddings, brain\_key="image\_sim" ) results.find\_unique(500) unique\_subset = dataset.select(results.unique\_ids) session.view = unique\_subset Then we visualize these unique samples in the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html).‍ ![](https://cdn.sanity.io/images/h6toihm1/production/1fa878e390c2f5e944baff11a1c1d4bcb45f9357-1400x734.png?auto=format&dpr=2&fit=max&q=75&w=1400) Note: You can use the FiftyOne Brain [compute\_visualization()](https://voxel51.com/docs/fiftyone/user_guide/brain.html#visualizing-embeddings) method to visualize an interactive plot of your embeddings to identify other patterns in your dataset.‍ These visually unique images give us a diverse subset of samples for training, while at the same time reducing the amount of annotation that needs to be performed. As a result, this procedure can significantly lower annotation costs. Of course, there is a lower bound to the number of samples needed to sufficiently train your model, so you will want to iteratively add more unique samples to your dataset as needed.‍ ## Dataset Annotation Now that we’ve decided on the subset of samples in the dataset that we want to annotate, it’s time to add some labels. We will be using CVAT, one of the leading open-source annotation tools, to create annotations on these samples. [CVAT and FiftyOne have a tight integration](https://voxel51.com/docs/fiftyone/integrations/cvat.html), allowing us to take the subset of unique samples in FiftyOne and load them into CVAT in just one Python command.‍ results = unique\_view.annotate( "annotation\_key", label\_type="detections", label\_field="ground\_truth", classes=\["airplane", "apple", …\], backend="cvat", launch\_editor=True, ) Since this annotation process can take some time, we will want to make sure that our dataset is [persisted in FiftyOne](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#dataset-persistence) so that we can load it again at some point in the future when the annotation process is complete.‍ dataset.persistent = True \# Optionally give it a custom name dataset.name = "cvat-fiftyone-demo" \# In the future in a new Python process dataset = fo.load\_dataset("cvat-fiftyone-demo") After our data is uploaded into CVAT, a web browser page should be opened. We will see the main annotation window, where we can create, modify, and delete annotations. There are different tools available on this window toolbar, so we can draw polygons, rectangles, masks, and several other figures. In the object detection task, the primary annotation type is the bounding box. Let’s draw one using the corresponding tool from the toolbar. ![](https://cdn.sanity.io/images/h6toihm1/production/896cf6d1a00466cd9403a4af77cef2bddc75fe1e-1400x813.png?auto=format&dpr=2&fit=max&q=75&w=1400) Now, we can set the label and other attributes for the created rectangle. In this case, we used the “cat” label. After we’ve finished annotating this object, we can continue annotating other objects and images the same way. After all objects are annotated, we save work by clicking the Save button. Then, we can click the Menu button above and the Open the task button in the menu to open the task overview page.‍ In CVAT, the data is organized into Projects, Tasks, and Jobs. Each Project represents a dataset with multiple subsets (or splits) and can have one or many Tasks. You can manage tasks inside a project, join them into subsets, and export and import the data. A Task represents an annotation assignment for a person or several people. The Task can be treated as a dataset, but its primary role is to organize and split the big workload into smaller chunks. Each Task is divided into Jobs to be annotated.‍ CVAT supports different scenarios. In typical scenarios the datasets are big — from hundreds to millions of images. Datasets like these are annotated in teams divided into squads with different assignments: annotating, and reviewing of the annotated data. In CVAT, we can do both these assignments. We can assign people to jobs using the Assignee field. If there is a person to review our work and we want the annotations to be reviewed, we need to change the job Stage to “validation”: ![](https://cdn.sanity.io/images/h6toihm1/production/783716c7099b8ac4ba2aebdc44e5eeac0451f747-1400x813.png?auto=format&dpr=2&fit=max&q=75&w=1400) ‍Now, the reviewer can open the job and comment on the problems found. The user interface now will allow us to create Issues. The issues are just comments in the free form, though CVAT provides several options to mark common problems with annotation with just a single click. ![](https://cdn.sanity.io/images/h6toihm1/production/53807dcee1c725aa6fdd4d5d624c797568cd5210-1400x813.png?auto=format&dpr=2&fit=max&q=75&w=1400) Once the review is finished, we click the Save button, and return back to the task page. If everything is annotated correctly, we can mark the job as accepted and move onto other tasks. If there are problems found during the review, we can switch the job back to the annotation stage and assign it back to the annotator again.‍ ![](https://cdn.sanity.io/images/h6toihm1/production/81e39f5252e8e0288a33304cdac4896c0ad7a030-1400x813.png?auto=format&dpr=2&fit=max&q=75&w=1400)![](https://cdn.sanity.io/images/h6toihm1/production/81a3c38bd30751dfa02386fd41d2c146b449f690-1400x813.png?auto=format&dpr=2&fit=max&q=75&w=1400) Now, the annotator will be able to fix the problems and leave comments on the issues. This process can take several turns before the dataset is annotated correctly.‍ When we finish annotating this batch of samples we can again use the [CVAT and FiftyOne integration](https://voxel51.com/docs/fiftyone/integrations/cvat.html#loading-annotations) to easily load the updated annotations back into FiftyOne.‍ unique\_view.load\_annotations("annotation\_key") ## Dataset Improvement With the annotations loaded into our FiftyOne dataset, we can make use of the [powerful querying and evaluation capabilities](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) that FiftyOne provides. You can use them in the FiftyOne Python SDK and the FiftyOne App to analyze the quality of the annotations and the dataset as a whole. Dataset quality is a fairly vague concept that can depend on several factors, such as the accuracy of labels, spatial tightness of bounding boxes, class hierarchy in the annotation schema, “difficulty” of samples, inclusion of edge cases, and more. However, with FiftyOne, you can easily analyze any number of different measures of “dataset quality”.‍ For example, in object detection datasets, having the same object annotated multiple times with duplicate bounding boxes is detrimental to model performance. We can use FiftyOne to [automatically find potential duplicate bounding boxes](https://voxel51.com/docs/fiftyone/recipes/remove_duplicate_annos.html) based on the IoU overlap between them, and then visually analyze if it actually is a duplicate or if it is just two closely overlapping objects.‍ import fiftyone.utils.iou as foui from fiftyone import ViewField as F foui.compute\_max\_ious( dataset, "ground\_truth", iou\_attr="max\_iou", classwise=True, ) dups\_view = dataset.filter\_labels( "ground\_truth", F("max\_iou") > 0.75 ) session.view = dups\_view We can then [tag these samples](https://voxel51.com/docs/fiftyone/user_guide/app.html#tags-and-tagging) in the FiftyOne App as needing reannotation in CVAT. ![](https://cdn.sanity.io/images/h6toihm1/production/71df15d4a3091b2e99be30b024356c4654468480-1400x721.png?auto=format&dpr=2&fit=max&q=75&w=1400) Note: Other workflows FiftyOne provides to assess your dataset quality include methods to [evaluate the performance of you model](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html), ways to [analyze embeddings](https://voxel51.com/docs/fiftyone/tutorials/image_embeddings.html), a [measure of the likelihood of annotation mistakes](https://voxel51.com/docs/fiftyone/user_guide/brain.html#label-mistakes), and [more](https://voxel51.com/docs/fiftyone/user_guide/brain.html).‍ Using the FiftyOne and CVAT integration, we can send only the tagged samples over to CVAT and reannotate them.‍ reannotate\_view = dataset.match\_tags("needs\_reannotation") results = reannotate\_view.annotate( "reannotation", label\_field="ground\_truth", backend="cvat", ) ![](https://cdn.sanity.io/images/h6toihm1/production/77dc29c7738cda83e49698b6fece71ee229940fa-702x516.png?auto=format&dpr=2&fit=max&q=75&w=702)![](https://cdn.sanity.io/images/h6toihm1/production/710529ef530e67cdcc6480b286a8d09d705d33ce-800x500.png?auto=format&dpr=2&fit=max&q=75&w=800) We can then load these annotations back into FiftyOne from CVAT with more confidence in the quality of our dataset. We can also [export the created dataset](https://opencv.github.io/cvat/docs/manual/advanced/export-import-datasets/) into any of [the common formats](https://opencv.github.io/cvat/docs/manual/advanced/formats/), including MS COCO, PASCAL VOC, and ImageNet, to be used in a model training framework directly from CVAT:‍ ## Next Steps Now that we have an annotated dataset of sufficiently high quality, the next step is to start training a model. [There](https://voxel51.com/docs/fiftyone/tutorials/detectron2.html) [are](https://voxel51.com/docs/fiftyone/integrations/lightning_flash.html) [many](https://towardsdatascience.com/stop-wasting-time-with-pytorch-datasets-17cac2c22fa8) [ways](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html#tfobjectdetectiondataset) you can train models by integrating FiftyOne datasets into your existing model training workflows or [using CVAT to create a dataset](https://opencv.github.io/cvat/docs/manual/advanced/export-import-datasets/) ready for use.‍ However, the process doesn’t stop after the model is trained. This is just the beginning. As you evaluate your model performance, you will find failure modes of the model that can indicate a need for further annotation improvements or for additional data to add to your datasets to cover a wider range of scenarios.‍ This process of dataset curation, annotation, training, and dataset improvement is the heart of data-centric AI and is a continuous cycle that will lead to improved model performance. Additionally, this process is necessary for any production models to prevent them from becoming out of date as the data distribution shifts over time‍ ## Summary In the current age of AI, and especially in the computer vision domain, data is king. Following a data-centric mindset and focusing on improving the quality of datasets is the most surefire way to improve the performance of your models. To that end, there are several open-source tools that have been built with data-centric AI in mind. [FiftyOne](https://fiftyone.ai/) and [CVAT](https://www.cvat.ai/) are two leading open-source tools in this space. On top of that, they are tightly integrated, allowing you to explore, visualize, and understand your datasets and their shortcomings, as well as to take action and efficiently annotate and improve your labels to start building better models.‍ _Originally published at [https://www.cvat.ai](https://www.cvat.ai/post/data-centric)._ [CVAT](https://voxel51.com/blog/tag/cvat) [data annotation](https://voxel51.com/blog/tag/data-annotation) [data-centric computer vision](https://voxel51.com/blog/tag/data-centric-computer-vision) [dataset curation](https://voxel51.com/blog/tag/dataset-curation) [dataset improvement](https://voxel51.com/blog/tag/dataset-improvement) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [integrations](https://voxel51.com/blog/tag/integrations) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ed751b5c0fbc3d2cb74f0b30e6418d3319564c-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — December 2022\\ \\ Event Recaps\\ \\ • \\ \\ Dec 13, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/4ac1a727dc192a21563cde51b6e345f620e09376-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0d2bbf0f002969f6fab0099d9e60be2fff2a256e-1400x887.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How to Curate, Annotate, and Improve Computer Vision Datasets with FiftyOne and Labelbox\\ \\ Tutorials\\ \\ • \\ \\ Jan 14, 2022](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-199-lllmstxt|> ## FiftyOne Aggregation Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Aggregation Tips and Tricks — Nov 25, 2022 Nov 26, 2022 • 5 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#d406b8abcaa1) [An Aggregations Primer](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#71011eebe767) [One Aggregation; Multiple Datasets](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#4a0c2bb45e82) [Multiple Aggregations](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#ea8fde6c3b21) [Unwinding Lists of Lists](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#4ac0ba2fe6ec) [Aggregations on Transformed Field Values](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#14ed505df624) [Going Beyond Aggregations](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#45c29c4be12b) [What’s next?](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#9c407093321d) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#d406b8abcaa1) [An Aggregations Primer](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#71011eebe767) [One Aggregation; Multiple Datasets](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#4a0c2bb45e82) [Multiple Aggregations](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#ea8fde6c3b21) [Unwinding Lists of Lists](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#4ac0ba2fe6ec) [Aggregations on Transformed Field Values](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#14ed505df624) [Going Beyond Aggregations](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#45c29c4be12b) [What’s next?](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022#9c407093321d) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/f0fe111e79e719eeb9100d13801921e34a778043-1200x673.png?auto=format&dpr=2&fit=max&q=75&w=1200) Welcome to our weekly FiftyOne tips and tricks blog where we give practical pointers for using FiftyOne on topics inspired by discussions in the open source community. In this Thanksgiving week installment, in the spirit of togetherness, we’ll cover [aggregations](https://voxel51.com/docs/fiftyone/user_guide/using_aggregations.html). ## **Wait, What’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## An Aggregations Primer [Datasets](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#using-datasets) are the core data structure in FiftyOne, allowing you to represent your raw data, labels, and associated metadata. When you query and manipulate a **Dataset** object using [dataset views](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#using-views), a DatasetView object is returned, which represents a filtered view into a subset of the underlying dataset’s contents. Complementary to this data model, one is often interested in computing aggregate statistics about datasets, such as label counts, distributions, and ranges, where each Sample is reduced to a single quantity in the aggregate results. The [`fiftyone.core.aggregations`](https://voxel51.com/docs/fiftyone/api/fiftyone.core.aggregations.html#module-fiftyone.core.aggregations) module offers a declarative and highly-efficient approach to computing summary statistics about your datasets and views. Continue reading for some tips and tricks to help you do just that. ## **One Aggregation; Multiple Datasets** If you want to compute a single aggregation on a single dataset, you can call the aggregation method directly on the dataset. For instance, to compute the [`Bounds`](https://voxel51.com/docs/fiftyone/api/fiftyone.core.aggregations.html#fiftyone.core.aggregations.Bounds) aggregation on the `uniqueness` field, you could write: dataset.bounds("uniqueness") However, if you want to use the same aggregation method (on the same field) on multiple datasets or views, you can define the aggregation on its own, and then compute that aggregation using the `aggregate()` method: mean\_agg = fo.Mean("predictions.detections.confidence") mean\_value1 = dataset1.aggregate(mean) mean\_value2 = dataset2.aggregate(mean) Learn more about the `aggregate()` method in the FiftyOne Docs. ## **Multiple Aggregations** Conversely, if you want to compute multiple aggregations on the same dataset, not necessarily on the same field, you can do so efficiently by batching them in the `aggregate` method: \# will count the number of samples in a dataset sample\_count = fo.Count() \# will retrieve the distinct labels in the \`ground\_truth\` field distinct\_labels = fo.Distinct("ground\_truth.detections.label") \# will compute a histogram of the \`uniqueness\` field histogram\_values = fo.HistogramValues("uniqueness") \# efficiently compute all three results aggs = \[sample\_count, distinct\_labels, histogram\_values\] count, labels, hist = dataset.aggregate(aggs) Learn more about [batching aggregations](https://voxel51.com/docs/fiftyone/user_guide/using_aggregations.html#batching-aggregations) in the FiftyOne Docs. ## **Unwinding Lists of Lists** If we want to compute aggregations that are not built into FiftyOne, such as the median, we can do so by first extracting the values for the relevant field and then applying our aggregation to this result. Due to the unstructured nature of computer vision data, where different samples may contain different numbers of detected objects, this result may be a list of lists of differing sizes. For instance, if we get the prediction confidence values, pred\_conf\_field = "predictions.detections.confidence" pred\_confs\_jagged = dataset.values(pred\_conf\_field) we can see that the first ten sublists all have different lengths by running: print(\[len(p) for p in pred\_confs\_jagged\[:10\]\]) If we wanted to compute the median from this jagged list of values, we would first need to flatten the list. However, FiftyOne does this for us when we pass the argument `unwind = True` into the `values()` method: pred\_conf\_field = "predictions.detections.confidence" pred\_confs = dataset.values(pred\_conf\_field, unwind=True) From there, we can pass the resulting flat array straight to numpy: median\_conf = np.median(pred\_confs) Learn more about unwinding and `values()` aggregation in the FiftyOne Docs. ## **Aggregations on Transformed Field Values** When we use the FiftyOne `Aggregations` class in conjunction with `ViewField`, we can easily perform aggregations over transformed field values. For instance, to compute the mean of squared prediction confidence values, we can write either of the following: \## Option 1 aggregation = fo.Mean(F("predictions.detections.confidence") \*\* 2) squared\_conf\_mean = dataset.aggregate(aggregation) \## Option 2 squared\_conf\_mean = dataset.mean(F("predictions.detections.confidence") \*\* 2)) Learn more about [expressions](https://voxel51.com/docs/fiftyone/api/fiftyone.core.expressions.html) and [`ViewField`](https://voxel51.com/docs/fiftyone/api/fiftyone.core.expressions.html#fiftyone.core.expressions.ViewField) in the FiftyOne Docs. ## **Going Beyond Aggregations** While aggregations over an entire `Dataset` or `DatasetView` are powerful, sometimes they are not sufficient to fully understand your model or your data. In fact, it is precisely this need for better transparency during data-model co-design that led to the creation of the FiftyOne App and FiftyOne Teams. Beyond dataset and view level aggregations, FiftyOne provides class-specific reports for classification and multi-class object detection techniques via the `print_report()` method. For multi-class detection tasks, this looks like: results = dataset.evaluate\_detections( "predictions", gt\_field="ground\_truth", eval\_key="eval", compute\_mAP=True, ) \# Get the 10 most common classes in the dataset counts = dataset.count\_values("ground\_truth.detections.label") classes\_top10 = sorted(counts, key=counts.get, reverse=True)\[:10\] \# Print a classification report for the top-10 classes results.print\_report(classes=classes\_top10) For a binary classification task, this looks like: results = dataset.evaluate\_classifications( "predictions", gt\_field="ground\_truth", eval\_key="eval", method="binary", classes=\["classA", "classB"\], ) results.print\_report() Learn more about `print_report()` and [evaluating detection](https://voxel51.com/docs/fiftyone/tutorials/evaluate_detections.html) and [classification](https://voxel51.com/docs/fiftyone/tutorials/evaluate_classifications.html) tasks in the FiftyOne Docs. ## **What’s next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [aggregations](https://voxel51.com/blog/tag/aggregations) [FAQ](https://voxel51.com/blog/tag/faq) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/3f54d0a45faa06a04b5d0244dd7c092603150cf0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Mar 10, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 11, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023) [FiftyOne Computer Vision Tips and Tricks – Mar 24, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 25, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-24-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4d37703d72d4b83a85bda19eb1999d5247915fc0-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 02, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 3, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-200-lllmstxt|> ## FiftyOne and Pandas Comparison [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Tutorials](https://voxel51.com/blog/category/tutorials) Why FiftyOne is the pandas of Computer Vision Nov 23, 2022 • 3 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#7689af5e5694) [How to perform pandas-style queries in FiftyOne](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#fdc9fddb5fad) [Example: getting the column or field schema](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#1079dce6e483) [Example: minimum and maximum values](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#414863795a5d) [Example: median and other aggregations](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#581a604bae2b) [Community contributions](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#f7d13b3999f4) [FiftyOne community updates](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#297c0bf44254) [What’s next?](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#d002f9b2eada) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#7689af5e5694) [How to perform pandas-style queries in FiftyOne](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#fdc9fddb5fad) [Example: getting the column or field schema](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#1079dce6e483) [Example: minimum and maximum values](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#414863795a5d) [Example: median and other aggregations](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#581a604bae2b) [Community contributions](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#f7d13b3999f4) [FiftyOne community updates](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#297c0bf44254) [What’s next?](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision#d002f9b2eada) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/cfa6067062cae206570d98a5e688951723545822-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=1200) [FiftyOne](https://github.com/voxel51/fiftyone) and [pandas](https://pandas.pydata.org/) are both open source Python libraries that make dealing with your data easy. While they serve different purposes — pandas is built for tabular data, and FiftyOne is built for unstructured data in computer vision tasks — their syntax and functionality are closely aligned. Specifically, the pandas DataFrame and FiftyOne Dataset share many similar functionalities. Because both pandas and FiftyOne are important components to many data science and machine learning workflows, we often hear the request in the FiftyOne community for a comparison of the two. The community requested it; we delivered it. Performing pandas-style queries on your computer vision data using FiftyOne has never been easier. In this blog post, its companion tutorial — [Perform pandas-style queries in FiftyOne](https://voxel51.com/docs/fiftyone/tutorials/pandas_comparison.html), and an accompanying [pandas vs FiftyOne Cheat Sheet](https://docs.voxel51.com/cheat_sheets/pandas_vs_fiftyone.html), we’ll show you how. Huge shout out to community Slack member [Kishan Savant](https://www.linkedin.com/in/kishan-savant/) for creating the cheat sheet! ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. ## **How to perform pandas-style queries in FiftyOne** - Are you a seasoned user of [pandas](https://pandas.pydata.org/), the Python library for tabular data analysis? Tired of scouring through FiftyOne’s API Reference pages in search of analogous operations? - Interested in seeing a new way to think about the FiftyOne Dataset and DatasetView? In this blog post we’ll show you how to perform some popular panda-style queries and operations in FiftyOne. To see the complete collection, check out our new pandas and FiftyOne Queries Comparison Guide! Starting with the basics, this tutorial covers everything you need to know to perform pandas-style queries and operations in FiftyOne. If you have more specific questions, you can go straight to the sections on [View Stages](https://voxel51.com/docs/fiftyone/tutorials/pandas_comparison.html#View-stages), [Aggregations](https://voxel51.com/docs/fiftyone/tutorials/pandas_comparison.html#Aggregations), [Structural Changes](https://voxel51.com/docs/fiftyone/tutorials/pandas_comparison.html#Structural-change-operations), or [Expressions](https://voxel51.com/docs/fiftyone/tutorials/pandas_comparison.html#Expressions). ## **Example**: getting the column or field schema In pandas, where all rows in a DataFrame share the same columns, we can get the names of the columns with the `columns` property. For instance, for a DataFrame `df`, we can get the columns via: df.columns In FiftyOne, the core field schema is shared among samples, but the structure within these first-level fields can vary. We can get the field schema by calling the `get_field_schema()` method: ds.get\_field\_schema() ## Example: m **inimum and maximum values** In pandas, you compute the minimum and maximum value of a Series separately. To get the minimum and maximum values in the “my\_col” column in a DataFrame `df` for instance, we can call: col\_min = df\[“my\_col”\].min() col\_max = df\[“my\_col”\].max() When working with a FiftyOne Dataset or DataView, the min and max are returned together in a tuple when the `bounds()` method is called on a field. To get the minimum and maximum in the field “my\_field” of a Dataset `ds`, we use: field\_min, field\_max = ds.bounds(“my\_field”) ## Example: m **edian and other aggregations** Some aggregations which are native to pandas, such as computing the median, are not native to FiftyOne. In these cases, the canonical way to compute the aggregation is by first extracting the values from the Dataset field, and then using native numpy or scipy functionality. To compute the median value in a column of a pandas DataFrame, we can write: col\_median = df\[“my\_col”\].median() In FiftyOne, fields can contain embedded fields or lists of values. To avoid dealing with these potentially jagged lists of values, we can pass the argument `unwind=True` into the Dataset `values()` method. We can then use this flattened list of values as input into numpy’s median method: import numpy as np pred\_confs\_flat = ds.values("predictions.detections.confidence", unwind = True) pred\_confs\_median = np.median(pred\_confs\_flat) ## Community contributions - Shoutout to [Kishan Savant](https://github.com/NeoKish), who created the [pandas vs FiftyOne Cheat Sheet](https://docs.voxel51.com/cheat_sheets/pandas_vs_fiftyone.html). ## **FiftyOne community updates** The FiftyOne community continues to grow! - 1,100+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,100+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 1,700+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 190+ repositories - 46+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## **What’s next?** - Check out the [Perform pandas-style queries in FiftyOne tutorial](https://voxel51.com/docs/fiftyone/tutorials/pandas_comparison.html). - Download the [pandas vs FiftyOne Cheat Sheet](https://docs.voxel51.com/cheat_sheets/pandas_vs_fiftyone.html). - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [pandas](https://voxel51.com/blog/tag/pandas) [pandas-style queries](https://voxel51.com/blog/tag/pandas-style-queries) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/047b21a97f6c858334f9f35ed89fa7655ebf5767-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ State-of-the-Art Object Detection with YOLO-NAS & FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ May 4, 2023](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/713e4352b25d3b4ee12eab92246ceff22f808471-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Spending My First Week With FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ Aug 21, 2023](https://voxel51.com/blog/spending-my-first-week-with-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/3ed3003518fbdb96f01122d1c652a225274d671b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Build Your Own AI Art Gallery\\ \\ Computer Vision, Plugins, Tips & Tricks, Tutorials\\ \\ • \\ \\ Aug 25, 2023](https://voxel51.com/blog/build-your-own-ai-art-gallery) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-201-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Nov 18, 2022 Nov 19, 2022 • 4 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#211e0c03e7ca) [Viewing the FiftyOne App with Interactive Plots](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#096b93d09f47) [Importing Data from CVAT](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#a2169278fb5b) [Selecting Samples with Attribute Set](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#3bfdb7e91c47) [Using FiftyOne with Images from the Internet](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#8b5fa3b424bd) [Sending Large Annotation Jobs to Label Studio](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#03858633e380) [What’s Next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#9e13ad8a353a) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#211e0c03e7ca) [Viewing the FiftyOne App with Interactive Plots](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#096b93d09f47) [Importing Data from CVAT](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#a2169278fb5b) [Selecting Samples with Attribute Set](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#3bfdb7e91c47) [Using FiftyOne with Images from the Internet](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#8b5fa3b424bd) [Sending Large Annotation Jobs to Label Studio](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#03858633e380) [What’s Next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022#9e13ad8a353a) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/47c4462ab8c14fe4728ce5901402986125d3f99d-1200x676.png?auto=format&dpr=2&fit=max&q=75&w=1200) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, What’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help Ok, let’s dive into this week’s tips and tricks! ## **Viewing the FiftyOne App with Interactive Plots** Community Slack member Patrick Rowsome asked, _“Is it possible to open the FiftyOne App alongside an interactive plot in JupyterLab?_ Yes! There are multiple ways of accomplishing this. By right-clicking the output of a cell in a JupyterLab notebook and selecting ‘Create New View for Output’, you can view the output — namely the FiftyOne GUI — in another JupyterLab tab. You can then drag the tab horizontally or vertically, as illustrated [here](https://datacomy.com/tools/jupyter/split_screen/), to create a split view. Alternatively, you can view the FiftyOne App in a separate browser window or tab. To do this, pass the option `auto = False` when you create a session: session = fo.launch\_app(…, auto=False) You can find the URL for the session by running `session.url`, which you can then copy and paste into your browser. Learn more about `InteractivePlot` and [interactive plotting in Jupyter notebooks](https://voxel51.com/docs/fiftyone/user_guide/plots.html) in the FiftyOne Docs. ## **Importing Data from CVAT** Community Slack member Joy Timmermans asked, _“I currently already have some data loaded into CVAT. is it possible to connect that directly to FiftyOne?”_ FiftyOne has a very tight integration with CVAT which makes it easy to both create a FiftyOne dataset from a list of CVAT tasks, and to [create CVAT annotation tasks directly from FiftyOne](https://voxel51.com/docs/fiftyone/integrations/cvat.html#requesting-annotations). If the media files are not downloaded to disk, then you can import annotations from CVAT using our provided `import_annotations` method, as in the example below. import fiftyone as fo import fiftyone.utils.cvat as fouc dataset = fo.Dataset("my-dataset") fouc.import\_annotations( dataset, task\_ids=\[...\], download\_media=True, ) Alternatively, if the media is downloaded to disk, you can provide the `data_path` [argument](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.cvat.html#fiftyone.utils.cvat.import_annotations) instead of `download_media`. Learn more about FiftyOne’s [CVAT integration](https://voxel51.com/docs/fiftyone/integrations/cvat.html#) in the FiftyOne Docs. ## **Selecting Samples with Attribute Set** Community Slack member Geoffrey Keating asked, _“I want to pull only the samples that have an attribute set on a Detections field. The field is a bool and will only belong to detections of a certain label. What is the best way to do this?”_ While this can also be accomplished using `filter_field`, it is best to use the `match` stage. This is the case any time you want to match samples based on a condition. The matching operation can be performed with the following syntax: view = dataset.match( F("ground\_truth.detections") .filter(F("condition") == True) .length() > 0 ) Learn more about `DatasetView` and [matching conditions](https://voxel51.com/docs/fiftyone/api/fiftyone.core.view.html#fiftyone.core.view.DatasetView.match) in the FiftyOne Docs. ## **Using FiftyOne with Images from the Internet** An anonymous user on Stack Overflow asked, _“Is it possible to use FiftyOne with images available at external URLs (in google images for example) without downloading the images first?_ This functionality is available in [FiftyOne Teams](https://voxel51.com/fiftyone-teams/), with which you can point samples directly to an `https://` image (as well as to media stored in Amazon S3, Google Cloud Storage, Azure, MinIO, etc.). If you’re using the open source FiftyOne library, then you should download the images first. Learn more about [loading a dataset from disk](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/datasets.html) in the FiftyOne Docs. ## **Sending Large Annotation Jobs to Label Studio** Community Slack member Brett Israelsen asked, _“I’m trying to send an annotation job out to Label Studio. My dataset is not currently labeled, so I’m sending the whole thing. Is there a limit to how many files can be sent in a single batch?”_ The [FiftyOne Label Studio integration](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.labelstudio.html) uploads all media files at once. If you have a large dataset (either large images, many images, or both), then it is probably best to batch the images you send for annotation, for example: for batch in range(0, len(dataset), batch\_size): view = dataset\[batch\*batch\_size : (batch+1) \* batch\_size\] results = view.annotate(..., backend="labelstudio", project\_name="my\_project") It is important that the annotation key passed in to `view.annotate()` is unique for each batch. Learn more about FiftyOne’s [Label Studio integration](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.labelstudio.html) in the FiftyOne Docs. Alternatively, [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) supports cloud-backed datasets, which allows for the creation of tasks with minimal moving around of media files. ## **What’s Next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help [CVAT](https://voxel51.com/blog/tag/cvat) [FAQ](https://voxel51.com/blog/tag/faq) [interactive plots](https://voxel51.com/blog/tag/interactive-plots) [Label Studio](https://voxel51.com/blog/tag/label-studio) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/f8c59b7ff0a53a527b7002a899ccee84e95dee0a-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 16, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 17, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/03107d477b7db4be03031293fa4fe15aaea806f0-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Sept 16, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 17, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/93682b6d528c21a04ca55782334799a0715028ea-1200x678.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 30, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 31, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-202-lllmstxt|> ## November 2022 Computer Vision Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup — November 2022 Nov 16, 2022 • 11 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#4f30567570b5) [Meetup Recap at a Glance](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#516d9c382a3d) [Talk #1: Scaling Autonomous Vehicles with End2end](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#4ba8856bbb70) [Talk #1 Video Replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#6d7cb43b35ad) [Talk #1 Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#935f07103c5d) [Talk #1 Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#863ec8dbe69a) [Talk #2: Synthetic Data Generators and Deploying Highly Accurate Retail Supply Chain Computer Vision Apps](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#74a7444126a7) [Talk #2 Video Replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#2149ee5b5206) [Talk #2 Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#f153218bad1a) [Talk #2 Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#e59dbae3c6e9) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#cb5f061578c4) [Upcoming Computer Vision Meetup Speakers & Schedule](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#12566c9c2bbe) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#be38be58aec9) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#4f30567570b5) [Meetup Recap at a Glance](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#516d9c382a3d) [Talk #1: Scaling Autonomous Vehicles with End2end](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#4ba8856bbb70) [Talk #1 Video Replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#6d7cb43b35ad) [Talk #1 Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#935f07103c5d) [Talk #1 Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#863ec8dbe69a) [Talk #2: Synthetic Data Generators and Deploying Highly Accurate Retail Supply Chain Computer Vision Apps](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#74a7444126a7) [Talk #2 Video Replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#2149ee5b5206) [Talk #2 Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#f153218bad1a) [Talk #2 Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#e59dbae3c6e9) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#cb5f061578c4) [Upcoming Computer Vision Meetup Speakers & Schedule](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#12566c9c2bbe) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022#be38be58aec9) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/bbb1d9add0b0b9aa12682acac795df7c2ba760a9-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=1200) Last week we hosted the November 2022 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/) and had a blast! The speakers were engaging, the virtual room was packed, and Q&A was lively. In this blog post we provide the recordings, as well as recap some highlights and Q&A from the presentations. We’ll also share the Meetup locations and upcoming schedule so that you can join us at a future event. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to vote for their favorite charity and help guide our monthly donation to charitable causes. The charity that received the highest number of votes was the [World Literacy Foundation](https://worldliteracyfoundation.org/). We are pleased to be making a donation of $200 to them on behalf of the computer vision community! ![](https://cdn.sanity.io/images/h6toihm1/production/6da3098b8f6eeb34e3674a5984affbc8a663de8e-300x300.png?auto=format&dpr=2&fit=max&q=75&w=300) ## Meetup Recap at a Glance - Talk #1 — Sri // Autonomous Vehicles with End2end - [Presentation recap](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366#e8a1) - [Video replay](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366#b563) - [Q&A recap](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366#41a7) - [Additional resources](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366#a862) - Talk #2 — Tarik // Retail Supply Chain & Computer Vision - [Presentation recap](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366#9692) - [Video replay](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366#5536) - [Q&A recap](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366#77d0) - [Additional resources](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366#e820) - [Computer Vision Meetup Locations](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366#ce0b) - [Computer Vision Meetup Speakers — December, January, February](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366#d77e) - [Get Involved!](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366#31cf) ## Talk \#1: Scaling Autonomous Vehicles with End2end The first talk was by [Sri Anumakonda](http://srianumakonda.com/), an autonomous vehicle developer focusing on creating Computer Vision software to help push the boundaries of self-driving cars. Computer Vision has been primarily used as a way to perform scene understanding when deploying autonomous vehicles. But recently, there has been more and more research into how we can leverage Deep Learning and neural networks to create learned mappings from raw data to control outputs. In this talk, Sri introduced a method to train self-driving cars completely on camera, through a technique known as end2end learning. He described how we can create internal mapping spaces to go from camera images to control, how we can interpret these “black-box networks”, and how we can use end2end learning to solve self driving. Additionally, he introduced some of the most promising research in the space and how we can create end2end systems that scale faster than any other method. If you’re interested in learning about these concepts and more, then Sri’s presentation is for you: - Self driving & autonomous vehicles (AVs) - Deep learning - Convolutional Neural Networks - End2end learning - Semantic segmentation - Transposed convolutions - Optical (motion) flow - Depth estimation - 3D reconstruction - Localization ## Talk \#1 Video Replay https://www.youtube.com/watch?v=8lLRZzFft5Y ## Talk \#1 Q&A Recap Here’s a recap of the live Q&A following the presentation during the virtual Computer Vision Meetup: **​​What are some examples of more modern autonomous vehicle architectures? More specifically, have there been major changes from the 2 stream 3D-CNN approach from the 2019 paper cited? Or are companies still widely using similar architectures?** Assuming this question is regarding the Wayve Urban Driving paper, Sri’s answer to that is yes and no. The approach for end-to-end learning is very straightforward. You’re given this image, you’re given this neural network that’s able to process this image, and then you have this output, which is your lateral and longitudinal control. That’s the main meat of deep learning or in this case end-to-end learning. But there’s a lot still being worked on in modern architectures. Check out a [recent blog post by Wayve](https://www.wayve.ai/blog/learning-a-world-model-and-a-driving-policy) that describes how to add other complexities to your model, such as bird’s-eye view and more, in order to improve performance. **Elon Musk predicted that AVs would be commonplace by 2021. What went wrong and what is your prediction about AVs?** Sri answers that self-driving itself is a very complex environment. Even though we’re able to create driving that works in certain environments of the world, it’s very hard to create generalizable driving if you don’t have cars that drive all over the place. That’s why companies like Tesla perform so well. Because of the fact that they have so many users that live across different places in America, they’re able to collect so much data about what’s going on in the real world. Regarding what went wrong? It’s likely something that’ll fix itself with respect to time. End-to-end learning itself is a very complex problem, especially with computer vision. The challenge in using deep neural networks and convolutional nets to go from input to output is essentially a data problem where the more and more diverse and high quality data that you have, the better your model can be and the more generalized it can be with respect to the real world. Sri states that we’ve come very far in the past 10 to 15 years, but predicts that this will take five to 10 more years to get it nailed down. We’re at the point where we have solved 99% of self-driving, but for every single 0.9 that we add to that 99.9, the challenge of self-driving becomes orders of magnitudes harder than what it was before. **Have you experimented with the ViT vision transformer in your work?** Sri has not so far, but notes that it is very interesting. [ViT transformers](https://en.wikipedia.org/wiki/Vision_transformer) and self-driving have had a really big boost over just the past 12 to 18 months, seeing how we can apply this idea that was originally made in NLP into self-driving. But Sri finds it really fascinating to think about how we can incorporate better scene understanding through transformers. So the answer to that is no, but it is a very interesting space to think about. **For optical flow, if the cameras are on the driving car, wouldn’t the background move faster than the other cars that are driving along with the self-driving car?** Sri answers no, mainly because of the fact that relative to the foreground, your background is more static, meaning that the change in the pixel values of tree movement in the background for example is much less compared to the change in the pixel values of other cars that are in the foreground. Especially if the car is in the opposite lane, then these cars move much faster than these trees in the background are moving. So therefore you can use that information and leverage it so that you can do better prediction. **How can I set up a similar environment to the one you’ve discussed?** Sri replies that there are a lot of really good simulators online for self driving, including a very popular one known as [CARLA](https://carla.org/), which is a really good simulator that has everything you need for self-driving. It has LiDAR values, semantic segmentation, instance segmentation, and more, in addition to being very fun to play around with. **Humans mainly rely on vision for driving, but sound is also important. For example in case of an ambulance coming, you may not see it, but you will hear it. How important a role do you think sound will play in autonomous vehicles and can this help computer vision in any way?** Sri notes that he’s been thinking of this as well because sound is very important, especially in the cases of needing to hear emergency vehicles. Sri has a hypothesis on how to add this to a computer vision model that involves having a separate network that does sound detection. For example, if you hear a police car, an ambulance, or fire truck, then maybe you apply a [one-hot](https://en.wikipedia.org/wiki/One-hot) encoding vector where you declare if sound == fire truck, then execute a command that has the car move to the right and start driving really slowly to increase caution. **Do end-to-end models with multiple sensor modalities perform better or worse than just implementing sensor fusion separately from the rest of the decision making?** Sri answers that there are two ways you could think about this. The first way that you could think about it is in sensor fusion the way that it works end-to-end is that you have it all running through one bigger network. This not only allows you to have a better representation or understanding of what’s going on in your scene, but the more important factor to think about is that now it’s not a human that’s determining what type of patterns you are looking for. So one of the problems with sensor fusion being implemented separately from the rest of the actual decision making is that humans tend to implement separate modalities into the actual sensor fusion itself. But with end-to-end learning, you can have the network create its own internal mapping so that it can pick up on patterns that humans are not able to pick up on, which allows it to not only have a better understanding of its world itself, but also have better performance. **How do models perform in different lighting conditions?** Sri replies yes, the lighting and weather are important to computer vision and self-driving particularly. Driving at night is much more complex than driving during the day because objects are much more clear in the day versus in the night and in adverse conditions. ## Talk \#1 Additional Resources Check out these additional resources on the presentation and the speaker: - [Presentation slides](https://docs.google.com/presentation/d/16RSGDTHsKWBKlC98lOtbZ2JUaL3WzH4sUWHYTXhDfeE/edit#slide=id.g18eeff7e1aa_0_819) - [Talk Transcript](https://www.rev.com/transcript-editor/shared/OHBxR2z7GKlxI59KB-uz9UDK8gwXHyWfW6YO8UuOVgx4tVO5OUgxB4TrRjQTPRdR3BGs2coIOHK3_p1ffTmWqWpt7E8?loadFrom=SharedLink) - [Sri’s website](http://srianumakonda.com/) - [Sri’s GitHub](https://github.com/srianumakonda) Thank you so much to Sri on behalf of the entire Computer Vision Meetup community for sharing your knowledge and inspiring us! ## Talk \#2: Synthetic Data Generators and Deploying Highly Accurate Retail Supply Chain Computer Vision Apps The second talk in the Computer Vision Meetup was by Tarik Hammadou, a Senior Developer Relations Manager at NVIDIA. Training data for product recognition within a large retail supply chain context is hard to get to scale as its dynamic in nature, with new products being introduced frequently. Supervised learning models rely on the training data and this problem becomes significant with the scale of the machine learning models. In this talk, Tarik presented a method based on creating a digital twin of the fulfillment or a distribution center facility and generating photorealistic digital assets to train and optimize the classification model to be deployed in the real world. The performance of the training process is then used in a feedback loop to adjust the synthetic data generator until an acceptable result is achieved. Furthermore, he shared his deployment orchestration methodology over a large number of compute nodes. This method can also be extended to product inspection and other more complex computer vision tasks. If you have scenario — such as material handling optimization — where you think trying it out in a “digital twin” to optimize performance would be a preferred step before attempting to implement it in your physical space (think depalletization, conveyor belts, picking and sorting stations), then this talk is for you! ## Talk \#2 Video Replay https://www.youtube.com/watch?v=aRDEi0Q\_hNo ## Talk \#2 Q&A Recap Here’s a recap of the live Q&A following the presentation during the virtual Computer Vision Meetup: **Did you combine real and synthetic data for training?** Tarik shared that in this specific example, they trained the classifier using a pre-trained model on synthetic data. **Which pretrained model did you use?** Tarik answered that they used Yolo v5. They have also done some instance segmentation with Mask R-CNN and Faster R-CNN as well. **How significant was the improvement from the automatic tuning of the dataset generation parameters?** To this question Tarik replied that regarding the use of Replicator and the synthetic data generator, as you are training your neural network in a feedback loop, you are tuning the parameters of your synthetic data generator, and you can get significant results — very close to state-of-the-art results — in terms of performance. **When using synthetic data for testing, is the feedback to the dataset generation parameters manual or automatic?** Tarik shared that when he performed these experiments, the feedback loop was manual but they are in the process of automating it. Tarik is working with a company, [Kinetic Vision](https://kinetic-vision.com/), and in their workflow there is an automatic feedback loop — when you are training a network, you use the accuracy of your result to go back into a feedback loop and change those parameters. **Do you use any differentiable rendering techniques inside Omniverse?** Tarik answered that, yes, there are two different types of methods and techniques in terms of rendering inside Omniverse. Stay tuned for some additional resources discussing this topic. **You said it took 4 months from the start of the project to its deployment. How many people were working on the project?** Tarik explains that the biggest challenge he has seen in the AI world is that POCs can be painful. It takes sometimes six to eight months just to conduct a POC. And in this four month project, they had six people working on it. It gave them an indication that if they have pre-trained models, and a workflow and pipeline that are well established, they can accelerate the enablement and the deployment of those applications. **What was your ratio of synthetic data to real image to train the advanced computer vision models or does the ratio matter?** Tarik shared that in this case, they trained those models only on synthetic data. The pretrained model was trained on real data from scratch and then it was fine tuned only with the synthetic data. ## Talk \#2 Additional Resources You can find the [talk transcript here](https://www.rev.com/transcript-editor/shared/Y802B6ad-40v24mrViiVhpsAD0k6LJhsPU9ghNZHdYc2y5-FifYpMd8XjtQWRi-zsTRhnx5Wk0Sn_UzvQdTSFXdMBnc?loadFrom=SharedLink). Thank you Tarik on behalf of the entire Computer Vision Meetup community for sharing your knowledge and the retail supply chain use case with us! ## Computer Vision Meetup Locations Computer Vision Meetup membership has grown to more than 1,600+ members in just a few months! The goal of the meetups is to bring together a community of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. If that’s you, we invite you to join the Meetup closest to your timezone: - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) ## Upcoming Computer Vision Meetup Speakers & Schedule We recently announced an exciting lineup of speakers for December, January, and February. Become a member of the Meetup closest to you, then register for the Zoom for the Meetups of your choice. ### December 8 - The Future of Data Annotation: Trends, Problems & Solutions [— Anna Petrovicheva](https://www.linkedin.com/in/anna-petrovicheva-44b24673/) (CVAT.AI) - Using Similarity Learning to Improve Data Quality — [Kacper Łukawski](https://www.linkedin.com/in/kacperlukawski/) (Qdrant) - [Zoom Link](https://us02web.zoom.us/webinar/register/4016674056664/WN_8u6oFjU7QN2StaQSWTw7HQ) ### January 12 - Hyperparameter Scheduling for Computer Vision — [Cameron Wolfe](https://www.linkedin.com/in/cameron-r-wolfe-04744a238/) (Alegion/Rice University) - Talk abstract coming! — [Julien Simon](https://www.linkedin.com/in/juliensimon/) (Hugging Face) - [Zoom Link](https://us02web.zoom.us/webinar/register/8016674059298/WN_XQIZlMP2RQuCFBwBylyRqA) ### February 9 - Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO — [Paula Ramos, PhD](https://www.linkedin.com/in/paula-ramos-41097319/) (Intel) - Understanding Speech Recognition with OpenAI’s Whisper Model — [Vishal Rajput](https://www.linkedin.com/in/vishal-rajput-999164122/) (AI-Vision Engineer) - [Zoom Link](https://us02web.zoom.us/webinar/register/3416674061154/WN_P8UHtAZGQWOx_A2HM1dcPA) ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if any of these describe you: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [autonomous vehicles](https://voxel51.com/blog/tag/autonomous-vehicles) [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [end2end learning](https://voxel51.com/blog/tag/end2end-learning) [retail use case](https://voxel51.com/blog/tag/retail-use-case) [supply chain use case](https://voxel51.com/blog/tag/supply-chain-use-case) [synthetic data](https://voxel51.com/blog/tag/synthetic-data) [synthetic data generator](https://voxel51.com/blog/tag/synthetic-data-generator) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ed751b5c0fbc3d2cb74f0b30e6418d3319564c-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — December 2022\\ \\ Event Recaps\\ \\ • \\ \\ Dec 13, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/09d19530030fd3df3cb247fdc77b8354b768f397-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup – January 2023\\ \\ Event Recaps\\ \\ • \\ \\ Jan 18, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/e56436d38d294978c25356ba4c5482b53a29afb3-960x540.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup – February 2023\\ \\ Event Recaps\\ \\ • \\ \\ Feb 14, 2023](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-203-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Nov 11, 2022 Nov 12, 2022 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#c745b9307f5f) [Filtering labels with ViewField](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#d3641e76b10b) [Filtering file paths for existing substrings](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#68acbd9d0b69) [Filtering labels based on detection IDs](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#cd904547e237) [Mistakenness probability and IoU default values](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#0d7ad93bd4f3) [Specifying colors for classes](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#92424d873b4f) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#cce37d0576f8) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#c745b9307f5f) [Filtering labels with ViewField](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#d3641e76b10b) [Filtering file paths for existing substrings](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#68acbd9d0b69) [Filtering labels based on detection IDs](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#cd904547e237) [Mistakenness probability and IoU default values](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#0d7ad93bd4f3) [Specifying colors for classes](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#92424d873b4f) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022#cce37d0576f8) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/95b57b1e85f9838a5ad0f07ee3b457331d0dd178-1200x674.png?auto=format&dpr=2&fit=max&q=75&w=1200) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help Ok, let’s dive into this week’s tips and tricks! ## **Filtering labels with ViewField** Community Slack member Geoffrey Keating asked, _“I have a function that takes a bounding box and margin of error to determine if the box is on the border of an image; could I use this in conjunction with `ViewField` to filter labels?”_ First, a little background on `ViewField`. When you create a `ViewField` using a string field like `ViewField(“$embedded.field.name”)`, the meaning of this field is interpreted relative to the context in which the `ViewField` object is used. For example, when passed to the `ViewExpression.map()` method, this object will refer to the `embedded.field.name` object of the array element being processed. In other cases, you may wish to create a `ViewField` that always refers to the root document. You can do this by prepending “ `$`” to the name of the field, as in `ViewField(“$embedded.field.name”)`. Here are two options that could work. The first one uses relative coordinates: import fiftyone as fo import fiftyone.zoo as foz from fiftyone import ViewField as F def is\_bordering\_box(margin=0.05): bbox = F("bounding\_box") margins = \[\ \ bbox\[0\],\ \ bbox\[1\],\ \ 1 - bbox\[0\] - bbox\[2\],\ \ 1 - bbox\[1\] - bbox\[3\],\ \ \] return F.any(\[m < margin for m in margins\]) dataset = foz.load\_zoo\_dataset("quickstart") view = dataset.select\_fields("ground\_truth") border\_boxes = view.filter\_labels("ground\_truth", is\_bordering\_box(margin=0.01)) session = fo.launch\_app(border\_boxes) And here’s one that works in pixels: import fiftyone as fo import fiftyone.zoo as foz from fiftyone import ViewField as F def is\_bordering\_box(margin=10): bbox = F("bounding\_box") margins = \[\ \ F("$metadata.width") \* bbox\[0\],\ \ F("$metadata.height") \* bbox\[1\],\ \ F("$metadata.width") \* (1 - bbox\[0\] - bbox\[2\]),\ \ F("$metadata.height") \* (1 - bbox\[1\] - bbox\[3\]),\ \ \] return F.any(\[m < margin for m in margins\]) dataset = foz.load\_zoo\_dataset("quickstart") dataset.compute\_metadata() view = dataset.select\_fields("ground\_truth") border\_boxes = view.filter\_labels("ground\_truth", is\_bordering\_box(margin=5)) session = fo.launch\_app(border\_boxes) Learn more about using [ViewFields and expressions](https://voxel51.com/docs/fiftyone/api/fiftyone.core.expressions.html?highlight=viewexpression#fiftyone.core.expressions.ViewExpression) (with examples) in the FiftyOne Docs. ## **Filtering file paths for existing substrings** Community Slack member Adrian Loy asked and answered his own question! _“Is it possible to filter file paths for existing substrings?”_ Yes! Use `contains_str ` which determines whether the expression, which must resolve to a string, contains the given string or string(s). Learn more about [`contains_str`](https://voxel51.com/docs/fiftyone/api/fiftyone.core.expressions.html?highlight=contains_str#fiftyone.core.expressions.ViewField.contains_str) in the FiftyOne Docs. ## **Filtering labels based on detection IDs** Community Slack member Guillaume Dumont asked, _“Is it possible to filter labels based on detection IDs?”_ Yes! Use `select_labels()`: view = dataset.select\_labels(fields="ground\_truth", ids=\["list", "of", "id", "strings"\]) ## **Mistakenness probability and IoU default values** Community Slack member Laura Lin asked, _“For `fiftyone.brain.compute_mistakenness`, how are missing objects calculated? Is there a certain probability threshold that a prediction has to reach? Also, is there a certain IoU or IoA threshold that a detection and prediction bounding needs to meet before it is marked as missing/spurious?”_ The confidence threshold for predictions to be marked as missing is currently hard coded at 0.95 and IoU at 0.5. In the future, it may make sense to expose these as parameters. Learn more about [computing mistakenness](https://voxel51.com/docs/fiftyone/api/fiftyone.brain.html#fiftyone.brain.compute_mistakenness) in the FiftyOne Docs. # **Specifying colors for classes** Community Slack member Benjamin Fenker asked, _“I’d like to annotate a bounding box dataset and use the same colors for each class every time. So, dogs are blue, cats are red, etc. Can someone point me to how to set this up in the configs?”_ At the moment, you can only provide a color pool to the App, from which colors are randomly pulled. However, this is a popular request! You can track this feature’s progress [here.](https://github.com/voxel51/fiftyone/issues/1763) If you are using our `draw_labels() ` functionality to render images to disk with labels drawn on them, then you could iteratively draw one label class at a time with a set color: import eta.core.annotations as etaa colors = \["0FFFFF", "FFFFFF"\] colormap\_config = etaa.ColormapConfig( { "type": "eta.core.annotations.ManualColormap", "config": etaa.ManualColormapConfig({"colors": colors}) } ) config = foua.DrawConfig( { "per\_object\_label\_colors": False, "colormap\_config": colormap\_config, } ) ## **What’s next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help [FAQ](https://voxel51.com/blog/tag/faq) [filtering](https://voxel51.com/blog/tag/filtering) [IoU](https://voxel51.com/blog/tag/iou) [mistakenness](https://voxel51.com/blog/tag/mistakenness) [ViewField](https://voxel51.com/blog/tag/viewfield) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/204bd847400b1534c494dae594f0797457bc4c90-1620x906.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Filtering Tips and Tricks — Dec 09, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 10, 2022](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/25a70e1c9c2cae7f9df18696e0e087dddf67707d-1200x679.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Sept 23, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 24, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/93682b6d528c21a04ca55782334799a0715028ea-1200x678.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 30, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 31, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-204-lllmstxt|> ## Computer Vision Meetup Update [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Computer Vision Meetup Update — November ‘22 Nov 9, 2022 • 2 min read Article content In this article [Membership is growing fast!](https://voxel51.com/blog/computer-vision-meetup-update-november-22#9e268dfee809) [RSVP for Upcoming Meetups](https://voxel51.com/blog/computer-vision-meetup-update-november-22#b48b2e66408f) [Get Involved!](https://voxel51.com/blog/computer-vision-meetup-update-november-22#6b86d1543dea) [Consider speaking at a future Meetup](https://voxel51.com/blog/computer-vision-meetup-update-november-22#f9d4d0bd7b50) [A quick word from our sponsor](https://voxel51.com/blog/computer-vision-meetup-update-november-22#f79b3f352744) In this article [Membership is growing fast!](https://voxel51.com/blog/computer-vision-meetup-update-november-22#9e268dfee809) [RSVP for Upcoming Meetups](https://voxel51.com/blog/computer-vision-meetup-update-november-22#b48b2e66408f) [Get Involved!](https://voxel51.com/blog/computer-vision-meetup-update-november-22#6b86d1543dea) [Consider speaking at a future Meetup](https://voxel51.com/blog/computer-vision-meetup-update-november-22#f9d4d0bd7b50) [A quick word from our sponsor](https://voxel51.com/blog/computer-vision-meetup-update-november-22#f79b3f352744) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/0d65f67f35ed7c2afbfc5e6978967e492b357c43-1250x705.png?auto=format&dpr=2&fit=max&q=75&w=1250) At the end of September [we announced](https://medium.com/voxel51/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51-9bb5235e0dd2) a virtual network of [12 Meetups focused on computer vision](https://www.meetup.com/pro/computer-vision-meetups/). It’s been a month, so we thought it was time to provide an update on how things are progressing. ## **Membership is growing fast!** The membership of the Meetup has grown to over 1,600+ members in just three short months! If you still haven’t joined, here’s the current list of Meetups: - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) ![](https://cdn.sanity.io/images/h6toihm1/production/6a2b391258c94d51b612b5ae8b8f0fc15c431f57-1200x687.png?auto=format&dpr=2&fit=max&q=75&w=1200) ## **RSVP for Upcoming Meetups** We’ve got an exciting lineup of speakers between Nov ’22 and Feb ’23. Make sure to register for the Zoom after RSVP-ing for the Meetup. **November 10** - _Scaling Autonomous Vehicles with Computer Vision_ — [Sri Anumakonda](https://www.linkedin.com/in/srianumakonda/) (Autonomous Vehicle Developer) - _Synthetic Data Generators and Deploying Highly Accurate Retail Supply Chain Computer Vision Apps_ — [Tarik Hammadou](https://www.linkedin.com/in/tarikhammadou/) (NVIDIA) - [Zoom Link](https://us02web.zoom.us/webinar/register/7216674056063/WN_HcwGoOtQSQSqRqParBtQXg) **December 8** - _The Future of Data Annotation: Trends, Problems & Solutions_ [— Anna Petrovicheva](https://www.linkedin.com/in/anna-petrovicheva-44b24673/) (CVAT.AI) - _Using Similarity Learning to Improve Data Quality_ — [Kacper Łukawski](https://www.linkedin.com/in/kacperlukawski/) (Qdrant) - [Zoom Link](https://us02web.zoom.us/webinar/register/4016674056664/WN_8u6oFjU7QN2StaQSWTw7HQ) **January 12** - _Hyperparameter Scheduling for Computer Vision_ — [Cameron Wolfe](https://www.linkedin.com/in/cameron-r-wolfe-04744a238/) (Alegion/Rice University) - _Talk abstract coming!_ — [Julien Simon](https://www.linkedin.com/in/juliensimon/) (Hugging Face) - [Zoom Link](https://us02web.zoom.us/webinar/register/8016674059298/WN_XQIZlMP2RQuCFBwBylyRqA) **February 9** - _Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO_ — [Paula Ramos, PhD](https://www.linkedin.com/in/paula-ramos-41097319/) (Intel) - _Understanding Speech Recognition with OpenAI’s Whisper Model_ — [Vishal Rajput](https://www.linkedin.com/in/vishal-rajput-999164122/) (AI-Vision Engineer) - [Zoom Link](https://us02web.zoom.us/webinar/register/3416674061154/WN_P8UHtAZGQWOx_A2HM1dcPA) ## **Get Involved!** This is a community-driven group of Meetups, so please consider getting involved to help make these events awesome! Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. ## **Consider speaking at a future Meetup** Are you working on an interesting computer vision problem at work or for your research? Are you the maintainer of an open source computer vision tool or library? If you answered “Yes” to either question, we encourage you to consider speaking at a future event! **Host a local Meetup** Do you know of a meeting space that we could use to host a local Meetup? If so, please reach out. **Become a co-organizer** Are you interested in finding local speakers, hunting down a meeting space, helping with logistics, and MC-ing local events? Then consider becoming a local co-organizer. **Sponsor a local Meetup** Does your employer have a marketing or DevRel budget to help support the Meetup with administrative costs, swag giveaways and monthly charitable donations on behalf of the members? Please do not hesitate to get in touch! ## **A quick word from our sponsor** \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop [Voxel51](https://voxel51.com/) is the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes. [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/6a2b391258c94d51b612b5ae8b8f0fc15c431f57-1200x687.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing the Computer Vision Meetups Network Sponsored by Voxel51\\ \\ Product & News\\ \\ • \\ \\ Sep 28, 2022](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51) [![](https://cdn.sanity.io/images/h6toihm1/production/622b7369c791083b44e3034b2b8772d3ecada8bb-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – November 2023\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Nov 1, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-november-2023) [FiftyOne Computer Vision Community Update – February 2024\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Feb 9, 2024](https://voxel51.com/blog/fiftyone-computer-vision-community-update-february-2024) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-205-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Nov 4, 2022 Nov 5, 2022 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#a22cb9eca993) [Why aren’t my objects appearing in the App?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#ebb829c0a55b) [Removing attributes when exporting a dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#f9c82a6256ed) [Overwriting a dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#55f0c10d979a) [Clipping bounding boxes to be within the image area only](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#08692acdc8f6) [Support for visualizing 2D point clouds](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#c0aa1125b237) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#8a0c42f61623) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#a22cb9eca993) [Why aren’t my objects appearing in the App?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#ebb829c0a55b) [Removing attributes when exporting a dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#f9c82a6256ed) [Overwriting a dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#55f0c10d979a) [Clipping bounding boxes to be within the image area only](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#08692acdc8f6) [Support for visualizing 2D point clouds](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#c0aa1125b237) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022#8a0c42f61623) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/d63c2ef9fed7cf00ea9af1c6dd3f2ee4d657c598-1200x685.png?auto=format&dpr=2&fit=max&q=75&w=1200) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help Ok, let’s dive into this week’s tips and tricks! ## **Why aren’t my objects appearing in the App?** Community Slack member Matthew Millendorf asked, _“My `Detections` are not showing up in the FiftyOne App. Any suggestions on where the issue lies in the formatting of my data?”_ , 'tags': \[\], 'metadata': , 'inspection\_id': , 'ground\_truth': ,\ \ FiftyOne detections use relative coordinates, so you’ll need to update the bounding box coordinates to be within \[0, 1\] x \[0, 1\] by dividing by the image height and width.\ \ Learn more about [Detections](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#object-detection) in the FiftyOne Docs.\ \ ## **Removing attributes when exporting a dataset**\ \ Community Slack member Xuân Huy Nguyễn asked,\ \ _“When exporting a dataset, some CVAT attributes are also exported which I don’t need. How can I remove these attributes when I run an export?”_\ \ There are a few options worth exploring depending on your use case. Taking COCO as an example, there is an optional `extra_attrs=False ` parameter you can pass when exporting so as not to include custom attributes in the export. Using COCO as an example:\ \ dataset.export(\ \ ...,\ \ dataset\_type=fo.types.COCODetectionDataset,\ \ extra\_attrs=False,\ \ ...,\ \ )\ \ Learn more about [the COCO exporter and its available parameters](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html#cocodetectiondataset) in the FiftyOne Docs.\ \ Another option is to [create a view](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#filtering-sample-contents) that excludes the unwanted attributes from your detections, then export that view to CVAT format.\ \ view = dataset.exclude\_fields(\[\ \ "your\_field.detections.bad\_attr1",\ \ "your\_field.detections.bad\_attr2",\ \ ...\ \ \])\ \ view.export(\ \ ...,\ \ dataset\_type=fo.types.COCODetectionDataset,\ \ ...,\ \ )\ \ The final option is to use the [CVAT integration](https://voxel51.com/docs/fiftyone/integrations/cvat.html) directly for your annotation tasks.\ \ ## **Overwriting a dataset**\ \ Community Slack member Oğuz Hanoğlu asked,\ \ _“What is the best way to overwrite a dataset?”_\ \ Two options. One approach is to first delete the existing dataset and then create the new dataset:\ \ dataset\_name = dataset.name\ \ dataset.delete()\ \ dataset = fo.Dataset(dataset\_name)\ \ Or you can combine these operations into a single step by passing the optional `overwrite` argument to the `Dataset` constructor:\ \ dataset = fo.Dataset(dataset.name, overwrite=True)\ \ Learn more about [common and custom formats](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html#common-formats) in the FiftyOne Docs.\ \ ## **Clipping bounding boxes to be within the image area only**\ \ Community Slack member Patrick Rowsome asked,\ \ _“I have an issue with an object detection dataset where some of the ground truth annotations are outside of the image just slightly, maybe 1–3 px. Is there a convenient way in FiftyOne to clip these bounding boxes to be within the image area only?”_\ \ If the bounding box coordinates are outside of 0–1, then you should clip them to that range. The easiest way to do this is to just loop your dataset and make the changes in a Python loop.\ \ for sample in dataset.iter\_samples(autosave=True):\ \ for detection in sample\[label\_field\].detections:\ \ tlx1, tly1, w1, h1 = detection.bounding\_box\ \ tlx2 = max(min(tlx1, 1), 0)\ \ w2 = w1 - (tlx2-tlx1)\ \ tly2 = max(min(tly1, 1), 0)\ \ h2 = h1 - (tly2-tly1)\ \ w2 = max(min(w2+tlx2, 1), 0) - tlx2\ \ h2 = max(min(h2+tly2, 1), 0) - tly2\ \ detection\["bounding\_box"\] = \[tlx2, tly2, w2, h2\]\ \ Note, the [autosave context](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#efficient-batch-edits) lets you avoid needing to call `sample.save() ` each time.\ \ ## **Support for visualizing 2D point clouds**\ \ Community Slack member Matthew Millendorf asked,\ \ _“I am working on a geospatial point cloud registration problem, curious can you use FiftyOne for visualizing 2D point clouds where the points are lat/lon or UTM?_\ \ Yes, as of [version 0.17](https://medium.com/voxel51/announcing-fiftyone-0-17-0-with-grouped-datasets-3d-geolocation-and-custom-plugins-339600ab73a1), when you load a dataset in the App that contains a `GeoLocation` field with `point` data populated, you can open the Map tab to visualize a scatterplot of the location data:\ \ ![](https://cdn.sanity.io/images/h6toihm1/production/25d861c4d8c6421850812f4f83c85923656322d9-917x821.gif?auto=format&dpr=2&fit=max&q=75&w=917)\ \ import fiftyone as fo\ \ import fiftyone.zoo as foz\ \ dataset = foz.load\_zoo\_dataset("quickstart-geo")\ \ session = fo.launch\_app(dataset)\ \ ## **What’s next?**\ \ - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone)\ - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes\ - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help\ \ [bounding boxes](https://voxel51.com/blog/tag/bounding-boxes) [exporting](https://voxel51.com/blog/tag/exporting) [FAQ](https://voxel51.com/blog/tag/faq) [point clouds](https://voxel51.com/blog/tag/point-clouds)\ \ MT Admin\ \ Bio\ \ ### Talk to a computer vision expert\ \ [Book a demo](https://voxel51.com/sales)\ \ ### Related posts\ \ Discover more insights and tips to boost your visual AI workflows.\ \ [View all](https://voxel51.com/blog)\ \ [![](https://cdn.sanity.io/images/h6toihm1/production/4d37703d72d4b83a85bda19eb1999d5247915fc0-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 02, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 3, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/c28199522446929ca5288c0559047354eb93f01a-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Oct 28, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Oct 29, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/a17b9ee7620741f8c3225d0256d174c857a575d5-1200x672.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Oct 14, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Oct 15, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022)\ \ [Talk to a CV expert](https://voxel51.com/sales)\ \ Product\ \ [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing)\ \ Solutions\ \ [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security)\ \ Developers\ \ [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community)\ \ Resources\ \ [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research)\ \ Company\ \ [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press)\ \ [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51)\ \ © 2025 Voxel51 All Rights Reserved\ \ [Terms of Service](https://voxel51.com/terms-of-service)\ \ [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-206-lllmstxt|> ## FiftyOne Community Rewards [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Announcing Open Source FiftyOne Community Rewards! Nov 3, 2022 • 3 min read Article content In this article [Share your FiftyOne success story and receive a limited edition hoodie](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#966ee0f4a602) [Help make open source FiftyOne awesome!](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#aedf0a16c6d4) [Love helping others level up their FiftyOne knowledge?](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#b5dab3d3e63d) [Enjoy writing about technical topics?](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#1dddcf7a9a1f) [Interested in helping plan, host, or present at a Meetup?](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#a7e5582461cc) [What’s next?](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#50d017ecdb94) In this article [Share your FiftyOne success story and receive a limited edition hoodie](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#966ee0f4a602) [Help make open source FiftyOne awesome!](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#aedf0a16c6d4) [Love helping others level up their FiftyOne knowledge?](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#b5dab3d3e63d) [Enjoy writing about technical topics?](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#1dddcf7a9a1f) [Interested in helping plan, host, or present at a Meetup?](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#a7e5582461cc) [What’s next?](https://voxel51.com/blog/announcing-open-source-fiftyone-community-rewards#50d017ecdb94) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/984cff2eeb406055cefa6dd4f89895435de9d5d4-1200x674.png?auto=format&dpr=2&fit=max&q=75&w=1200) At [Voxel51](https://voxel51.com/), our mission is to bring transparency and clarity to the world’s data. We do this by offering open source ( [FiftyOne](https://voxel51.com/docs/fiftyone/)) and commercial ( [FiftyOne Teams](https://voxel51.com/fiftyone-teams/)) software to help engineers and the organizations they work for solve their most challenging computer vision problems. Beyond developing software, community is the core of how we operate at Voxel51. The FiftyOne community is all about bringing together people from different backgrounds and skill levels who share a common interest in computer vision, and then finding ways to help each other. In the coming weeks we’ll be rolling out the official “FiftyOne Community Rewards” program designed to inspire community members and recognize their efforts and achievements within the community. But, in the spirit of “release early, release often”, here’s a glimpse into contributions that are swag worthy. And there’s no need to wait for the official program to start contributing! We’re already on the lookout to reward contributions like these and are ready to send some swag your way to say thanks. ![](https://cdn.sanity.io/images/h6toihm1/production/300002f06cd511483702ede5ef076418f27c718b-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) ## **Share your FiftyOne success story and receive a limited edition hoodie** Is your company using FiftyOne to solve interesting computer vision problems? If so, [consider sharing your success story](https://www.voxel51.com/fiftyone-computer-vision-success-story-submission/) and claim a box of FiftyOne swag as a thank you! For example, check out how [Forsight uses FiftyOne](https://medium.com/voxel51/forsight-finds-a-centralized-dataset-management-solution-in-fiftyone-teams-1220b445c8dc) in combination with cutting edge AI vision technology to solve some of the construction industry’s greatest challenges regarding safety, security, management and more. ![](https://cdn.sanity.io/images/h6toihm1/production/db39339dd1e14edb9a9de8deb8ba2b9a91c83eb3-1400x828.png?auto=format&dpr=2&fit=max&q=75&w=1400) **_What qualifies as a success story?_** If you are using FiftyOne to deliver production computer vision projects at your company either internally or for clients and customers, that’s a success story! Are you using FiftyOne in your research project in industry or at a university? That’s a success story too! **_How will my success story be used?_** We are in the process of collecting success stories to showcase on our landing page. The goal of the landing page is to communicate not only the FiftyOne community’s diversity and international reach, but also its ability to help solve today’s most challenging computer vision problems across a variety of industries. There are plenty of other ways you can get involved in the community. If you do decide to get involved, don’t be surprised if someone from the Voxel51 developer relations team reaches out looking to send you some cool FiftyOne swag! ## **Help make open source FiftyOne awesome!** - Found a [bug](https://github.com/voxel51/fiftyone/issues/new?assignees=&labels=bug&template=bug_report_template.md&title=%5BBUG%5D) in FiftyOne? - Like to write code? Have a look at one of our [“Good First Issues”](https://github.com/voxel51/fiftyone/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22). - Found a [typo or error](https://github.com/voxel51/fiftyone/issues/new?assignees=&labels=documentation&template=documentation_issue_template.md&title=%5BDOCUMENTATION%5D) in our Docs? - Have a [feature request](https://github.com/voxel51/fiftyone/issues/new?assignees=&labels=enhancement&template=feature_request_template.md&title=%5BFR%5D)? - Experiencing issues getting [FiftyOne installed](https://github.com/voxel51/fiftyone/issues/new?assignees=&labels=bug&template=installation_issue_template.md&title=%5BSETUP-BUG%5D)? ## **Love helping others level up their FiftyOne knowledge?** - Answer questions related to [open FiftyOne issues](https://github.com/voxel51/fiftyone/pulls) - Answer questions and provide suggestions in [FiftyOne community Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) - Answer questions about FiftyOne anywhere the community, like Reddit or Stack Overflow ## **Enjoy writing about technical topics?** - Interested in creating a blog post about FiftyOne? Ping @Jacob Marks or @Jimmy Guerrero in [FiftyOne Community Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) and we’d be happy to help with topic selection, editing and encouragement! - Help translate FiftyOne content into other languages to help reach members who’d prefer to consume content in their native language. Ping us on FiftyOne Community Slack. ![](https://cdn.sanity.io/images/h6toihm1/production/0d65f67f35ed7c2afbfc5e6978967e492b357c43-1250x705.png?auto=format&dpr=2&fit=max&q=75&w=1250) ## **Interested in helping plan, host, or present at a Meetup?** We are always looking for co-hosts, co-sponsors, and presenters for the [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/) network. We currently meet virtually every month, but would love to find local computer vision enthusiasts to help us start hosting in-person events. Drop us a line at [community@voxel51.com](mailto:community@voxel51.com) to get the conversation started. ## **What’s next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help [community rewards](https://voxel51.com/blog/tag/community-rewards) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/286bca6ba83c8386a9b92750a9249e2e8aba7d8a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – June ‘23\\ \\ Product & News\\ \\ • \\ \\ Jun 2, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/99dc871d875865fc200f14931a3cfb3118ded624-1020x1007.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Tunnel vision in computer vision: can ChatGPT see?\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 16, 2022](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see) [![](https://cdn.sanity.io/images/h6toihm1/production/a735267ad7effa9f799f850ab7c8ffa241088710-1024x1024.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Why 2022 was the most exciting year in computer vision history (so far)\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 14, 2022](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-207-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Oct 28, 2022 Oct 29, 2022 • 3 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#c778779b931a) [Counting bounding boxes and visualizing anchor boxes](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#686df00b0d53) [Computing embeddings on labels](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#876b02ae3c0f) [Working with features metadata](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#b7a01a8e1a9e) [Remapping labels](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#4c7112c8d598) [Matching tags within a Detections field](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#64c332f09dac) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#21583022fa91) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#c778779b931a) [Counting bounding boxes and visualizing anchor boxes](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#686df00b0d53) [Computing embeddings on labels](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#876b02ae3c0f) [Working with features metadata](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#b7a01a8e1a9e) [Remapping labels](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#4c7112c8d598) [Matching tags within a Detections field](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#64c332f09dac) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022#21583022fa91) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/c28199522446929ca5288c0559047354eb93f01a-1200x673.png?auto=format&dpr=2&fit=max&q=75&w=1200) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help Ok, let’s dive into this week’s tips and tricks! ## **Counting bounding boxes and visualizing anchor boxes** Community Slack member Sidney Guaro asked a two part question, _“How do you count the bounding boxes at different area scales? (e.g. `area_small, area_medium, area_large`) And is it possible to visualize the anchor boxes?”_ If you need to filter detections by the bounding box area use `ViewExpression` You can visualize any number of sets of boxes you like if you extract the relevant information from your model and add them as `Detections` to your FiftyOne dataset. Learn more about the [`ViewExpression`](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#filtering-detections-by-area) and [adding object detections to a dataset](https://voxel51.com/docs/fiftyone/recipes/adding_detections.html?highlight=detections) in the FiftyOne Docs. ## **Computing embeddings on labels** Community Slack member Nadav Ben-Haim asked, _“Is it possible to compute embeddings on labels rather than images? Referencing this [tutorial](https://voxel51.com/docs/fiftyone/tutorials/image_embeddings.html), it appears you can do it on full images, but I am interested in computing such embeddings on crops from bounding box labels.”_ Yes! You can pass in the `patches_field` parameter to the [visualization](https://voxel51.com/docs/fiftyone/user_guide/brain.html#object-embeddings-example) and the [similarity functions](https://voxel51.com/docs/fiftyone/user_guide/brain.html#object-similarity) available in FiftyOne Brain to use object-level embeddings rather than image-level. Learn more about [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html#fiftyone-brain) in the FiftyOne Docs. ## **Working with features metadata** Community Slack member Adrian Loy asked, _“Can I see the distributions of metadata (e.g. frame count) features, similar to how I can see the distribution over Labels?”_ Yes! You can easily pull out any specific metadata fields of interest and add them as their own field with the following: dataset.set\_values(“width”, dataset.values(“metadata.width”)) Alternatively, check out some of our [Plotly integrations](https://voxel51.com/docs/fiftyone/user_guide/plots.html#id11) to see how to create custom plots on any fields of interest. ( **Note**: If you work in notebooks, then they can also be interactive! But we also plan to add plugins to do this right in the App. Stay tuned!) ## **Remapping labels** Community Slack member Patrick Rowsome asked, _“I am trying to merge datasets and would like to map class names to be the same. For example, one dataset uses the person label and another uses pedestrian, ideally I would like to map all instances of pedestrians to the person label.”_ Use the `map_labels()` function. It lets you create a view where labels are remapped, but if you want to make it permanent you can just call `view.save()` Learn more about the [map\_labels()](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html?highlight=map_labels#fiftyone.core.collections.SampleCollection.map_labels) function in the FiftyOne Docs. ## **Matching tags within a Detections field** Community Slack member Geoffrey Keating asked, _“What would be the best way to match tags within a `fo.Detections` field? I have individual detections tagged and would like to recall them.”_ Try `select_labels()` which selects only the specified labels from the collection, with the returned view omitting samples, sample fields, and individual labels that do not match the specified selection criteria. Alternatively, you can always write a view expression with `filter_labels()`: from fiftyone import ViewField as F view = dataset.filter\_labels("detections\_field", F("tags").contains("my-tag")) ## **What’s next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help [anchor boxes](https://voxel51.com/blog/tag/anchor-boxes) [bounding boxes](https://voxel51.com/blog/tag/bounding-boxes) [embeddings](https://voxel51.com/blog/tag/embeddings) [FAQ](https://voxel51.com/blog/tag/faq) [metadata](https://voxel51.com/blog/tag/metadata) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/25a70e1c9c2cae7f9df18696e0e087dddf67707d-1200x679.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Sept 23, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 24, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/4d37703d72d4b83a85bda19eb1999d5247915fc0-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 02, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 3, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/d63c2ef9fed7cf00ea9af1c6dd3f2ee4d657c598-1200x685.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Nov 4, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 5, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) FiftyOne Computer Vision Tips and Tricks — Oct 28, 2022 - Voxel51 <|firecrawl-page-208-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Oct 21, 2022 Oct 22, 2022 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#c9e3c673cbf7) [Using sample tags to get a subset of images by class](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#0116b5f27734) [Persisting datasets and loading them directly](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#d1deb3f3f50b) [Exploring and filtering video datasets by a specific label](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#8f5c935baa58) [Speeding up dataset load times with FiftyOne Brain](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#f856ee915af7) [Importing and exporting datasets to the cloud](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#36ef9e49b65a) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#1048b32cc3e9) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#c9e3c673cbf7) [Using sample tags to get a subset of images by class](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#0116b5f27734) [Persisting datasets and loading them directly](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#d1deb3f3f50b) [Exploring and filtering video datasets by a specific label](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#8f5c935baa58) [Speeding up dataset load times with FiftyOne Brain](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#f856ee915af7) [Importing and exporting datasets to the cloud](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#36ef9e49b65a) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-21-2022#1048b32cc3e9) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/32ce8f9c9ced60f3c38cb5c950e34abc6dccdd86-1200x735.png?auto=format&dpr=2&fit=max&q=75&w=1200) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help Ok, let’s dive into this week’s tips and tricks! ## **Using sample tags to get a subset of images by class** Community Slack member Gaurav Savlani asked, _“How do I do something like `view.groupby([groupbycolumn]).apply(somefunction)` like I would in pandas? For example: if I have 100 classes and I want to subset 10 images from each class. Is there functionality to this?_ Although there is a [`group_by()`](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.group_by) view stage, it doesn’t currently support filtering the elements of each group. We’d recommend using sample tags to encode the sampling of each group. For example: import fiftyone as fo import fiftyone.zoo as foz from fiftyone import ViewField as F dataset = foz.load\_zoo\_dataset("cifar10", split="test") for label in dataset.distinct("ground\_truth.label"): view = dataset.match(F("ground\_truth.label") == label).take(10) view.tag\_samples("sample") view = dataset.match\_tags("sample") print(view.count\_values("ground\_truth.label")) \# {'airplane': 10, 'ship': 10, ..., 'cat': 10} ## **Persisting datasets and loading them directly** Community Slack member Sidney Guaro asked, _“Whenever I start a new program, do I need to process the dataset again or can I directly load it using `fo.load_dataset()`?”_ Once you make your dataset persistent with `dataset.save()`, you can load it directly with `fo.load_dataset()`. Note that `dataset.save()` is only required when you edit a property like `dataset.info` in-place. For example: dataset.info\["new\_field"\] = "new-value" dataset.save() # required in order for changes to save \# Save automatically occurs in these cases dataset.info = {"new-field": "new-value"} dataset.persistent = True Learn more about [saving changes to your dataset](https://voxel51.com/docs/fiftyone/faq/index.html#why-didn-t-changes-to-my-dataset-save) in the FiftyOne Docs. ## **Exploring and filtering video datasets by a specific label** Community Slack member Adian Loy asked, _“I have a dataset consisting of videos and frame-by-frame classification labels. I want to explore this dataset by filtering on snippets of a specific label. Is there a way for FiftyOne to do that automatically or on the fly?_ Yes, there is! Here’s the steps: **Step 1:** [Load your videos](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html#custom-formats) into a dataset and add your frame-level classification \# Pseudocode for what that looks like, parsing depends on your raw data format samples = \[\] for video\_filepath, frame\_classifications in my\_raw\_data: sample = fo.Sample(filepath=video\_filepath) # Add frame classifications for frame\_number, frame\_classification in frame\_classifications: classification = fo.Classification(label=frame\_classification) sample.frames\[frame\_number\]\["ground\_truth"\] = fo.Classifications( classifications=\[classification\] ) samples.append(sample) dataset = fo.Dataset("my-dataset") dataset.add\_samples(samples) **Step 2:** Filter the frame labels to [create a view](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#filtering) for the specific label of interest: from fiftyone import ViewField as F filtered\_view = dataset.filter\_labels( "frames.ground\_truth", F("label") == "your-label" ) **Step 3:** You can then turn this filtered video view into a [temporary clips view](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#clip-views) where every contiguous set of frames that has the classification label you filtered by gets turned into its own clip on the fly: clips\_view = filtered\_view.to\_clips(“frames.ground\_truth”) When you then visualize the clips\_view in the App, each clip is shown separately in the grid and when you click on it, you only see the frames associated with that clip. session = fo.launch\_app(clips\_view) If you then [export the clip dataset to disk](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html#video-clips) in the future, only then are the clip media actually extracted from the video and saved. ## **Speeding up dataset load times with FiftyOne Brain** Community Slack member Oğuz Hanoğlu asked, _“ `find_duplicates()` takes about 5 min to run with 180k samples. Is it possible to save its `results(neighbors_map...)` and load them the next time I open the notebook?”_ We’d recommend taking advantage of the [FiftyOne Brain](https://voxel51.com/docs/fiftyone/api/fiftyone.brain.html#) in this scenario. To start, these are the attributes that are set by `find_duplicates():` results.\_thresh = thresh results.\_unique\_ids = unique\_ids results.\_duplicate\_ids = duplicate\_ids results.\_neighbors\_map = neighbors\_map So, in your case you’ll want to make use of a `brain_key`: results = dataset.load\_brain\_results(brain\_key) \# Load results from a previous call to \`find\_duplicates()\` results.\_thresh = thresh results.\_unique\_ids = unique\_ids results.\_duplicate\_ids = duplicate\_ids results.\_neighbors\_map = neighbors\_map plot = results.visualize\_duplicates(...) plot.show() ## **Importing and exporting datasets to the cloud** Community Slack member George Pearse asked, _“Do dataset exports currently support loading to the cloud? e.g. like model checkpoints that just allow you to put a cloud path in there?”_ This capability is a feature of the [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) product. The Teams Python SDK fully supports cloud paths everywhere that you would be using local paths using the OSS library (for example when importing and exporting datasets.) FiftyOne Teams enables multiple users to securely collaborate on the same datasets and models, either on-premises or in the cloud, all built on top of the open source FiftyOne workflows that you’re already relying on. Learn more about [FiftyOne Teams](https://voxel51.com/fiftyone-teams/). ## **What’s next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne Brain](https://voxel51.com/blog/tag/fiftyone-brain) [persisting datasets](https://voxel51.com/blog/tag/persisting-datasets) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/03107d477b7db4be03031293fa4fe15aaea806f0-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Sept 16, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 17, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/ecb6afb20436d0f0e68fbb25bcfc7443657d7b91-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Feb 10, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Feb 10, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/3f54d0a45faa06a04b5d0244dd7c092603150cf0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Mar 10, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 11, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-209-lllmstxt|> ## Voxel51 Turns Four [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) It’s Our Birthday! (Voxel51 Turns Four) Oct 19, 2022 • 4 min read Article content In this article [Voxel51 Throughout the Years](https://voxel51.com/blog/its-our-birthday-voxel51-turns-four#0dbc40ac040b) [Wait, Who’s Voxel51 and What Do You Do?](https://voxel51.com/blog/its-our-birthday-voxel51-turns-four#eb58aa88d08e) [Celebrating Our 4 Years — and More!](https://voxel51.com/blog/its-our-birthday-voxel51-turns-four#cf4be6c6ffb4) [What’s Next](https://voxel51.com/blog/its-our-birthday-voxel51-turns-four#3c215fabf8cd) In this article [Voxel51 Throughout the Years](https://voxel51.com/blog/its-our-birthday-voxel51-turns-four#0dbc40ac040b) [Wait, Who’s Voxel51 and What Do You Do?](https://voxel51.com/blog/its-our-birthday-voxel51-turns-four#eb58aa88d08e) [Celebrating Our 4 Years — and More!](https://voxel51.com/blog/its-our-birthday-voxel51-turns-four#cf4be6c6ffb4) [What’s Next](https://voxel51.com/blog/its-our-birthday-voxel51-turns-four#3c215fabf8cd) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/610db1faf979759ee38cd4aec9714c4cd03d1d81-1509x1025.png?auto=format&dpr=2&fit=max&q=75&w=1509) Four years ago today, Voxel51 Inc. was born! It’s been an incredible journey so far and we wanted to take a moment to reflect on where we’ve been and celebrate what we’ve accomplished so far as a FiftyOne community. ## Voxel51 Throughout the Years Here are some key dates and milestones from our first four years as a company: October 18, 2018: Started Voxel51 Inc. to enable developers, scientists, and organizations to build high-quality datasets and computer vision models August 8, 2019: Announced Seed Funding June 1, 2020: Released FiftyOne 0.1 to a few dozen private-beta users August 11, 2020: Open sourced FiftyOne, making it **_the_** open source tool for building high-quality datasets and computer vision models Late July, 2021: Began working with dozens of startups and Fortune 500 enterprises as early adopters of FiftyOne Teams September 21, 2022: Announced Series A funding and the public availability of FiftyOne Teams Today: Voxel51 turns four! ## Wait, Who’s Voxel51 and What Do You Do? If you’re new to Voxel51 and FiftyOne, no worries — here’s a little bit about us. We’re [Voxel51](https://voxel51.com/) and our mission is to bring transparency and clarity to the world’s data. We’re the company behind the [open source FiftyOne project](https://github.com/voxel51/fiftyone), the open source tool for building high-quality datasets and computer vision models. We also build [FiftyOne Teams](https://voxel51.com/fiftyone-teams/), our commercial product that enables teams to securely collaborate on their datasets and models. Tens of thousands of engineers and scientists have integrated open source FiftyOne into their ML workflows. It’s easy to get up and running in minutes — learn how [in the docs](https://voxel51.com/docs/fiftyone/index.html). ## Celebrating Our 4 Years — and More! Not only do we love celebrating our years of incorporated bliss with [sugary treats](https://www.instagram.com/p/B3wzXnfpMd2/), but we also — and more importantly — love celebrating our community and customers. ### Celebrating FiftyOne & Our Community Today we are proud to share that **[FiftyOne](https://github.com/voxel51/fiftyone) has crossed 2000 stars on GitHub!** Open source software doesn’t happen without an amazing community supporting it, so a big thank you to all our stargazers and everyone in the FiftyOne community for your support and contributions over the years. Since open sourcing FiftyOne in August 2020, we have released tons of new capabilities and features. The latest big FiftyOne release was last month — v0.17 added support for 3D datasets, grouped datasets, geolocation data, custom plugins, and more. You can read more about it in the [release blog post](https://medium.com/voxel51/announcing-fiftyone-0-17-0-with-grouped-datasets-3d-geolocation-and-custom-plugins-339600ab73a1). Many of the new goodies over the years have come from you — through your PRs, GitHub issues, feedback, and questions in the community Slack. Thank you for making FiftyOne what it is today and as it continues to evolve! Now, with full support for images, video, 3D, and geolocation, we hope FiftyOne can become a helpful hub for all your datasets throughout your ML workflows. Other FiftyOne community milestones that we’re celebrating today include: - 1,075+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 1,300+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used in](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 178 repositories - 43 open source [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) Thanks again to our growing and enthusiastic community, we look forward to reaching many more milestones with you! ### Celebrating FiftyOne Teams & Our Customers A little over a year ago, we began working with dozens of startups and Fortune 500 enterprises as early adopters of FiftyOne Teams, our commercial product that enables teams to securely collaborate on their datasets and models. By providing us with real-world usage, input, and feedback, our early adopters helped us shape and harden the [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) solution that we announced last month. FiftyOne Teams helps customers across a variety of industries, including automotive, autonomous vehicles, robotics, security, retail, healthcare, and more. Companies big, small, and everywhere in between have seen tremendous value in extending FiftyOne into a team-centered implementation for their team or entire organization. A huge shoutout and thank you to all of our early adopters for making FiftyOne Teams what it is today, including these customers who have this to share about their experiences: > “We’ve used FiftyOne Teams at ADT Commercial for a year now. It has helped us manage our huge datasets, collaborate on model evaluation, tighten our production schedule, and ultimately deliver solutions that help our customers better manage their risk. FiftyOne Teams has added tremendous value to our computer vision processes.” > > \- Philippe Sawaya, Director of Artificial Intelligence, ADT Commercial > “We use FiftyOne Teams to organize, select, display, and share our data which has led to better collaboration with and understanding of our large volume of data. FiftyOne Teams enables us to gain insights such as identifying and understanding data problems early, hypothesis validation, and dataset management overall. This has led to better solution engineering and better testing for the products and services we deliver to our customers.” – Lanny Lin, Senior Director of AI and Data Science, Vivint ### Celebrating People at Voxel51 Voxel51 wouldn’t be here today without the brilliant and thoughtful people that make up our team. Thank you to everyone who has shaped Voxel51 so far on our incredible journey. And to those who will be part of the future of that journey: we’re just getting started — the best is yet to come! We’re building a fully-remote team of people who want to help us achieve our mission of bringing transparency and clarity to the world’s data, and we’re going to need many more exceptional and diverse people to help us. If our mission excites you, check out our [open positions across product, engineering, community, and more](https://voxel51.com/jobs/). ## What’s Next - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started](https://voxel51.com/docs/fiftyone/index.html), we’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help - Check out our [open positions](https://voxel51.com/jobs/) and consider joining us on our mission [Voxel51 birthday](https://voxel51.com/blog/tag/voxel51-birthday) [Voxel51 milestone](https://voxel51.com/blog/tag/voxel51-milestone) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/2e1ee531befe139f840d795b4bab00275367f6be-1200x675.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Celebrating Three Years of FiftyOne!\\ \\ Product & News\\ \\ • \\ \\ Aug 19, 2023](https://voxel51.com/blog/celebrating-three-years-of-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/0f4cab7e19991301a33a2691bbfeb2d9a22024fc-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Four Years of Open Source FiftyOne!\\ \\ Product & News\\ \\ • \\ \\ Aug 12, 2024](https://voxel51.com/blog/four-years-of-open-source-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/a92a7366fc91c4f31c0883f77d0add2ff4abadb8-1841x963.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Happy 5th Birthday, Voxel51!\\ \\ Product & News\\ \\ • \\ \\ Oct 18, 2023](https://voxel51.com/blog/happy-5th-birthday-voxel51) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-210-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Oct 14, 2022 Oct 15, 2022 • 5 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#41b493b970e0) [Sorting samples in a collection by fields or expressions](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#a34b0f3c0810) [Connecting the FiftyOne client to MongoDB](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#6d614b5d0d72) [Mapping ground truth labels in a bounding box problem](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#5bc3da943ee0) [Merging samples and updating labels](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#77bbddd87688) [Using FiftyOne datasets with the PyTorch dataloader](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#995d5ec9c12e) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#449a4fb0104c) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#41b493b970e0) [Sorting samples in a collection by fields or expressions](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#a34b0f3c0810) [Connecting the FiftyOne client to MongoDB](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#6d614b5d0d72) [Mapping ground truth labels in a bounding box problem](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#5bc3da943ee0) [Merging samples and updating labels](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#77bbddd87688) [Using FiftyOne datasets with the PyTorch dataloader](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#995d5ec9c12e) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022#449a4fb0104c) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/a17b9ee7620741f8c3225d0256d174c857a575d5-1200x672.png?auto=format&dpr=2&fit=max&q=75&w=1200) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Ok, let’s dive into this week’s tips and tricks! ## **Sorting samples in a collection by fields or expressions** Community Slack member Sybil Lyu asked, _“Can I use `SortBy` with an expression in FiftyOne via the `add stage` button in the FiftyOne App?”_ Today, the best way to sort by an expression is via Python. If you run some of [the examples](https://voxel51.com/docs/fiftyone/api/fiftyone.core.stages.html#fiftyone.core.stages.SortBy) from the Docs, you’ll see the equivalent JSON that you’d need to type into the `SortBy` stage in the view bar of the App to create the view: \# Sort by number of GT objects view = dataset.sort\_by(F("ground\_truth.detections").length(), reverse=True) \# Click into view bar to see equivalent JSON session = fo.launch\_app(view) ## **Connecting the FiftyOne client to MongoDB** Community Slack member Naman Gupta asked, _“Is it possible to connect to an already running local FiftyOne instance from a Jupyter notebook without having to spin up another FiftyOne server? I want to connect to the locally running instance and either load, filter or export datasets on the same machine.”_ You can run MongoDB in a separate container and then configure your FiftyOne client in the Jupyter container to [connect to it](https://voxel51.com/docs/fiftyone/user_guide/config.html#configuring-a-mongodb-connection). Other than MongoDB, there’s no FiftyOne “server” in the open source package. [FiftyOne Teams](https://voxel51.com/fiftyone-teams/), on the other hand, provides a centralized MongoDB database and a FiftyOne App server allowing everyone on your team to easily load the same datasets in Python and the App. Learn more about working with [data](https://voxel51.com/docs/fiftyone/environments/index.html#local-data) and [notebooks](https://voxel51.com/docs/fiftyone/environments/index.html#notebooks) in the FiftyOne Docs. ## **Mapping ground truth labels in a bounding box problem** Community Slack member Raghav Mecheri asked, _“I’m trying to map a set of ground truth labels to a broader set of categories for a bounding box problem. For example turning bounding boxes that have labels for “audi”, “bmw”, “mercedes” all into “car”. I could iterate through each image as I load it, but I feel that there’s probably a “right” way to do this in FiftyOne — any good starting points?”_ You can use [`map_labels()`](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.map_labels) for this! `view = dataset.map_labels(...)` This will give you a view that dynamically renames the labels when you iterate over/visualize it in the App. If you want to save the changes to the actual dataset, just add: `view.save()` ## **Merging samples and updating labels** Community Slack member Jason Barbee asked, _“I cloned a dataset, changed the ground\_truth labels on samples, but running main\_dataset.merge(working\_dataset) doesn’t seem to overwrite my existing labels. Is there a replace\_sample type API?”_ By default, when using `merge_samples()`, the `merge_lists` attribute is `True`, meaning that for lists of labels like detections, the two lists will be merged based on label ID rather than the working dataset overwriting all main dataset labels. If you set `merge_lists=False`, then it will discard all existing labels and keep only the labels from the dataset being merged in. Learn more about [`merge_samples`](https://voxel51.com/docs/fiftyone/api/fiftyone.core.dataset.html?highlight=merge_samples#fiftyone.core.dataset.Dataset.merge_samples) in the FiftyOne Docs. ## **Using FiftyOne datasets with the PyTorch dataloader** Community Slack member Sidney Guaro asked, _“Is it possible to use a FiftyOne dataset in PyTorch dataloader?”_ We do have some integrations with [PyTorch Lightning Flash](https://voxel51.com/docs/fiftyone/integrations/lightning_flash.html), as well as a [Detectron2 tutorial](https://voxel51.com/docs/fiftyone/tutorials/detectron2.html). But you can also always integrate FiftyOne datasets right into PyTorch dataloaders [(check out this blog)](https://towardsdatascience.com/stop-wasting-time-with-pytorch-datasets-17cac2c22fa8). Here is an example from the blog that sets up a torch dataset from FiftyOne: import torch import fiftyone.utils.coco as fouc from PIL import Image class FiftyOneTorchDataset(torch.utils.data.Dataset): """A class to construct a PyTorch dataset from a FiftyOne dataset. Args: fiftyone\_dataset: a FiftyOne dataset or view that will be used for training or testing transforms (None): a list of PyTorch transforms to apply to images and targets when loading gt\_field ("ground\_truth"): the name of the field in fiftyone\_dataset that contains the desired labels to load classes (None): a list of class strings that are used to define the mapping between class names and indices. If None, it will use all classes present in the given fiftyone\_dataset. """ def \_\_init\_\_( self, fiftyone\_dataset, transforms=None, gt\_field="ground\_truth", classes=None, ): self.samples = fiftyone\_dataset self.transforms = transforms self.gt\_field = gt\_field self.img\_paths = self.samples.values("filepath") self.classes = classes if not self.classes: # Get list of distinct labels that exist in the view self.classes = self.samples.distinct( "%s.detections.label" % gt\_field ) if self.classes\[0\] != "background": self.classes = \["background"\] + self.classes self.labels\_map\_rev = {c: i for i, c in enumerate(self.classes)} def \_\_getitem\_\_(self, idx): img\_path = self.img\_paths\[idx\] sample = self.samples\[img\_path\] metadata = sample.metadata img = Image.open(img\_path).convert("RGB") boxes = \[\] labels = \[\] area = \[\] iscrowd = \[\] detections = sample\[self.gt\_field\].detections for det in detections: category\_id = self.labels\_map\_rev\[det.label\] coco\_obj = fouc.COCOObject.from\_label( det, metadata, category\_id=category\_id, ) x, y, w, h = coco\_obj.bbox boxes.append(\[x, y, x + w, y + h\]) labels.append(coco\_obj.category\_id) area.append(coco\_obj.area) iscrowd.append(coco\_obj.iscrowd) target = {} target\["boxes"\] = torch.as\_tensor(boxes, dtype=torch.float32) target\["labels"\] = torch.as\_tensor(labels, dtype=torch.int64) target\["image\_id"\] = torch.as\_tensor(\[idx\]) target\["area"\] = torch.as\_tensor(area, dtype=torch.float32) target\["iscrowd"\] = torch.as\_tensor(iscrowd, dtype=torch.int64) if self.transforms is not None: img, target = self.transforms(img, target) return img, target def \_\_len\_\_(self): return len(self.img\_paths) def get\_classes(self): return self.classes ## **What’s next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help [bounding boxes](https://voxel51.com/blog/tag/bounding-boxes) [FAQ](https://voxel51.com/blog/tag/faq) [MongoDB](https://voxel51.com/blog/tag/mongodb) [PyTorch](https://voxel51.com/blog/tag/pytorch) [sorting](https://voxel51.com/blog/tag/sorting) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/4d37703d72d4b83a85bda19eb1999d5247915fc0-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 02, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 3, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/d63c2ef9fed7cf00ea9af1c6dd3f2ee4d657c598-1200x685.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Nov 4, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 5, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/c28199522446929ca5288c0559047354eb93f01a-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Oct 28, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Oct 29, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-211-lllmstxt|> ## Hacktoberfest FiftyOne Contributions [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Hack on FiftyOne in Hacktoberfest 2022 Oct 14, 2022 • 3 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/hack-on-fiftyone-in-hacktoberfest-2022#c4a20a2cd7e6) [Open Source Is in Our DNA](https://voxel51.com/blog/hack-on-fiftyone-in-hacktoberfest-2022#c125a30a7a77) [How to Contribute to FiftyOne](https://voxel51.com/blog/hack-on-fiftyone-in-hacktoberfest-2022#f171ca515f84) [What’s Next](https://voxel51.com/blog/hack-on-fiftyone-in-hacktoberfest-2022#8444a7f70327) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/hack-on-fiftyone-in-hacktoberfest-2022#c4a20a2cd7e6) [Open Source Is in Our DNA](https://voxel51.com/blog/hack-on-fiftyone-in-hacktoberfest-2022#c125a30a7a77) [How to Contribute to FiftyOne](https://voxel51.com/blog/hack-on-fiftyone-in-hacktoberfest-2022#f171ca515f84) [What’s Next](https://voxel51.com/blog/hack-on-fiftyone-in-hacktoberfest-2022#8444a7f70327) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/8f7e5027ba5809a1d70d1def7ee6a8a4e0d75870-1140x599.png?auto=format&dpr=2&fit=max&q=75&w=1140) It’s [Hacktoberfest](https://hacktoberfest.com/), the month-long hack celebration thrown annually by DigitalOcean that is all about getting more people involved in open source. Although we’re a little (fashionably?!) late to the party, we wanted to let you know we added the ”hacktoberfest” topic to the FiftyOne project on GitHub again this year and are open for Hacktoberfest contributions! For those of you participating in Hacktoberfest, we hope you’ll consider hacking on open source FiftyOne and we look forward to seeing and accepting your pull/merge requests. If you’re looking for a good place to get started, we tagged a collection of “ [good first issues](https://github.com/voxel51/fiftyone/issues?q=is%3Aopen+is%3Aissue+label%3A%22good+first+issue%22)” in GitHub for you to consider. If you need assistance, reach out anytime in the [FiftyOne Community Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ). We’re standing by to help you be successful with FiftyOne this Hacktoberfest and beyond! ## Wait, What’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. ## Open Source Is in Our DNA We first released the open source FiftyOne project a little over two years ago, and since then we’ve added a ton of new features in partnership with the community, which has now grown to tens of thousands of engineers and their forward-looking companies, 1900+ GitHub stars, 1000+ Slack members, and 1000+ Meetup members. All-year-round we welcome community contributions of all kinds — whether you submit GitHub issues to see new features you’d like added or make contributions to the project itself, everyone is welcome. As a fellow group of open source maintainers and enthusiasts ourselves, we love everything that Hacktoberfest stands for: encouraging more people to get involved in open source, encouraging collaboration, encouraging contributions of all types at all levels, especially making it super friendly for people who are just getting started. In fact, looking at the Hacktoberfest values, we feel a lot of natural alignment with our own values in the FiftyOne community: ![](https://cdn.sanity.io/images/h6toihm1/production/99db548d117d50128561daac1d4dd658accd2a65-1054x626.png?auto=format&dpr=2&fit=max&q=75&w=1054) These are just some of the reasons we are excited to be part of Hacktoberfest again this year. ## How to Contribute to FiftyOne Community contributions are always welcome to the open source FiftyOne project! Check out the [contribution guide](https://github.com/voxel51/fiftyone/blob/develop/CONTRIBUTING.md) to learn how to get involved. The guide covers: - GitHub issues - Pull requests - Contribution guidelines - Developer guide - Documentation - More! In addition, we tagged a collection of “ [good first issues](https://github.com/voxel51/fiftyone/issues?q=is%3Aopen+is%3Aissue+label%3A%22good+first+issue%22)” in GitHub for you to consider working on. Here are just a few examples, but please do browse the entire collection to find the one you’d like to contribute to. ![](https://cdn.sanity.io/images/h6toihm1/production/87ec4cce45002449975ffe85c74297f8af16ca1e-1400x568.png?auto=format&dpr=2&fit=max&q=75&w=1400) Finally, join the **#contributors** channel in the Slack community anytime to get help, discuss ideas, or ask questions about contributing to open source FiftyOne. Hope to see you there! ## What’s Next - Find out more about [participating in the Hacktoberfest](https://hacktoberfest.com/participation/) event. - [Get started with FiftyOne!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. Take it for a spin and get a sense for what contributions you want to make. - Make your contributions using [this guide](https://github.com/voxel51/fiftyone/blob/develop/CONTRIBUTING.md) as a foundation. - If you need any assistance along the way, join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ); we’re always happy to help, during Hacktoberfest and beyond. [Hacktoberfest](https://voxel51.com/blog/tag/hacktoberfest) [open source](https://voxel51.com/blog/tag/open-source) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/29fa790b13637c569efda3fa1a797042c04d6bad-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – Feb ‘23\\ \\ Product & News\\ \\ • \\ \\ Feb 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/2e1ee531befe139f840d795b4bab00275367f6be-1200x675.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Celebrating Three Years of FiftyOne!\\ \\ Product & News\\ \\ • \\ \\ Aug 19, 2023](https://voxel51.com/blog/celebrating-three-years-of-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/0f4cab7e19991301a33a2691bbfeb2d9a22024fc-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Four Years of Open Source FiftyOne!\\ \\ Product & News\\ \\ • \\ \\ Aug 12, 2024](https://voxel51.com/blog/four-years-of-open-source-fiftyone) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-212-lllmstxt|> ## Webinar Recap: FiftyOne Updates [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Webinar Recap: What’s New in FiftyOne & FiftyOne Teams Oct 8, 2022 • 8 min read Article content In this article [Donating $200 to World Literacy Foundation](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#81a8c4226f65) [Presentation Highlights](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#ff74ba8bf187) [Introduction to open source FiftyOne](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#0c3652242d88) [What’s new in the latest version (v.17) of FiftyOne](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#b9a02ab692da) [FiftyOne Open Source — Live Demo Time!](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#8c2e5ac6eb97) [FiftyOne Teams — Live Demo Time!](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#479b5339a68f) [Q&A from the Webinar](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#caca906cebc6) [What’s Next](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#86ec1360ef5c) In this article [Donating $200 to World Literacy Foundation](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#81a8c4226f65) [Presentation Highlights](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#ff74ba8bf187) [Introduction to open source FiftyOne](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#0c3652242d88) [What’s new in the latest version (v.17) of FiftyOne](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#b9a02ab692da) [FiftyOne Open Source — Live Demo Time!](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#8c2e5ac6eb97) [FiftyOne Teams — Live Demo Time!](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#479b5339a68f) [Q&A from the Webinar](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#caca906cebc6) [What’s Next](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams#86ec1360ef5c) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/17422a76c76f14096dce21e43da51945f498a811-1200x673.png?auto=format&dpr=2&fit=max&q=75&w=1200) _Editor’s note: FiftyOne Teams is now [FiftyOne Enterprise](https://voxel51.com/enterprise/)._ We recently announced the availability of [FiftyOne .17](https://medium.com/voxel51/announcing-fiftyone-0-17-0-with-grouped-datasets-3d-geolocation-and-custom-plugins-339600ab73a1) and [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) for collaborating securely on datasets. Voxel51 Co-Founder and CTO [Brian Moore](https://www.linkedin.com/in/brimoor/) walked us through all the new features in a live webinar. You can watch the playback on [YouTube](https://www.youtube.com/watch?v=QDgsXTJ7YZA), take a look at the [slides](https://docs.google.com/presentation/d/1APm9X41K_929Glq0ZRej4bOwIgTLwl0IWwj52eozBzM/edit?_hsmi=2&_hsenc=p2ANqtz-80cSfPm2m6lo6dsckAI4bXD70q9IiOUKNJB0ABWDoGWWdOm1AwAgbh5FDUFLmz1mvrjaUQE4z1ObtWd1mpVo_0SgBwWw#slide=id.g15ecaa5ed7e_1_0), see the full [transcript](https://www.rev.com/tc-editor/shared/ZWIc8nJsk5eOUEBmw_UW1b9kbV-zWk7ZN5VuczmZlRgyp-ZrPijsP9c5XCm1Zk_1g7yBMMaYg3f-vMcvrsYXYAej-qM?loadFrom=SharedLink), and read the recap below for the highlights. https://www.youtube.com/watch?v=QDgsXTJ7YZA ## Donating $200 to World Literacy Foundation In lieu of swag, we gave attendees the opportunity to vote for their favorite charity and help guide our monthly donation to charitable causes. The charity that received the highest number of votes was the [World Literacy Foundation](https://worldliteracyfoundation.org/). We are pleased to be making a donation of $200 to them on behalf of the FiftyOne community. ![](https://cdn.sanity.io/images/h6toihm1/production/6da3098b8f6eeb34e3674a5984affbc8a663de8e-300x300.png?auto=format&dpr=2&fit=max&q=75&w=300) ## Presentation Highlights To set the scene, Brian took us on a journey starting 10 years ago, at a time when you could work with a dataset manually because the dataset size was relatively small. But fast forward to now, datasets are much bigger and span more modalities — images, video, 3d, and more. While you might think working on models is what takes up a lot of an ML engineer’s time, it’s really the data quality. If you have poor quality data, it leads to problems — model bias, physical danger, reduced model performance. Our company, [Voxel51](https://voxel51.com/), is on a mission to bring transparency and clarity to the world’s data. We focus on building tools to help improve data quality and data-centric workflows. We do that through the [open source FiftyOne project](https://github.com/voxel51/fiftyone). ## Introduction to open source FiftyOne Brian described what you can do with FiftyOne — the open-source tool for building high-quality datasets and computer vision models — including all the workflows it supports, all the computer vision tasks it supports, and all the integrations, too: **FiftyOne helps you with these workflows and dozens more:** - Curate, visualize, and analyze datasets - Streamline annotation workflows - Find and fix labeling mistakes - Identify and correct model failures **FiftyOne supports all popular computer vision tasks:** - Classification - Detection - Instance segmentation - Semantic segmentation - Polygons and polylines - Keypoints - Point clouds and annotations - Geolocation - Embeddings - Multiview datasets - Image, video, and 3D data **FiftyOne integrates with all your favorite ML tools:** \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop **FiftyOne comes with the FiftyOne Brain:** In the [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html) you’ll find a bunch of interesting data-centric workflows designed to go beyond just the straight visualization and query capabilities of the tool, but really go into the next level where you’re trying to identify specific insights into the data sets, like automatically finding potential mistakes that your model is making, automatically computing embeddings, providing your own custom embeddings, visualizing low dimensional representations of the embeddings, and interacting with them to pull out and observe patterns in your data. ## What’s new in the latest version (v.17) of FiftyOne We recently released [FiftyOne .17](https://medium.com/voxel51/announcing-fiftyone-0-17-0-with-grouped-datasets-3d-geolocation-and-custom-plugins-339600ab73a1) and there were a number of new goodies that were added based on community input, including: - 3D + Multiview datasets [(docs)](https://voxel51.com/docs/fiftyone/user_guide/groups.html) - Geolocation [(docs)](https://voxel51.com/docs/fiftyone/user_guide/app.html#map-tab) - Custom plugins [(docs)](https://github.com/voxel51/fiftyone/blob/develop/app/packages/plugins/README.md) ## FiftyOne Open Source — Live Demo Time! Brian covered all the awesomeness available in open source FiftyOne from [~14:34 to ~40:20](https://www.youtube.com/watch?v=QDgsXTJ7YZA#t=14m34s) in the presentation, including these how-to’s: - Install FiftyOne in a Python terminal window with pip install fiftyone and import the library - Connect to a database (a MongoDB database automatically spins up when you import the library, although you can configure it to connect to an existing database if you’d like) - Make a dataset persistent (datasets are ephemeral by default) - Load a dataset (to show this Brian loads CIFAR-100, then later loads KITTI, COCO, MNIST, and others) - Visualize a dataset in the FiftyOne App - Interactively work with data across the App and code - Filter samples in the App, whether they are images, videos, 3D, location data, etc. - Use the FiftyOne Brain to search for visually similar images or find annotation mistakes (and much more!) - Work with FiftyOne in a Jupyter Notebook - Load in datasets packaged in the Dataset Zoo - Generate scatter plots on samples - Use models out of the Model Zoo, like MobileNet - And so much more! A lot of what Brian demoed is from the [Images Embeddings tutorial](https://voxel51.com/docs/fiftyone/tutorials/image_embeddings.html) if you’d like to have a look. Plus, everything Brian showed until now is all part of open source FiftyOne, and you can find everything you need to get started on [GitHub](https://github.com/voxel51/fiftyone). If you like what you see, consider giving the project a star. ## FiftyOne Teams — Live Demo Time! Next, Brian dove into [FiftyOne Teams](https://voxel51.com/fiftyone-teams/), to enable multiple users to securely collaborate on the same datasets and models, either on-premises or in the cloud, all built on top of the open source FiftyOne. **FiftyOne Teams includes:** - Native cloud storage - Centralized web portal with SSO - Dataset/user permissions - Dataset versioning - Dataset listing and management - Custom dashboards Brian covered all the enterprise features in FiftyOne Teams starting at [~45:31 in the presentation](https://www.youtube.com/watch?v=QDgsXTJ7YZA#t=45m31s), including how to: - Log into FiftyOne Teams - Manage access to datasets - Manage user roles — admins, members, guests - Pin datasets - Find datasets based on a variety of filters - Install, connect to, and work with datasets through Python - More! ## Q&A from the Webinar There was a lively Q&A all throughout the presentation, covering open source FiftyOne, FiftyOne Teams, and more! Here’s a recap: **Do the FiftyOne features work with video datasets?** Yes! FiftyOne fully supports video datasets just like image datasets and even allows you to work with clips or even individual frames in your video datasets. Learn more about [working with video datasets](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#video-views) in the FiftyOne Docs. **How does the session object work with FiftyOne Teams when multiple users have multiple sessions in the app?** The hosted FiftyOne Teams App doesn’t currently support connection to Python sessions (but stay tuned!) Instead, Teams users can achieve interactive Python sessions by launching the localhost App via the Python SDK, just like Brian showed in the open source demo allowing users to interact through code. Learn more about [FiftyOne Teams.](https://voxel51.com/fiftyone/) **Can I work with embeddings directly in the App like I can with geolocation through the new Map tab?** Currently, embeddings work is done through Jupyter notebooks with plotly plots, but we plan to add support for interactive embeddings workflows directly in the App (like geolocation maps) in the near future! Check out this tutorial for how to [interact with embeddings](https://voxel51.com/docs/fiftyone/user_guide/plots.html#visualizing-embeddings) and the FiftyOne App. **Can you share the notebook used in the image embeddings demo?** The demo is based largely on [this tutorial](https://voxel51.com/docs/fiftyone/tutorials/image_embeddings.html). For background, FiftyOne provides a powerful [embeddings visualization](https://voxel51.com/docs/fiftyone/user_guide/brain.html#visualizing-embeddings) capability that you can use to generate low-dimensional representations of the samples and objects in your datasets. [This notebook](https://gitcdn.link/cdn/voxel51/fiftyone/v0.17.2/docs/source/tutorials/image_embeddings.ipynb) highlights several applications of visualizing image embeddings, with the goal of motivating some of the many possible workflows that you can perform. **Are we able to use geolocations in the to\_patches() view as well?** Yep! If you store `fo.GeoLocation` labels on your `fo.Detection` labels, then when you convert `to_patches()` you’ll be able to interact with the map in the patches view. There is one more step required in between, but you can check out this [GitHub gist](https://gist.github.com/ehofesmann/41481dd0b91e0b1a42435b8a07a76cf9) for an example of how to do this. **To what extent will the FiftyOne project’s features continue to remain open source?** FiftyOne will always be free open source for individual users to install locally. For teams of users who need to collaborate on their datasets, we offer commercial [FiftyOne Teams](https://voxel51.com/fiftyone-teams/). **Can the FiftyOne Brain be used for finding incorrect detection annotations?** Yes! We have a mistakenness computation in the brain to do just that. Check out the docs to learn more about [label mistakes](https://voxel51.com/docs/fiftyone/user_guide/brain.html#label-mistakes). **Does FiftyOne currently support 3D mesh data?** In 0.17, we added support for [3D point clouds](https://voxel51.com/docs/fiftyone/user_guide/groups.html#point-cloud-slices) stored in .pcd format. We’re also discussing adding support for .ply and .obj files. Alternatively, you can always write a [custom plugin](https://github.com/voxel51/fiftyone/blob/develop/app/packages/plugins/README.mdhttps://github.com/voxel51/fiftyone/blob/develop/app/packages/plugins/README.md) that supports visualizing other 3D formats like meshes. **How can we analyze statistical metrics with FiftyOne, like accuracy, IoU, etc?** FiftyOne [natively supports](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#detections) evaluation of regressions, classifications, detections, segmentations, and temporal detections. Specifically, for each task, FiftyOne implements best practice protocols like COCO for object detection. (Note that FiftyOne is [the recommended way](https://cocodataset.org/#download) to evaluate COCO datasets on the COCO website!) Uniquely, FiftyOne’s evaluation methods store individual TP/FP/FN results on your data, as well as aggregate metrics like accuracy/mAP so you can really dig in and see how individual samples/predictions performed. **Can I connect my trained model to a dataset management/version control system?** [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) comes with dataset versioning builtin so you can track changes, tag revisions, and rollback to previous dataset versions as needed. You can use this feature in concert with experiment tracking tools like Weights & Biases or MLflow to record the specific dataset revision on which a model/experiment was trained Make sure to [follow Voxel51 on Linkedin](https://www.linkedin.com/company/voxel51) for an upcoming blog post showing best practices for using FiftyOne together with your favorite experiment tracking tools! **What are some interesting non-machine learning use cases for FiftyOne that you’ve seen?** We are constantly amazed by the variety of applications that FiftyOne users showcase in our [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ). We’ve recently seen datasets including agriculture, satellite imagery, robotics, retail, healthcare, and more. Who knows, maybe somebody is using the new [custom plugins feature](https://github.com/voxel51/fiftyone/blob/develop/app/packages/plugins/README.mdhttps://github.com/voxel51/fiftyone/blob/develop/app/packages/plugins/README.md) as we speak to do something truly creative, like say … a photo editing plugin! **How does FiftyOne Teams differentiate itself from products like Scale AI and Labelbox?** FiftyOne Teams is built on top of open source FiftyOne, which offers a number of key benefits to users: - Lower barrier to entry: FiftyOne is 100% free, forever, with unlimited data volumes - FiftyOne Teams is backwards-compatible with OSS FiftyOne - The FiftyOne ecosystem is purpose built with flexibility and extensibility in-mind - No vendor lock-in for annotation - Significantly more full-featured Python library for building data-centric workflows in your ML environment ## What’s Next - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone)! - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help! [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne 0.17](https://voxel51.com/blog/tag/fiftyone-0-17) [FiftyOne Teams](https://voxel51.com/blog/tag/fiftyone-teams) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/308698a5aece1d5b1b95ee1bf52811b24448458c-1200x672.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne 0.18 for Computer Vision\\ \\ Event Recaps\\ \\ • \\ \\ Dec 6, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/2be41b07bd86d7efc5916442ca5b94aa5115234d-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Greatest Hits of 2022: FiftyOne & Voxel51\\ \\ Product & News\\ \\ • \\ \\ Jan 10, 2023](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51) [![](https://cdn.sanity.io/images/h6toihm1/production/a625547e8b6b712e9a9bd5b1300cf9686c11cf20-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Meetup Recap: How to Build High-Quality Machine Learning Datasets and Computer Vision Models\\ \\ Event Recaps\\ \\ • \\ \\ Jul 21, 2022](https://voxel51.com/blog/meetup-recap-how-to-build-high-quality-machine-learning-datasets-and-computer-vision-models) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-213-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Oct 7, 2022 Oct 8, 2022 • 3 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#d0bbbbf4b8c1) [Importing previously exported FiftyOne datasets](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#9cd5d6a732c4) [Support for TIFF images in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#72dde34e8cac) [MSCOCO analysis on FiftyOne datasets](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#93a7083b8a8d) [Working with absolute and relative paths](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#e32c9ef93571) [Importing and exporting custom datasets](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#5bade32bda56) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#267dc010301b) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#d0bbbbf4b8c1) [Importing previously exported FiftyOne datasets](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#9cd5d6a732c4) [Support for TIFF images in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#72dde34e8cac) [MSCOCO analysis on FiftyOne datasets](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#93a7083b8a8d) [Working with absolute and relative paths](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#e32c9ef93571) [Importing and exporting custom datasets](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#5bade32bda56) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-7-2022#267dc010301b) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/4ec33331c3696c67ff0c56acd1bf45adf9416980-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=1200) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help Ok, let’s dive into this week’s tips and tricks! ## **Importing previously exported FiftyOne datasets** Community Slack member Dan Erez asked, _“How can I import previously exported FiftyOne datasets?”_ The basic syntax is: import fiftyone as fo \# The directory containing the dataset to import dataset\_dir = "/path/to/dataset" \# The type of the dataset being imported dataset\_type = fo.types.COCODetectionDataset # for example \# Import the dataset dataset = fo.Dataset.from\_dir( dataset\_dir=dataset\_dir, dataset\_type=dataset\_type, ) Learn more about [loading datasets from disk](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/datasets.html) in the FiftyOne Docs. ## **Support for TIFF images in the FiftyOne App** Stackoverflow member Joanna B asked, _“Can I load TIFF images using FiftyOne and Python in .ipynb notebook?”_ The image types that FiftyOne supports are based on the underlying support of the browser you are using to run the FiftyOne App. The TIFF format is only supported out of the box by a few browsers [like Safari.](https://en.wikipedia.org/wiki/Comparison_of_web_browsers#Image_format_support) For other browsers, you will need to check if there is an extension that adds TIFF support, like this one for [Firefox.](https://addons.mozilla.org/en-US/firefox/addon/tiff-viewer/) Learn more about the [image types supported in FiftyOne](https://voxel51.com/docs/fiftyone/faq/index.html#what-image-file-types-are-supported) in the FiftyOne Docs. ## **MSCOCO analysis on FiftyOne datasets** Community Slack member Sidney Guaro asked, _“Does Fiftyone have a MSCOCO analysis implementation that include things like localization error, background error, and similarity?”_ Yes! By default, `evaluate_detections()` will use COCO-style evaluation to analyze predictions when the specified label fields are `Detections` or `Polylines`. Although FiftyOne doesn’t directly expose a function that computes these metrics, you can construct dataset views that will allow you to compute many of these cases. For example [stratifying by bounding box area](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#filtering-detections-by-area). Another option is to export the relevant labels in COCO format and [pass them to pycocotools](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html#cocodetectiondataset). Learn more about [COCO-style evaluation](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#coco-style-evaluation-default-spatial) in the FiftyOne Docs. ## **Working with absolute and relative paths** Community Slack member Teemu Sormuen asked, _“I can’t seem to reference images correctly due to issues with the root path. Currently, it seems that `Sample.filepath` should be an absolute path, instead of a relative path with regards to the current working directory. I would like to use the current working directory as a relative base path for `Sample.filepath`. Is there some way to define the relative path to which `Sample.filepath` is compared to?”_ The best way to do that is to modify the filepaths based on your current working directory. However, you can also use [`dataset.set_field()`](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#transforming-fields) to create a temporary view that updates filepaths using a `ViewExpression` rather than changing them permanently with `set_values()` each time. For example: from fiftyone import ViewField as F from fiftyone import ViewExpression as E view = dataset.set\_field( "filepath", E("/path/to/dir/").concat(F("filepath").split("/")\[-1\]), ) Learn more about [ViewExpressions](https://voxel51.com/docs/fiftyone/api/fiftyone.core.expressions.html#fiftyone.core.expressions.ViewExpression) in the FiftyOne Docs. ## **Importing and exporting custom datasets** Community Slack member Sagar Kalburgi asked, _“Can FiftyOne import a dataset that has been labeled using the [LabelMe tool](http://labelme.csail.mit.edu/Release3.0/)? Furthermore, is it possible to export the dataset in a different format after importing?”_ Yes! You can always import custom labeled data into FiftyOne through a simple [Python loop](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html#custom-formats) (here’s how to load custom label types like [detections](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#instance-segmentations) and [instance segmentations](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#object-detection)), or by writing a [custom dataset importer](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/datasets.html#custom-formats). Once your data is loaded, you can then export it over two dozen formats. Learn more about [exporting datasets in different formats](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html) in the FiftyOne Docs. ## **What’s next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help [custom datasets](https://voxel51.com/blog/tag/custom-datasets) [FAQ](https://voxel51.com/blog/tag/faq) [TIFF images](https://voxel51.com/blog/tag/tiff-images) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/4d37703d72d4b83a85bda19eb1999d5247915fc0-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 02, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 3, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/204bd847400b1534c494dae594f0797457bc4c90-1620x906.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Filtering Tips and Tricks — Dec 09, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 10, 2022](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/f0fe111e79e719eeb9100d13801921e34a778043-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Aggregation Tips and Tricks — Nov 25, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 26, 2022](https://voxel51.com/blog/fiftyone-aggregation-tips-and-tricks-nov-25-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-214-lllmstxt|> ## Funding Announcement [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Announcing Our $12.5M Series A Funding to Bring Transparency and Clarity to the World’s Data Sep 22, 2022 • 3 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/1f1e14f53157ef2324e17e6e2978722f4aa0be66-1200x674.png?auto=format&dpr=2&fit=max&q=75&w=1200) I’m delighted to announce that we raised $12.5 million in Series A funding from new investors Drive Capital, Top Harvest Capital, and Shasta Ventures as well as existing investors eLab Ventures and ID Ventures. This financing allows us to accelerate the next phase of our growth in bringing data-centric machine learning to the world. Since we started [Voxel51](https://voxel51.com/) in October 2018, we’ve been building open source and commercial software that enables developers, scientists, and organizations to build high-quality datasets and computer vision models that are powering some of today’s most remarkable machine learning and artificial intelligence. Our software provides the infrastructure for these users to analyze and modulate their datasets allowing them to address critical issues like data bias. One of our first big milestones was in August 2020 with the launch of the [open source FiftyOne project](https://github.com/voxel51/fiftyone). Since then, FiftyOne has seen incredible growth and adoption. Today FiftyOne is used by tens of thousands of engineers and scientists and has reached 150,000+ monthly active machines, 1900+ GitHub stars, and 1000+ members in our Slack community. We’re grateful for everyone in the enthusiastic and growing FiftyOne community — for your support, contributions, and for being a part of this journey with us! Late last year, we began working with dozens of startups and Fortune 500 enterprises as early adopters of FiftyOne Teams, our commercial product that enables teams to securely collaborate on their datasets and models. Our early adopters span a variety of industries — automotive, robotics, security, retail, healthcare, and more — including large organizations, which are typically risk averse; a testament to the utility and value that FiftyOne Teams provides! By providing us with real-world usage, input, and feedback, our early adopters helped us shape and harden the [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) solution that we’re publicly announcing today. (You can find the full press release [here](https://www.prnewswire.com/news-releases/voxel51-raises-12-5m-series-a-to-bring-transparency-and-clarity-to-computer-vision-data-301629679.html).) A huge shoutout and thank you to all of our early adopters for making FiftyOne Teams ready for the broader ML/AI community to enjoy! So… what’s next for Voxel51? It’s no secret that there has been an explosion of visual data. For example, there are an estimated [45 billion cameras](https://www.ldv.co/blog/2017/8/8/45-billion-cameras-by-2022-fuel-business-opportunities) in the world today, and the growth will continue to accelerate for decades. This creates a tremendous opportunity for computer vision applications, but only if the data can be properly organized, indexed, and labeled. As Voxel51 enters its next phase of growth, we’re doubling down on our commitment to build FiftyOne alongside the community so that it remains the leading open source tool for building high-quality datasets and computer vision models. We’re also accelerating the development of FiftyOne Teams as critical and trusted infrastructure for managing organization’s visual data, enabling them to build machine learning systems based on high quality, transparent data that brings their ML-powered products to market faster. We’re going to need many more exceptional and diverse people to achieve our mission of bringing transparency and clarity to the world’s data. If our mission excites you, check out our [open positions across product, engineering, community, and more](https://voxel51.com/jobs/). Thank you to everyone who has supported Voxel51 over the years — our amazing team, investors, customers, partners, and open source community. Today is a significant milestone on our journey and I’m excited for many more to come! _We presented a webinar with interactive demos of FiftyOne v.17 and FiftyOne Teams. **[Catch the recap](https://medium.com/voxel51/webinar-recap-whats-new-in-fiftyone-fiftyone-teams-c2c0425c8f91).**_ [FiftyOne Teams](https://voxel51.com/blog/tag/fiftyone-teams) [funding](https://voxel51.com/blog/tag/funding) [Series A](https://voxel51.com/blog/tag/series-a) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/61b9a72c1362d209b3cb768c4ddc7f84cdf22452-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Automatically Set Up a New ML Project, Pain Free\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Feb 8, 2023](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free) [![](https://cdn.sanity.io/images/h6toihm1/production/ab58c64891ed83a26cfa9c69c8f801bbb5a2675c-3840x2160.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Announcing Voxel51’s Series B Led By Bessemer Venture Partners\\ \\ Product & News\\ \\ • \\ \\ May 16, 2024](https://voxel51.com/blog/announcing-series-b-led-by-bessemer-venture-partners) [![](https://cdn.sanity.io/images/h6toihm1/production/749a0f8c0da5f5f3ca4f8d8a6b6bf3a216551023-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Announcing FiftyOne 0.23.3 and FiftyOne Teams 1.5.4\\ \\ Product & News\\ \\ • \\ \\ Jan 24, 2024](https://voxel51.com/blog/fiftyone-0-23-3-and-fiftyone-teams-1-5-4) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-215-lllmstxt|> ## FiftyOne 0.17 Release [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Announcing FiftyOne 0.17 with Grouped Datasets, 3D, Geolocation, and Custom Plugins Sep 21, 2022 • 4 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#4134a7f6016c) [What’s New in 0.17?](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#198104ffb6fe) [Grouped Datasets: Working with Multiview Data, Including 3D!](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#83309d305b4e) [Visualize and Interact with Geolocation Data](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#63ce328fc95b) [Custom App Plugins](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#df3c6b52568d) [Training Detectron2 Models in FiftyOne](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#2ac71711bbdd) [Community Contributions](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#890eda705ab6) [FiftyOne Community Updates](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#8b0385262644) [See FiftyOne 0.17 in Action!](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#b0baac9942bd) [What’s Next?](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#c4dc9ec32947) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#4134a7f6016c) [What’s New in 0.17?](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#198104ffb6fe) [Grouped Datasets: Working with Multiview Data, Including 3D!](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#83309d305b4e) [Visualize and Interact with Geolocation Data](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#63ce328fc95b) [Custom App Plugins](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#df3c6b52568d) [Training Detectron2 Models in FiftyOne](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#2ac71711bbdd) [Community Contributions](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#890eda705ab6) [FiftyOne Community Updates](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#8b0385262644) [See FiftyOne 0.17 in Action!](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#b0baac9942bd) [What’s Next?](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins#c4dc9ec32947) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Voxel51, in conjunction with the FiftyOne community, are excited to announce the release of FiftyOne 0.17! ## **Wait, What’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into the release! ## **What’s New in 0.17?** This release includes enhancements and fixes to the FiftyOne App, core library, annotation integrations, three new datasets/models in the FiftyOne Zoo, and updated documentation. In total, there are 16 new features and 6 bug fixes. You can check out all the details in the official [release notes](https://voxel51.com/docs/fiftyone/release-notes.html). Here’s a quick tl;dr to highlight some of the new features in this release. ## **Grouped Datasets: Working with Multiview Data, Including 3D!** FiftyOne now supports the creation of grouped datasets, which contain multiple slices of samples of possibly different modalities (image, video, or point cloud) that are organized into groups. Grouped datasets can be used to represent multiview scenes, where data for multiple perspectives of the same scene can be stored, visualized, and queried in ways that respect the relationships between the slices of data. Grouped datasets may contain 3D samples, including point cloud data stored in .pcd format and associated 3D annotations (detections and polylines). As expected, you can work with grouped datasets in the FiftyOne App! ![](https://cdn.sanity.io/images/h6toihm1/production/ee65c9e4f168d7ce673205036c9e5a22fde08017-2560x1370.gif?auto=format&dpr=2&fit=max&q=75&w=1600) In the FiftyOne App you can perform a variety of operations with grouped datasets: - View all samples in the current group in the modal - Samples can include image, video, and point cloud slices - Browse images and videos in a scrollable carousel and maximize them in the visualizer - For point cloud slices, you can make use of a new interactive 3D visualizer - View statistics across all slices Check out [the docs](https://voxel51.com/docs/fiftyone/user_guide/groups.html) to learn how to get started with grouped datasets and interact with them in the FiftyOne App. ## **Visualize and Interact with Geolocation Data** The FiftyOne App has a new Map tab that appears whenever your dataset has a `GeoLocation` field with `point` data populated. You can use the Map tab to see a scatterplot of your dataset’s location data: import fiftyone as fo import fiftyone.zoo as foz dataset = foz.load\_zoo\_dataset("quickstart-geo") session = fo.launch\_app(dataset) ![](https://cdn.sanity.io/images/h6toihm1/production/cbd91b951db70788c004b58581d79ac894544380-2292x2052.gif?auto=format&dpr=2&fit=max&q=75&w=1600) What can you do with the Map tab? You can: - Lasso points in the map to show the corresponding samples in the grid - Choose between available map types (dark, light, satellite, road, etc.) - Configure your own custom default settings for the Map tab Check out the [Map tab docs](https://voxel51.com/docs/fiftyone/user_guide/app.html#app-map-tab) to learn how to get started with visualizing and interacting with geolocation data. ## **Custom App Plugins** FiftyOne now supports a plugin system that you can use to customize and extend the App’s behavior! For example if you need a unique way to visualize individual samples, plot entire datasets, or fetch FiftyOne data, a custom plugin just might be the ticket! ![](https://cdn.sanity.io/images/h6toihm1/production/1f1bea9c5dad46794c819e54e2b8dd2ee63327d0-1400x366.png?auto=format&dpr=2&fit=max&q=75&w=1400) Check out [this tutorial](https://github.com/voxel51/fiftyone/blob/develop/app/packages/plugins/README.md) on GitHub that walks you through how to develop and publish a custom plugin. ## **Training Detectron2 Models in FiftyOne** [Detectron2](https://github.com/facebookresearch/detectron2) is Facebook AI Research’s next generation library that provides state-of-the-art detection and segmentation algorithms. It supports a number of computer vision research projects and production applications in Facebook. New in this release is [a tutorial](https://voxel51.com/docs/fiftyone/tutorials/detectron2.html) that shows how, with two simple functions, you can integrate FiftyOne into your Detectron2 model training and inference pipelines. ![](https://cdn.sanity.io/images/h6toihm1/production/e94f20fa81716294c7a6caccf7256e9106cb6e89-967x800.png?auto=format&dpr=2&fit=max&q=75&w=967) ## **Community Contributions** We’d like to take a moment to give a few shout outs to FiftyOne community members who contributed to this release. **OpenAI’s CLIP Model Now in the FiftyOne Model Zoo** [Rustem Galiullin](https://github.com/Rusteam) contributed PRs [#1691](https://github.com/voxel51/fiftyone/pull/1691) and [#2072](https://github.com/voxel51/fiftyone/pull/2072), which added a CLIP ViT-Base-32 model to the [FiftyOne Model Zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html) for zero-shot classification and embedding generation. The CLIP model was [announced by](https://openai.com/blog/clip/) researchers at OpenAI in 2021 and is a breakthrough in efficiently learning visual concepts from natural language supervision. The model can be used in FiftyOne, for example, to classify images according to an arbitrary set of classes in a zero-shot manner: ![](https://cdn.sanity.io/images/h6toihm1/production/0f42e52597a2db2a318bb1afd1fdf0c8d846f0cd-1999x1475.png?auto=format&dpr=2&fit=max&q=75&w=1600) **Additional Community Contributions** Shoutout to the following community members who contributed the following PRs to the FiftyOne project over the past few weeks! - [George Pearse](https://github.com/GeorgePearse) contributed [#2068 — Update install.rst](https://github.com/voxel51/fiftyone/pull/2068) - [Odd Eirik Igland](https://github.com/oddeirikigland) contributed [#2066 — Bugfix: task\_map expected string got int](https://github.com/voxel51/fiftyone/pull/2066) - [Victor1cea](https://github.com/victor1cea) contributed [#2016 — Fix Issue #1903 path variable](https://github.com/voxel51/fiftyone/pull/2016) and [#1884 — Eliminate non-XML or non-TXT files from CVAT, KITTI, CVAT Video](https://github.com/voxel51/fiftyone/pull/1884) - [Geoffrey Keating](https://github.com/geoffrp) contributed [#1973 — CVAT Annotate attribute documentation update](https://github.com/voxel51/fiftyone/pull/1973) - [Idow09](https://github.com/idow09) contributed [#1909 — Fix custom\_parser implementation in recipe](https://github.com/voxel51/fiftyone/pull/1909) ## **FiftyOne Community Updates** The FiftyOne community continues to grow! - 1,000+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 1,900+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 1,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 166 repositories - 36 [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## See FiftyOne 0.17 in Action! Join me (Voxel51 Co-Founder and CTO) for a live webinar, where I’ll give an interactive demo of FiftyOne 0.17 and introduce our new FiftyOne Teams offering. [Sign up here](https://us02web.zoom.us/webinar/register/WN_nXb9aYN2S8KHYXu7oC4OVg)! ## **What’s Next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone)! - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help! [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne 0.17](https://voxel51.com/blog/tag/fiftyone-0-17) [product release](https://voxel51.com/blog/tag/product-release) ![](https://cdn.sanity.io/images/h6toihm1/production/8d61ff90b31d151405f9e21a33c2802509f34651-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Brian Moore Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a4c2bee9ed053c5be2a1c161e5abf758c9a12ff8-1400x923.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.18 with App Performance Improvements, Sidebar Modes, and Custom Attributes\\ \\ Product & News\\ \\ • \\ \\ Nov 15, 2022](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes) [![](https://cdn.sanity.io/images/h6toihm1/production/ed0b9dc4072cfa3d1d1bd2837c0b834fccec607f-1400x775.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Introducing FiftyOne: A Tool for Rapid Data & Model Experimentation\\ \\ Product & News\\ \\ • \\ \\ Sep 12, 2020](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation) [![](https://cdn.sanity.io/images/h6toihm1/production/8651d28f4a2978eb72cca25ef09e1f01f81847ea-2968x2042.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.19 with Spaces, In-App Embeddings Visualization, Saved Views, and More!\\ \\ Product & News\\ \\ • \\ \\ Feb 16, 2023](https://voxel51.com/blog/announcing-fiftyone-0-19) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-216-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Sept 16, 2022 Sep 17, 2022 • 3 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#6d3c808ca4f6) [Checking and filtering for valid annotations](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#be9a31b6ddc4) [Loading CVAT projects into FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#72d0c752b79c) [Selecting multiple samples in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#e0497ed84e45) [Viewing field information at the frame level](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#c2fa17e42a1e) [Working with object-level embeddings](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#7a4c3dd71fd5) [What’s Next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#5679563ebbb1) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#6d3c808ca4f6) [Checking and filtering for valid annotations](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#be9a31b6ddc4) [Loading CVAT projects into FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#72d0c752b79c) [Selecting multiple samples in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#e0497ed84e45) [Viewing field information at the frame level](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#c2fa17e42a1e) [Working with object-level embeddings](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#7a4c3dd71fd5) [What’s Next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022#5679563ebbb1) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/03107d477b7db4be03031293fa4fe15aaea806f0-1200x677.png?auto=format&dpr=2&fit=max&q=75&w=1200) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/docs/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## Checking and filtering for valid annotations Community Slack member Yicheng He asked, **_“Is there a way to check for (and possibly filter) valid annotations when loading COCO-styled object detection data into FiftyOne? Some examples include checking if the bounding box annotation is non-empty and if the coordinates are valid.”_** One approach is to write your own custom dataset importer that is able to perform whatever validation you need when loading labels into FiftyOne in the `COCODetectionDataset` format, rather than from one of our integrations. Alternatively, you could filter your dataset after the fact to perform this verification: \# Only omits None-valued samples view = dataset.exists("ground\_truth") \# Also omits samples with empty Detections(\[\]) view = dataset.match(F("ground\_truth.detections").length() > 0) Learn more about writing your own [custom dataset importer](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/datasets.html#custom-formats) and [filtering your dataset](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#filtering) in the FiftyOne Docs. ## Loading CVAT projects into FiftyOne Stackoverflow member MosQuan asked, **_“Can I load a previously created CVAT project from a server? If so, what arguments should I specify besides the CVAT project name?”_** Yes, you can use the `fiftyone.utils.cvat.import_annotations()` method to import labels that are already in a CVAT project or task into a FiftyOne Dataset. Note that in order to use `fo.load_dataset()`, the dataset needs to already exist in FiftyOne. You can initialize an empty dataset as shown below: dataset = fo.Dataset("my-dataset-name") Then, you can call `import_annotations()`, providing a project name and optionally a `data_path` and `export_media=True` to download all of the media from your project to a local directory as well as all of the labels in your project, then import them into the dataset you just created. dataset = fo.Dataset("my-dataset-name") fouc.import\_annotations( dataset, project\_name=project\_name, data\_path="/tmp/cvat\_import", download\_media=True, ) If your media already exists on disk, then check out [this example](https://voxel51.com/docs/fiftyone/integrations/cvat.html#importing-existing-tasks) for how to provide a `data_map` mapping the CVAT filename to filepath of the media on the local disk. Learn more about [CVAT integration](https://voxel51.com/docs/fiftyone/integrations/cvat.html?highlight=fiftyone%20utils%20cvat%20import_annotations) and the [`fiftyone.utils.cvat.import_annotations()`](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.cvat.html?highlight=fiftyone%20utils%20cvat%20import_annotations#fiftyone.utils.cvat.import_annotations) function in the FiftyOne Docs. ## Selecting multiple samples in the FiftyOne App Community Slack member Nadav Ben-Haim asked, **_“Is there a way to batch select lots of videos? In the FiftyOne App I am only able to click on the check box in the top left to select a video, but there doesn’t appear to be a select with some kind of drag & drop capability.”_** Yes! You can use _shift_ and _ctrl_ to select ranges of samples in the grid view of the FiftyOne App. Click to select one sample, scroll down, hold _shift_ and click another sample, then all samples in between those two will be selected. ## Viewing field information at the frame level Community Slack member Eric Ng asked, **_“Is there a way to view the field information at the frame level? Currently when I add a field at the frame level it shows up in the FiftyOne Apps OTHER tab as `frames.field_name`, but doesn’t show the value.”_** One way to see frame-level field values in the FiftyOne App is to open the JSON viewer by pressing _j_ or clicking the _{…}_ button in the menu at the bottom. If the value is stored as a `Classification` instance, it will be overlaid on the video itself like this: ![](https://cdn.sanity.io/images/h6toihm1/production/02944f4e75c6d403992ba5936e9f55f7b9b5755f-1268x714.png?auto=format&dpr=2&fit=max&q=75&w=1268) ## Working with object-level embeddings Community Slack member Yashovardhan Chaturvedi asked, **_“Can I store object-level embeddings on the `Detection` instances in my dataset?”_** Yes! The FiftyOne Brain component has support for using object-level embeddings for various purposes. To learn more about [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html?highlight=brain) and an [object embeddings example](https://voxel51.com/docs/fiftyone/user_guide/brain.html#object-embeddings-example), check out the FiftyOne Docs. ## What’s Next? - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [CVAT](https://voxel51.com/blog/tag/cvat) [embeddings](https://voxel51.com/blog/tag/embeddings) [FAQ](https://voxel51.com/blog/tag/faq) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a3a918e30b0553723b9392ea90763379f98480a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – April 7, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Apr 7, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/5eca7f455938fa6e0b4755f60529b4123ae91046-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – May 26, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ May 26, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/47c4462ab8c14fe4728ce5901402986125d3f99d-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Nov 18, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 19, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-217-lllmstxt|> ## FiftyOne Vision Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Sept 23, 2022 Sep 24, 2022 • 3 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#77967260074b) [Filtering for media that doesn’t contain a tag](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#fb51591e28c4) [Visualizing embeddings in FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#13b75a193013) [Exporting labeled images](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#b20d81add5a2) [Grouping multiple slices of image, video or point cloud samples](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#3cd1f90fd6b2) [Creating views for samples that contain just the bounding box](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#4bea92d0880e) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#3d510d7d4340) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#77967260074b) [Filtering for media that doesn’t contain a tag](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#fb51591e28c4) [Visualizing embeddings in FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#13b75a193013) [Exporting labeled images](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#b20d81add5a2) [Grouping multiple slices of image, video or point cloud samples](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#3cd1f90fd6b2) [Creating views for samples that contain just the bounding box](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#4bea92d0880e) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-23-2022#3d510d7d4340) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/25a70e1c9c2cae7f9df18696e0e087dddf67707d-1200x679.png?auto=format&dpr=2&fit=max&q=75&w=1200) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/docs/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help Ok, let’s dive into this week’s tips and tricks! ## **Filtering for media that doesn’t contain a tag** Community Slack member Nadav Ben-Haim asked, **_“Is there an easy way to filter media in the FiftyOne App for all the media that does NOT contain a tag?”_** Indeed there is! In the [_view_ bar](https://voxel51.com/docs/fiftyone/user_guide/app.html#using-the-view-bar) use `match_tags()` where `bool=False` ![](https://cdn.sanity.io/images/h6toihm1/production/8a9bbbeee4f976a0d5b73219c15e38d3bc63921d-904x94.png?auto=format&dpr=2&fit=max&q=75&w=904) ## **Visualizing embeddings in FiftyOne** Community Slack member Suphanut Jamonnak asked, **_“Is it possible to visualize embeddings using the FiftyOne App?”_** The [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html) component can help you visualize embeddings. So, instead of combing through individual images/videos and staring at aggregate performance metrics trying to figure out how to improve the performance of your model, you can use the FiftyOne Brain to visualize your dataset in a low-dimensional embedding space to reveal patterns and clusters in your data that can help you answer many important questions about your data, from identifying the most critical failure modes of your model, to isolating examples of critical scenarios, to recommending new samples to add to your training dataset. ![](https://cdn.sanity.io/images/h6toihm1/production/71fd22deb539a204db046c2da2c9732b1e59b15b-1205x755.png?auto=format&dpr=2&fit=max&q=75&w=1205) Learn more about how to [visualize embeddings](https://voxel51.com/docs/fiftyone/user_guide/brain.html#brain-embeddings-visualization) in the FiftyOne Docs. ## **Exporting labeled images** Community Slack member Tiffany Chen asked, **_“How do I export labeled images from a dataset of keypoints?”_** You’ll want to use `draw_labels()`. FiftyOne provides native support for rendering annotated versions of image and video samples with label fields overlaid on the source media. The interface for drawing labels on samples is exposed via the Python library and the CLI. You can easily annotate one or more label fields on entire datasets or arbitrary subsets of your datasets that you have identified by constructing a `DatasetView`. Learn more about [drawing labels on samples](https://voxel51.com/docs/fiftyone/user_guide/draw_labels.html) in the FiftyOne Docs. ## **Grouping multiple slices of image, video or point cloud samples** Community Slack member Marco Dal Farra Krsitensen asked, **_“I have multiple slices of an MRI, is there a way to add multiple slices to one image and ideally see all (8 in my case) slices at the same time grouped in one observation?”_** With the latest [FiftyOne 0.17.0 release](https://medium.com/voxel51/announcing-fiftyone-0-17-0-with-grouped-datasets-3d-geolocation-and-custom-plugins-339600ab73a1), there is now support for the creation of grouped datasets, which contain multiple slices of samples of possibly different modalities (image, video, or point cloud) that are organized into groups. Grouped datasets can be used to represent multiview scenes, where data for multiple perspectives of the same scene can be stored, visualized, and queried in ways that respect the relationships between the slices of data. Grouped datasets may contain 3D samples, including point cloud data stored in .pcd format and associated 3D annotations (detections and polylines). ![](https://cdn.sanity.io/images/h6toihm1/production/e20f3b7156b0c31048015d206168f9075ee16903-1324x709.gif?auto=format&dpr=2&fit=max&q=75&w=1324) Learn more about [grouped datasets](https://voxel51.com/docs/fiftyone/user_guide/groups.html) in the FiftyOne Docs. ## **Creating views for samples that contain just the bounding box** Community Slack member Alex Thaman asked, **_“I would like to have a view that, for each bounding box in the dataset, there is a sample that contains just that bounding box. However, unlike to\_patches, I would like the entire image. So basically, for any image in the dataset, there should be N samples of each image, where each sample is the entire image, but contains only one bounding box label. Any tips on how to do this?”_** For background, it can be beneficial to view every object as an individual sample, especially when there are multiple overlapping detections in an image. In FiftyOne this is called a “patches view” and can be created through Python or directly in the App. In regards to Alex’s question, if you clone the patches view, you’ll have an ordinary dataset with the content that he describes: import fiftyone.zoo as foz dataset = foz.load\_zoo\_dataset("quickstart") patches\_view = dataset.to\_patches("ground\_truth") non\_patch\_dataset = patches\_view.clone() print(len(non\_patch\_dataset)) \# 1232 ![](https://cdn.sanity.io/images/h6toihm1/production/296435b5ba7ab5a5687dc9154abce4456fd1d3cd-967x800.png?auto=format&dpr=2&fit=max&q=75&w=967) Learn more about using the [patches view](https://voxel51.com/docs/fiftyone/user_guide/app.html#viewing-object-patches) in the FiftyOne Docs. ## **What’s next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help [bounding boxes](https://voxel51.com/blog/tag/bounding-boxes) [embeddings](https://voxel51.com/blog/tag/embeddings) [FAQ](https://voxel51.com/blog/tag/faq) [filtering](https://voxel51.com/blog/tag/filtering) [grouped datasets](https://voxel51.com/blog/tag/grouped-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/c28199522446929ca5288c0559047354eb93f01a-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Oct 28, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Oct 29, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-28-2022) [FiftyOne Computer Vision Tips and Tricks – Mar 24, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 25, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-24-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4d37703d72d4b83a85bda19eb1999d5247915fc0-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 02, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 3, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-218-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Sept 30, 2022 Oct 1, 2022 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#020dee11341b) [Custom Plugins for the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#d98f9a615e47) [Exporting visualization options from the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#a13db5f9546d) [Retrieving aggregations per video](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#107fb0c30e28) [Working with polylines and labels using the CVAT integration](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#c79e858c46f7) [Exporting only landscape orientation images](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#b284ba2b6816) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#87cdad35d3e8) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#020dee11341b) [Custom Plugins for the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#d98f9a615e47) [Exporting visualization options from the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#a13db5f9546d) [Retrieving aggregations per video](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#107fb0c30e28) [Working with polylines and labels using the CVAT integration](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#c79e858c46f7) [Exporting only landscape orientation images](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#b284ba2b6816) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022#87cdad35d3e8) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/049aec7818dc6abca36e9a736a9aa4e944a2b4ce-1200x671.png?auto=format&dpr=2&fit=max&q=75&w=1200) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help Ok, let’s dive into this week’s tips and tricks! ## **Custom Plugins for the FiftyOne App** Community Slack member Gerard Corrigan asked, _“I’d like to integrate FiftyOne with another app. How can I do that?”_ With the latest [FiftyOne 0.17.0](https://medium.com/voxel51/announcing-fiftyone-0-17-0-with-grouped-datasets-3d-geolocation-and-custom-plugins-339600ab73a1) release you can now customize and extend the FiftyOne App’s behavior. For example, if you need a unique way to visualize individual samples, plot entire datasets, or fetch FiftyOne data, a custom plugin just might be the ticket! ![](https://cdn.sanity.io/images/h6toihm1/production/1f1bea9c5dad46794c819e54e2b8dd2ee63327d0-1400x366.png?auto=format&dpr=2&fit=max&q=75&w=1400) Learn more about [how to develop custom plugins](https://github.com/voxel51/fiftyone/blob/develop/app/packages/plugins/README.md) for the FiftyOne App on GitHub. ## **Exporting visualization options from the FiftyOne App** Community Slack member Murat Aksoy asked, _“Is there a way to export videos from the FiftyOne App? Specifically after an end-user makes adjustments to visualization options such as which labels to show, opacity, etc?”_ The best workflow to accomplish this would be: - Interactively select/filter in the FiftyOne App in IPython/notebook - Press the “bookmark” icon to save your current filters into the view bar - Render the labels for the current view from Python via: session.view.draw\_labels(...) The `drawlabels()` method provides a bunch of options for configuring the look-and-feel of the exported labels, including font sizes, transparency, etc. In a nutshell, you can use the FiftyOne App to visually filter, but you must use `drawlabels()` in Python to trigger the rendering and provide any look-and-feel customizations you want, like transparency. Learn more about the [`drawlabels()`](https://voxel51.com/docs/fiftyone/user_guide/draw_labels.html) [method](https://voxel51.com/docs/fiftyone/user_guide/draw_labels.html) in the FiftyOne Docs. ## **Retrieving aggregations per video** Community Slack member Tadej Svetina asked, _“I have a video dataset and I am interested in getting some aggregations (let’s say count of detections) per video. How do I do that?”_ You can actually use the `values()` aggregation for this along with some pretty advanced view expression usage. Specifically, you can reduce the frames in each video to a single value based on the length of the detections in a field of each video. For example, here’s a way to get the number of detections in every video of the dataset. import fiftyone as fo import fiftyone.zoo as foz from fiftyone import ViewField as F from fiftyone.core.expressions import VALUE dataset = foz.load\_zoo\_dataset("quickstart-video") \# Expression that computes the number of objects in a video field\_name = "detections" num\_objects = F("frames").reduce(VALUE + F(field\_name + ".detections").length()) dataset.values(num\_objects) \# \[493, 1051, 1728, 1375, 248, 381, 2880, 602, 579, 2008\] Or you can modify this code slightly to get a dictionary mapping of video IDs to the number of objects in each video: id\_num\_objects\_map = dict(zip(\*dataset.values(\["id", num\_objects\]))) Learn more about how to use [`reduce()`](https://voxel51.com/docs/fiftyone/api/fiftyone.core.expressions.html#fiftyone.core.expressions.ViewExpression.reduce) in the FiftyOne Docs. ## **Working with polylines and labels using the CVAT integration** Community Slack member Guillaume Dumont asked, _“I am using the CVAT integration with a local CVAT server and somehow, in the cases where the polylines have the same `label_id`, the last one would override the previous one when downloading annotations. This ended up leaving a single polyline where I expected there to be many. Any ideas what’s going on here?”_ When calling `to_polylines()` you want to make sure to use the `mask_types="thing"` rather than the default `mask_types="stuff"` which will give each segment a unique ID. You can also directly annotate semantic segmentation masks and let FiftyOne manage the conversion to polylines for you. Here’s the relevant snippet from [the Docs](https://voxel51.com/docs/fiftyone/api/fiftyone.core.labels.html?highlight=mask#fiftyone.core.labels.Segmentation.to_detections) in regards to `mask_types`: > `mask_types`(“stuff”) — whether the classes are “stuff” (amorphous regions of pixels) or “thing” (connected regions, each representing an instance of the thing). > > Can be any of the following: > > \- “stuff” if all classes are stuff classes > > \- “thing” if all classes are thing classes > > \- a dict mapping pixel values to “stuff” or “thing” for each class Learn more about the [CVAT integration](https://voxel51.com/docs/fiftyone/integrations/cvat.html?highlight=cvat) and `to_polylines()` and [semantic segmentation](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#semantic-segmentation) in the FiftyOne Docs. ## **Exporting only landscape orientation images** Community Slack member Stan asked, _“How would I go about exporting only images in a certain orientation, for example landscape vs portrait? I have a script to tag images as landscape by checking if width is greater than height and then removing all images and annotations for the images that are not in the landscape orientation. What would be the FiftyOne approach for this?”_ Here’s a way to isolate landscape image samples, and then remove all other samples, using for example the quickstart dataset: import fiftyone as fo import fiftyone.zoo as foz from fiftyone import ViewField as F dataset = foz.load\_zoo\_dataset("quickstart") dataset.compute\_metadata() \# Create a view that only contains landscape images view = dataset.match(F("metadata.width") > F("metadata.height")) print(view) # note that size shrunk from 200 to 147 samples \# Visualize only landscape images in the App session = fo.launch\_app(view) \# If you want to delete the portrait images from the dataset view.keep() Learn more about the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html) quickstart dataset in the FiftyOne Docs. ## **What’s next?** - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone) - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help [custom plugins](https://voxel51.com/blog/tag/custom-plugins) [CVAT](https://voxel51.com/blog/tag/cvat) [integrations](https://voxel51.com/blog/tag/integrations) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a3a918e30b0553723b9392ea90763379f98480a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – April 7, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Apr 7, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/f8c59b7ff0a53a527b7002a899ccee84e95dee0a-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 16, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 17, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-16-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/47c4462ab8c14fe4728ce5901402986125d3f99d-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Nov 18, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 19, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-219-lllmstxt|> ## Webinar Recap: Pandas Queries [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Webinar Recap: Pandas-Style Queries for Computer Vision Data Dec 20, 2022 • 10 min read Article content In this article [Presentation Highlights](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#fbdf8ee54534) [Overview & intro to the unstructured nature of computer vision data](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#466ca03a8b63) [What is FiftyOne?](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#5aa865976056) [A note on setup](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#b45b7f48022e) [Live demo time!](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#c4cbb1bd38b6) [1 — The basics: understanding your computer vision dataset and working with samples](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#7b8b97a99665) [2 — Calculating aggregate statistics](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#a272359a995d) [3 — Filtering and matching](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#cce0665ce701) [Summary](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#0bf69742f77f) [Q&A from the Webinar](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#41dee6dd96ab) [Additional Resources](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#bbb8a1011391) [What’s Next](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#3684c13e873e) In this article [Presentation Highlights](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#fbdf8ee54534) [Overview & intro to the unstructured nature of computer vision data](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#466ca03a8b63) [What is FiftyOne?](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#5aa865976056) [A note on setup](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#b45b7f48022e) [Live demo time!](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#c4cbb1bd38b6) [1 — The basics: understanding your computer vision dataset and working with samples](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#7b8b97a99665) [2 — Calculating aggregate statistics](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#a272359a995d) [3 — Filtering and matching](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#cce0665ce701) [Summary](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#0bf69742f77f) [Q&A from the Webinar](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#41dee6dd96ab) [Additional Resources](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#bbb8a1011391) [What’s Next](https://voxel51.com/blog/webinar-recap-pandas-style-queries-for-computer-vision-data#3684c13e873e) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/02ddf540125c56fc7e43080d8c64b02965a840d1-1200x681.png?auto=format&dpr=2&fit=max&q=75&w=1200) [FiftyOne](https://github.com/voxel51/fiftyone) and [pandas](https://pandas.pydata.org/) are both open source Python libraries that make dealing with your data easy. While they serve different purposes — pandas is built for tabular data, and FiftyOne is built for computer vision data — their syntax and functionality are closely aligned. Members of the FiftyOne community asked for a comparison of the two tools, and so we made a collection of materials available. [Jacob Marks](https://www.linkedin.com/in/jacob-marks/), Machine Learning Engineer and Developer Evangelist at Voxel51, recently presented a live webinar on how to perform pandas-style queries for computer vision data with FiftyOne. You can watch the playback on [YouTube](https://www.youtube.com/watch?v=0MhuMQuhYSw), take a look at the [slides](https://docs.google.com/presentation/d/1x_HwxE70sxBD2pExHkWs_UfYdQRdWJDmXsXAFj7MBiE/edit?usp=sharing), see the [transcript](https://www.rev.com/transcript-editor/shared/lSgt05k8OCVQZY0T0j0HS41HAZcWX5WZxFXZPMpetAWuCFg7zIKWPdocuKa8qUAR7-Bvk1zLVyVThq5ZhruLA0DwiFk?loadFrom=SharedLink), and read the recap below for the summary. We also include a section with links to other resources comparing pandas and FiftyOne. Enjoy! https://www.youtube.com/watch?v=0MhuMQuhYSw ## Presentation Highlights ## Overview & intro to the unstructured nature of computer vision data Whether you’re dealing with images, videos, satellite imagery, LiDAR, or other 3D data, the metadata that goes along with the computer vision data is unstructured; detections, tags, key points, segmentation masks, etc. — all of these are more flexible than would typically fit in a tabular format. For example, if we take a look at detections, when you have a dataset of images, not every image is going to have the exact same number of detections. Some images could have zero or one detections, while others have 30 or 40. This means you can’t have a set number of rows to actually structure them. We need a way to be able to work with these types of possibilities. Enter FiftyOne, the open source toolset that brings pandas-style queries to computer vision data. ## What is FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## A note on setup In the webinar, Jacob starts from the point where he already has a dataset (in fact, it’s an image dataset of pandas, you know, “big, fuzzy bamboo-eating bears“), and has performed some follow-on activities to get straight into the demos. Therefore, prerequisites to performing pandas-like functionality as Jacob demonstrates in FiftyOne are to: - Download FiftyOne - Download the [dataset](https://images.cv/download/giant_panda/1300/CALL_FROM_SEARCH/%22giant_panda%22) - Import FiftyOne - Load & format the data into FiftyOne - Generate predictions (classification & detection) — In Jacob’s example, he’s using ResNet-50 for classification predictions and Faster RCNN for detection predictions - Compute metadata, uniqueness, and mistakenness ## Live demo time! From now through the Q&A, Jacob demonstrates operations we would be performing on tabular data in pandas, along with their analogies for working with unstructured computer vision data in FiftyOne. He does this in three broad categories: the basics, aggregate statistics, and filtering and matching. Bonus: everything Jacob demos can also be found in [this Colab notebook](https://colab.research.google.com/drive/10tQmsROZqRvZ9sTYebEGhaIEWQ-80-2D?usp=sharing) so you can follow along. ## 1 — The basics: understanding your computer vision dataset and working with samples The pandas `DataFrame` and FiftyOne `Dataset` classes share many similar functionalities. Jacob covers a number of comparisons of common operations in the two libraries, including starting with printing information on the dataset in pandas with `df.info()` and printing similar information on the dataset in FiftyOne with `print(dataset)`. He then covers performing these basic operations and their accompanying outputs: ![](https://cdn.sanity.io/images/h6toihm1/production/834d7ca3834af07c95ea35bcc555f5261d184511-1846x1244.png?auto=format&dpr=2&fit=max&q=75&w=1600) Jacob also shows how it’s possible to: - Identify what all of the possible values are for classifications in this dataset — `ds.distinct(*)` in FiftyOne is the equivalent of `df[*].unique()` in pandas - Identify all of the distinct detections values, too - Find scenarios where classification values include “giant panda” but detection values do not — Wait, what’s going on here? Find out in just a few paragraphs, and in the final example in section 3. While up until now, we’ve been getting a high-level understanding of what’s in our dataset using our notebook, we haven’t yet taken the next step and explored our computer vision dataset by actually looking through the images. FiftyOne makes it very easy to visualize your computer vision data via the FiftyOne App. ![](https://cdn.sanity.io/images/h6toihm1/production/1b46082a0c1fe04c79ddd7c91a5856f471e43096-1999x1068.png?auto=format&dpr=2&fit=max&q=75&w=1600) Jacob fires up the App and first explores some of the fields that this image dataset contains (a `field` in FiftyOne is the equivalent of a `column` in pandas). Because these are computer vision fields, in the App we can see ground truth labels, classifications, detections, bounding boxes, file paths, uniqueness, mistakenness, and many other goodies. Going back to the issue Jacob identified, where classification values include “giant panda” but detection values do not … it’s easy to go from a notebook to the App with `session.show()` to see what’s going on. Jacob looks at a few samples in the App and sees that the bounding boxes classify a “bear” (but not “panda”). In this scenario, maybe the detection model wasn’t trained on “panda”, it was trained on “bear” as one of the categories. More on this in the last example in section 3. ## 2 — Calculating aggregate statistics FiftyOne makes it easy to get summary statistics of your dataset, in a very similar way to pandas as well. Jacob covers how to compute these aggregations: ![](https://cdn.sanity.io/images/h6toihm1/production/dcfb47bbb0cae63e474421a9d6cdee92e4260e02-1863x946.png?auto=format&dpr=2&fit=max&q=75&w=1600) Talking through some characteristics of the example dataset, Jacob looks at the bounds for his detection model and computes that there are very high confidence detections, but low confidence predictions. He also computes the average confidence for detections, which is very low, but the average confidence for classifications is very high. And because this is computer vision data, we can compute aggregate statistics over fields, and do so across an entire dataset using ViewField. Jacob imports this using: ```python 1from fiftyone import ViewField as F ``` He uses this to compute the length of the mean number of detections across the dataset and find that the average sample has nearly 14 detections, which is a lot. And a lot of those are very low confidence detections, which is something we might want to work on. (In fact, FiftyOne exists to help you build high-quality datasets and computer vision models! Insights like this into your computer vision datasets can help you do just that.) Jacob computes the bounds and finds that the image with the fewest detections has two; and the image with the most detections has 40. ## 3 — Filtering and matching Next, Jacob moves into demonstrating filtering and matching, which he notes are “some of the more interesting, more complex” activities. Jacob explains a few basics including how to: - Look at the most unique images in our dataset using the `match()` expression applied to the uniqueness field - Sort our images by uniqueness using the `ds.sort_by(*)` method (the equivalent of `df.sort_values()` in pandas) - View images that are very crowded, with more than 10 predicted detections in them Jacob then goes into a deeper scenario to identify images that have at least one very high confidence detection in them, and notes that there are only a few and when he looks at these images, he can see some issues emerging, and ways to improve the dataset and models moving forward. Jacob shows in the FiftyOne App the samples that had a very high confidence detection, and it was on a person, not a panda. ![](https://cdn.sanity.io/images/h6toihm1/production/cc21efedf9908794072916a894bfa9a5582c29d0-1999x1068.png?auto=format&dpr=2&fit=max&q=75&w=1600) This is a question to all the ML/CV engineers out there — how would you handle these images in your real-world scenarios? You could send them for reevaluation. You could get more images like this then fine tune the model to still be able to detect in scenarios like this with high confidence. Or you could say this is an edge case that you don’t want to include in your dataset to begin with. Another scenario is to sort by “how wrong our predictions are” (the predictions where our model was most confident, but predicted something that was different than the ground truth classification). Jacob describes how to do this using mistakenness to view the most mistaken images first: ```python 1mistaken_view = dataset.sort_by(F("mistakenness"), reverse=True) ``` Doing this results in this view in the App: ![](https://cdn.sanity.io/images/h6toihm1/production/1e9c5672868fd2333eee56ef1d8ca2fbbe3f924d-1999x1068.png?auto=format&dpr=2&fit=max&q=75&w=1600) Looking at the very first sample in the grid, it’s understandable why it was classified as a `prison` because the panda is in a cage. The panda is there, but the cage is a very dominant feature. We might want to send that back for reevaluation. The final scenario Jacob walks us through is how to compare our classifications and detections. When you have a very large bounding box where the box takes up almost the entire canvas of the image, then the information that goes into the bounding box (the detection model’s classification) is very similar to the information that goes into the classification of the entire image. ```python 1## Get images with very large bounding boxes - where we expect detections and classifications to align 2bbox_area = F("bounding_box")[2] * F("bounding_box")[3] 3 4# Only includes predictions whose bounding boxes have an area of at 5# least 80% of the image, and only include samples with at least 6# one prediction after filtering 7large_boxes_view = dataset.filter_labels("faster_rcnn", bbox_area >= 0.8) ``` Then we can further refine our large box samples by looking for a matches that contain at least one bear classification: ```python 1## Get images with large bounding boxes that contain bears 2large_boxes_bear_view = large_boxes_view.match( 3 F("faster_rcnn.detections.label").contains(["bear"]) 4) 5 6print(large_boxes_bear_view.count()) ``` Finally, we can say that we’re very sure that images of giant pandas are samples that our detection algorithm computes as a large bounding box that is a bear _and_ that our classification model predicts the whole image is a giant panda. So we can start from large boxes that are bears and match for our classification being giant panda, as follows: ```python 1definite_panda_view = large_boxes_bear_view.match(F("resnet50.label") == "giant panda") 2print(definite_panda_view.count()) ``` When we look at these results, these are images that we can use as a starting point for whatever algorithm or procedure we want to iterate on and improve our models. ## Summary FiftyOne allows you to extend the pandas-type querying functionality available for tabular data to the much more flexible, unstructured data that you might expect and encounter in computer vision workflows. In particular, all of the filtering, matching, and querying operations in this presentation have helped us find mistakes in the ground truth and failure modes for our prediction models. We can then use these insights to decide how to handle edge cases. These are just a few ways that some very simple querying and visualization can help you build high-quality datasets and computer vision models. ## Q&A from the Webinar **I’m new to the filter expressions like the ones you demoed. Do you have recommendations on ways I can expand my knowledge in this area?** Yes, this is a popular request and there is a cheat sheet on filtering coming out in a week or two with examples of how to perform these types of filters and matching operations — stay tuned! In the meantime, there are examples in the [FiftyOne view expressions documentation](https://voxel51.com/docs/fiftyone/recipes/creating_views.html#View-expressions) with some filtering and matching examples. We also published a Tips & Tricks blog post focused on [filtering with FiftyOne](https://medium.com/voxel51/fiftyone-filtering-tips-and-tricks-dec-09-2022-58ba13500253). **How many different statistical properties are available there?** There are many different ways available in FiftyOne to compute aggregate statistics about datasets! You can find an overview of the [builtin aggregations](https://voxel51.com/docs/fiftyone/user_guide/using_aggregations.html), covering 15+ examples, in the docs. Also, all builtin aggregations are subclasses of the Aggregation class, each encapsulating the computation of a different statistic about your data. Visit the [fiftyone.core.aggregations](https://voxel51.com/docs/fiftyone/api/fiftyone.core.aggregations.html#module-fiftyone.core.aggregations) module which offers a declarative and highly-efficient approach to computing summary statistics about your datasets and views. **Is there a limit to the number of fields you can put into the data?** There are no limits — you can add as many fields as you want (unless and until your database runs out of memory). You can add detections, classifications, relationships, key points — those are all examples of well defined fields that have structure to them. You can also add your own custom fields, whether those are strings or numbers or something else. We also recently added the capabilities to add dynamic fields, adding more flexibility there. You may want to check out these docs resources: - [FiftyOne dataset fields](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#fields), including adding, editing, and removing fields - [Adding dynamic attributes](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#dynamic-attributes) to fields in your FiftyOne datasets **Where is FiftyOne API documented?** The full API in all its glory is documented [here in the FiftyOne docs](https://voxel51.com/docs/fiftyone/api/fiftyone.html). **Is there any source on ground truth prediction match analysis in object detection such as recall, true positive, false negative, etc.?** Yes, FiftyOne provides a variety of builtin methods for evaluating your model predictions, including regressions, classifications, detections, polygons, instance and semantic segmentations, on both image and video datasets. Therefore it is very easy to evaluate predictions with respect to a ground truth.You can learn more about [evaluating models](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#evaluating-models) in the docs. There are also some [tutorials on model evaluation](https://voxel51.com/docs/fiftyone/tutorials/index.html) in the docs. **Is the FiftyOne open source primary purpose is to improve dataset quality and what other things we can do with images/videos?** At a high level, yes! More specifically, FiftyOne’s primary purpose is to help engineers, data scientists, and others who work with computer vision data (including images, videos, geolocation, and 3D) to increase the transparency and clarity of their data, and have a data-centric approach to machine learning. It exists to help you improve the performance of your computer vision models by helping you curate high quality computer vision datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. ## Additional Resources - [This Colab notebook](https://colab.research.google.com/drive/10tQmsROZqRvZ9sTYebEGhaIEWQ-80-2D?usp=sharing) accompanies the presentation - It leverages the [FiftyOne](https://github.com/voxel51/fiftyone/tree/develop/fiftyone) open source computer vision library to efficiently query unstructured computer vision data - For a high-level motivation, see the blog post [Why FiftyOne is the pandas of computer vision](https://medium.com/voxel51/why-fiftyone-is-the-pandas-of-computer-vision-87618c3f1c3) - For a comprehensive walk-through, see the tutorial [pandas-style queries in FiftyOne](https://voxel51.com/docs/fiftyone/tutorials/pandas_comparison.html) - For the quick and dirty query commands, check out the [pandas vs FiftyOne Cheat Sheet](https://docs.voxel51.com/cheat_sheets/pandas_vs_fiftyone.html) ## What’s Next - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone)! - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help! [pandas](https://voxel51.com/blog/tag/pandas) [pandas-style queries](https://voxel51.com/blog/tag/pandas-style-queries) [webinar recap](https://voxel51.com/blog/tag/webinar-recap) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/308698a5aece1d5b1b95ee1bf52811b24448458c-1200x672.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne 0.18 for Computer Vision\\ \\ Event Recaps\\ \\ • \\ \\ Dec 6, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ed751b5c0fbc3d2cb74f0b30e6418d3319564c-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — December 2022\\ \\ Event Recaps\\ \\ • \\ \\ Dec 13, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/cfa6067062cae206570d98a5e688951723545822-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why FiftyOne is the pandas of Computer Vision\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ Nov 23, 2022](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-220-lllmstxt|> ## Computer Vision Meetups Network [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Announcing the Computer Vision Meetups Network Sponsored by Voxel51 Sep 28, 2022 • 4 min read Article content In this article [What topics will the Meetup focus on?](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#1b08e9fef117) [Meetup locations](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#60455ff353ac) [An exciting line up for the Nov 10 Meetup](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#098f5b862883) [How often will the Meetup meet?](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#0dec4b6c49a2) [Giving back](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#dee4846588c6) [Get involved!](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#d9175e0905e1) [A quick word from our sponsor](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#f29b6dd44620) In this article [What topics will the Meetup focus on?](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#1b08e9fef117) [Meetup locations](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#60455ff353ac) [An exciting line up for the Nov 10 Meetup](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#098f5b862883) [How often will the Meetup meet?](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#0dec4b6c49a2) [Giving back](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#dee4846588c6) [Get involved!](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#d9175e0905e1) [A quick word from our sponsor](https://voxel51.com/blog/announcing-the-computer-vision-meetups-network-sponsored-by-voxel51#f29b6dd44620) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/6a2b391258c94d51b612b5ae8b8f0fc15c431f57-1200x687.png?auto=format&dpr=2&fit=max&q=75&w=1200) At Voxel51 we are excited to announce a virtual network of [12 Meetups focused on computer vision](https://www.meetup.com/pro/computer-vision-meetups/)! Although there are already a variety of machine learning and data science Meetups, we felt that there was an untapped opportunity to establish a group of Meetups that are focused exclusively on the intersection of computer vision, open source, real world applications, and the cutting edge research happening at universities. So, in August we started setting up the Meetups. Curiously, without any promotion on our part, confirmed speakers or dates, the Meetups have managed to organically draw 1,100+ members in a little over a month! To us, this looks like a positive signal that there currently appears to be an unmet desire for this type of content in the Meetup universe. ## **What topics will the Meetup focus on?** These Meetups are geared towards data scientists, machine learning engineers, and open source enthusiasts who want to expand their knowledge of computer vision and complementary technologies. We are putting an emphasis on open source software, and speakers who are computer vision practitioners or academics doing research in the field. We’ll aim to have a mix of technical topics that are approachable to both folks just getting started with computer vision, but also seasoned professionals. ![](https://cdn.sanity.io/images/h6toihm1/production/0d65f67f35ed7c2afbfc5e6978967e492b357c43-1250x705.png?auto=format&dpr=2&fit=max&q=75&w=1250) ## **Meetup locations** No need to suffer from FOMO. Join a Meetup that is most time friendly to your location so you don’t miss out on any events coming to your neck of the woods, either in-person or virtually. Here’s the location of the 12 Meetups: - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) ## **An exciting line up for the Nov 10 Meetup** We’ve already lined up two great speakers for the first Meetup happening on Nov 10. **Scaling Autonomous Vehicles with Computer Vision** [Sri Anumakonda](https://www.linkedin.com/in/srianumakonda/) is an autonomous vehicle developer who has built more than 20 computer vision projects, ranging from lane detection all the way to synthetic data generation for autonomous vehicle training. Sri will be diving deep into the world of self-driving cars and explore how computer vision is used in autonomous vehicles, how companies like Tesla, Wayve.ai and Comma.ai tackle the problem, plus emerging technology and trends in this space. Buckle up, this is sure to be a great talk! **Synthetic Data Generators and Deploying Highly Accurate Retail Supply Chain Computer Vision Apps** [Tarik Hammadou](https://www.linkedin.com/in/tarikhammadou/) has been a Senior Developer Relations Manager at NVIDIA since 2019. He’s the author/co-author of over 20 journal and conference papers, and holds several patents in the area of image processing and sensors. In this talk, Tarik will present a method based on creating a digital twin of the fulfillment or a distribution center facility and generating photorealistic digital assets to train and optimize the classification model to be deployed in the real world. The performance of the training process is then used in a feedback loop to adjust the synthetic data generator until an acceptable result is achieved. Furthermore, he’ll share his deployment orchestration methodology over a large number of compute nodes. ## **How often will the Meetup meet?** Until we can find local co-organizers and spaces to meet at, we’ll be meeting virtually over Zoom. We are committed to hosting two (sometimes three!) diverse speakers every second Thursday of the month. Here’s the schedule for the next three Meetups: - Nov 10, 2022 at 10 AM Pacific - Dec 8, 2022 at 10 AM Pacific - Jan 12, 2023 at 10 AM Pacific ## **Giving back** Swag is nice and a given at these types of events. But, we think we can do better. Bringing folks together every month is a great excuse to use that collective energy to put some goodness back out into the world. At every Meetup, Voxel51 will be giving attendees the opportunity to vote on one of three charities with the winning charity receiving a donation on behalf of the Meetup members. ## **Get involved!** This is a community-driven group of Meetups, so please consider getting involved to help make these events awesome, month after month! Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. **Consider speaking at a future Meetup** Are you working on an interesting computer vision problem at work? Are you the maintainer of an open source computer vision tool or library? If you answered “Yes” to either question, we encourage you to consider speaking at a future event! **Host a local Meetup** Do you know of a meeting space that we could use to host a local Meetup? If so, please reach out. **Become a co-organizer** Are you interested in finding local speakers, hunting down a meeting space, helping with logistics and MC-ing local events? Then consider becoming a local co-organizer. **Sponsor a local Meetup** Does your employer have a marketing or DevRel budget to help support the Meetup with administrative costs, swag giveaways and monthly charitable donations on behalf of the members? Please do not hesitate to get in touch! ## **A quick word from our sponsor** \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop [Voxel51](https://voxel51.com/) is the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. They’ve made it easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes. [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0d65f67f35ed7c2afbfc5e6978967e492b357c43-1250x705.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Computer Vision Meetup Update — November ‘22\\ \\ Product & News\\ \\ • \\ \\ Nov 9, 2022](https://voxel51.com/blog/computer-vision-meetup-update-november-22) [![](https://cdn.sanity.io/images/h6toihm1/production/622b7369c791083b44e3034b2b8772d3ecada8bb-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – November 2023\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Nov 1, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-november-2023) [FiftyOne Computer Vision Community Update – February 2024\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Feb 9, 2024](https://voxel51.com/blog/fiftyone-computer-vision-community-update-february-2024) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-221-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Dec 30, 2022 Dec 31, 2022 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#287c8c715459) [Converting FiftyOne label schema to CVAT](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#1b1ace5781bf) [Understanding the support metric in evaluations](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#d166cebfc266) [Lazy-loading large media files in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#7ebb98b349eb) [Filtering bounding boxes by width and height](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#1612e676d4ce) [Selecting IDs from session](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#9680eb4d93f0) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#eabdd26b468c) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#5afc21188661) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#287c8c715459) [Converting FiftyOne label schema to CVAT](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#1b1ace5781bf) [Understanding the support metric in evaluations](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#d166cebfc266) [Lazy-loading large media files in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#7ebb98b349eb) [Filtering bounding boxes by width and height](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#1612e676d4ce) [Selecting IDs from session](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#9680eb4d93f0) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#eabdd26b468c) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-30-2022#5afc21188661) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ![](https://cdn.sanity.io/images/h6toihm1/production/93682b6d528c21a04ca55782334799a0715028ea-1200x678.png?auto=format&dpr=2&fit=max&q=75&w=1200) ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **Converting FiftyOne label schema to CVAT** Community Slack member Vinicius Madureira asked, _“I’m trying to create a project with FiftyOne’s CVAT utils and initialize it with certain classes. I tried to use a FiftyOne `label_schema`, but it is not working. What would be a correct `label_schema`?”_ To create a CVAT project using [FiftyOne’s CVAT utils](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.cvat.html), you need to pass in a `label_schema` matching CVAT’s label schema. You can generate this from a FiftyOne `label_schema` using the `get_cvat_schema()` function in the CVAT utils: ```python 1import fiftyone as fo 2import fiftyone.utils.cvat as fouc 3cvat = fouc.CVATAnnotationAPI(...) 4 5project_name = “my_project_name” 6project_id = “my_project_id” 7 8### example FiftyOne label schema 9label_schema = { 10 "detections": { 11 "type": "detections", 12 "classes": [\ 13 "class1",\ 14 "class2",\ 15 "class3",\ 16 ] 17 } 18} 19 20# conversion from FiftyOne to CVAT 21# attrs embed concepts like occlusion and group ids from CVAT into FiftyOne 22cvat_schema, assign_scalar_attrs, occluded_attrs, group_id_attrs = cvat._get_cvat_schema( 23 label_schema, 24 project_id=project_id, 25 occluded_attr=occluded_attr, 26 group_id_attr=group_id_attr, 27) 28 29cvat.create_project(project_name,schema=cvat_schema) ``` Once the label schema has been converted, you can pass this into the `create_project()` function to create the project in CVAT. ```python 1cvat.create_project(project_name, schema = cvat_schema) ``` Learn more about FiftyOne’s [integration with CVAT](https://voxel51.com/docs/fiftyone/integrations/cvat.html#cvat-integration) in the FiftyOne Docs. ## **Understanding the support metric in evaluations** Community Slack member Isaac Padberg asked, _“I’m hoping to get some insight into the `support` metric shown when printing out an evaluation report. What does this mean?”_ The `support` in FiftyOne’s evaluation API refers to the number of ground truth instances of each class in a classification or multi-class object detection task. As such, the `support` will not depend on the model predictions you are evaluating. Nevertheless, it can be useful in assessing the model’s performance. When the support for a particular class is small, evaluation metrics for that class can have high variance. Learn more about FiftyOne’s [evaluation API](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html) in the FiftyOne Docs. ## **Lazy-loading large media files in the FiftyOne App** Community Slack member Mohamed Serrari asked, _“Is there a way to lazy-load high-resolution media files like satellite images into the FiftyOne App?”_ The way the FiftyOne App is set up, a media file associated with each sample is displayed in the grid view. The default behavior is for the media file that displays in the grid view to be the same as the media file used to define the `Sample` (located at `sample.filepath` ). This works well for small-to-medium sized images. If you’re dealing with ultra-HD images or satellite imagery, however, this may lead to slow loading in the FiftyOne App. The best practice in these scenarios is to generate a lower resolution thumbnail image for each sample, and configure the app to display the thumbnails in the grid view: ```python 1import fiftyone as fo 2import fiftyone.utils.image as foui 3import fiftyone.zoo as foz 4 5dataset = foz.load_zoo_dataset("quickstart") 6 7# Generate some thumbnail images 8foui.transform_images( 9 dataset, 10 size=(-1, 32), 11 output_field="thumbnail_path", 12 output_dir="/tmp/thumbnails", 13) 14 15dataset.app_config.media_fields = ["filepath", "thumbnail_path"] 16dataset.app_config.grid_media_field = "thumbnail_path" 17dataset.save() ``` The thumbnails can be any size, so long as the aspect ratio matches the original media file. With these settings, the media file that appears in the expanded modal will still be the original, full-resolution version. Learn more about [multiple media fields](https://voxel51.com/docs/fiftyone/user_guide/app.html#multiple-media-fields) and [configuring the FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html#multiple-media-fields) in the FiftyOne Docs. ## **Filtering bounding boxes by width and height** Community Slack member George Pearse asked, _“Is there a way to filter bounding boxes based on their absolute width and height?”_ Yes! In FiftyOne, object detections are stored on samples in `Detection` objects with `bounding_box` attributes, which store the two-dimensional area information in the `[top-left-x, top-left-y, width, height]` format. These dimensions are represented in relative coordinates, with size relative to the total image size, in the range `[0,1]`. To transform this into absolute values for width and height, you can use the absolute width and height for the image in pixels, which is stored in the `metadata` field on the sample. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset(“quickstart”) 5## One-time computation of the metadata for all samples in dataset 6dataset.compute_metadata() 7 8### get abs width and height for first sample 9sample = dataset.first() 10img_width = sample.metadata.width 11img_height = sample.metadata.height ``` Relative to detections on an image, this metadata is stored at the parent level. Fortunately, with FiftyOne it is possible to use this parent level information in filters by prepending the field name with the `$` character and passing it into the ViewField. For example, to get ground truth bounding boxes with `width < 400` and `height < 600` pixels, we can create the following filter: ```python 1from fiftyone import ViewField as F 2 3rel_width = F("bounding_box")[2] 4img_width = F("$metadata.width") 5abs_width = rel_width * img_width 6 7rel_height = F("bounding_box")[3] 8img_height = F("$metadata.height") 9abs_height = rel_height * img_height 10 11size_filter = (abs_width < 400) & (abs_height < 600) ``` Finally, we can apply this filter to our dataset: ```python 1small_boxes_view = dataset.filter_labels( 2 "ground_truth", size_filter 3) ``` Learn more about the [MongoDB syntax underlying FiftyOne expressions](https://voxel51.com/docs/fiftyone/api/fiftyone.core.expressions.html#fiftyone.core.expressions.to_mongo) in the FiftyOne Docs. ## **Selecting IDs from session** Community Slack member Dan Erez asked, _“When using the selection plot tool, is there a way to extract the list of sample ids that I selected?”_ If you’re using the FiftyOne App to plot some of your samples, for instance visualizing a [uMAP](https://umap-learn.readthedocs.io/en/latest/) embedding, then you can access the collection of currently lassoed samples using `plot.selected_ids`. This could be useful if, for instance, you want to [pre-annotate a selection of samples by cluster](https://voxel51.com/docs/fiftyone/tutorials/image_embeddings.html#Pre-annotation-of-samples). This feature exemplifies the ease with which it is possible to move back and forth between the Python SDK and the FiftyOne App. As another example, if you click on an image in the sample grid and select a detection bounding box, you can pass the sample and prediction info to the Python SDK using `session.selected_labels`. Learn more about the [sessions](https://voxel51.com/docs/fiftyone/user_guide/app.html#sessions) and the FiftyOne App in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,200+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,300+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 208+ repositories - 51+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## What’s next? - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [CVAT](https://voxel51.com/blog/tag/cvat) [FAQ](https://voxel51.com/blog/tag/faq) [filtering](https://voxel51.com/blog/tag/filtering) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/204bd847400b1534c494dae594f0797457bc4c90-1620x906.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Filtering Tips and Tricks — Dec 09, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 10, 2022](https://voxel51.com/blog/fiftyone-filtering-tips-and-tricks-dec-09-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/47c4462ab8c14fe4728ce5901402986125d3f99d-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Nov 18, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 19, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-18-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/95b57b1e85f9838a5ad0f07ee3b457331d0dd178-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Nov 11, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 12, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-11-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-222-lllmstxt|> ## FiftyOne Importing Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Importing and Exporting Tips and Tricks — Dec 23, 2022 Dec 24, 2022 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#ebdce9d7b4ae) [A primer on importing and exporting](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#c19754d38ec5) [Always check the zoo](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#b67e34bad56d) [Pattern match to common formats](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#a2585dee0935) [Save space on export](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#b1b683704723) [Don’t skip class](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#89b607b88513) [Cloud-backed media with FiftyOne Teams](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#1f77c379e575) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#dbdcdc10595a) [What’s next?](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#72f56292bf81) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#ebdce9d7b4ae) [A primer on importing and exporting](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#c19754d38ec5) [Always check the zoo](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#b67e34bad56d) [Pattern match to common formats](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#a2585dee0935) [Save space on export](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#b1b683704723) [Don’t skip class](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#89b607b88513) [Cloud-backed media with FiftyOne Teams](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#1f77c379e575) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#dbdcdc10595a) [What’s next?](https://voxel51.com/blog/fiftyone-importing-and-exporting-tips-and-tricks-dec-23-2022#72f56292bf81) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we give practical pointers for using FiftyOne on topics inspired by discussions in the open source community. This week we’ll cover [importing](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/datasets.html#) and [exporting](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html) datasets. ![](https://cdn.sanity.io/images/h6toihm1/production/2849b01b384dc5e7288098649419713d5f0572e3-1200x673.png?auto=format&dpr=2&fit=max&q=75&w=1200) ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **A primer on importing and exporting** [FiftyOne Datasets](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#using-datasets) allow you to easily load, modify, visualize, and evaluate your data along with any related labels (classifications, detections, etc). They provide a consistent interface for loading images, videos, annotations, and model predictions into a format that can be visualized in the FiftyOne App, synced with your annotation source, and shared with others. If you have your own collection of data, [loading it as a Dataset](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html) will allow you to easily search and sort your samples. If you have created a custom dataset in FiftyOne with predictions, annotations, and embeddings, you can export this data for further processing. Continue reading for some tips and tricks to help you master the importing and exporting of datasets in FiftyOne! ## **Always check the zoo** If you want to work with a relatively common computer vision dataset, and you do not yet have the dataset downloaded to disk, your first step should always be to check the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html). The Zoo contains a variety of common datasets across multiple computer vision domains, and new datasets are constantly being added. To load in the desired dataset, find the name of the dataset [here](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html) under the “Details” section, and use the FiftyOne Zoo’s `load_zoo_dataset()` method. To load in the BDD100K dataset, for example, you can do so with: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset(“bdd100k”) ``` Learn more about [load\_zoo\_dataset](https://voxel51.com/docs/fiftyone/api/fiftyone.zoo.datasets.html#fiftyone.zoo.datasets.load_zoo_dataset) and the FiftyOne Dataset Zoo in the FiftyOne Docs. ## **Pattern match to common formats** Before building a custom data importer, it is also worth checking FiftyOne’s built-in [supported data formats](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/datasets.html#supported-import-formats) for loading data from disk. FiftyOne has basic importers for constructing datasets from an input directory for various media types and vision tasks, such as `ImageSegmentationDirectoryImporter` and `VideoClassificationDirectoryImporter`. Additionally, FiftyOne provides import classes for various common dataset formats. If you have a copy of the BDD100K dataset already stored on disk, you can load it into FiftyOne with the `BDDDatasetImporter` class. Many datasets that do not have custom importers in FiftyOne are still structured using common formats. For instance, a new image detection dataset might have its class labels stored in MS COCO format. In cases like this, the `COCODetectionDatasetImporter` will work without modification! Learn more about the FiftyOne [DatasetImporter](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.data.importers.html#fiftyone.utils.data.importers.DatasetImporter) class in the FiftyOne Docs. ## **Save space on export** When exporting a dataset in FiftyOne using the `export()` method, you can use the `export_media()` argument to specify how the media files should be handled. With this parameter, you can choose whether to copy, move, symlink, or omit the media files from the export. If you have limited space, you can pass in `export_media = False` to export the dataset _without_ copying the media files, as in the following example. ```python 1import fiftyone as fo 2 3export_dir = "/path/for/fiftyone-dataset" 4 5# The dataset or view to export 6dataset_or_view = fo.Dataset(...) 7 8# Export the dataset without copying the media files 9dataset_or_view.export( 10 export_dir=export_dir, 11 dataset_type=fo.types.FiftyOneDataset, 12 export_media=False, 13) ``` Learn more about exporting datasets and the FiftyOne [DatasetExporter](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.data.exporters.html#fiftyone.utils.data.exporters.DatasetExporter) class in the FiftyOne Docs. ## **Don’t skip class** Some media dataset formats, such as [COCO](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html#cocodetectiondataset-export) and [YOLO](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html#yolov5dataset-export), require that a list of classes is stored for the label field. If you use the `export()` method without passing in a list of classes via the `classes` argument, the exported list will be generated based on the observed classes in the dataset. While this can be convenient, if you are not careful, it can lead to unexpected behavior. In particular, if the collection of samples being exported does not have any detections or classifications for a certain class (that is present in the list of allowed classes), that class will not appear in the list of classes for the exported dataset. It is best practice to set the `default_classes` property for a dataset while it is in use, and to then pass the argument `classes = dataset.default_classes` into `export()`. For most datasets from the FiftyOne Dataset Zoo, the `default_classes` property is pre-populated. As an example, suppose we create a dataset from COCO samples that contain “cat” or “dog” detections: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5# Load 10 samples containing cats and dogs (among other objects) 6dataset = foz.load_zoo_dataset( 7 "coco-2017", 8 split="validation", 9 classes=["cat", "dog"], 10 shuffle=True, 11 max_samples=10, 12) ``` The `default_classes` property for this dataset still contains all of the COCO classes, even though not every class will show up in the collection: ```python 1# Loading zoo datasets generally populates the `default_classes` attribute 2print(len(dataset.default_classes)) # 91 ``` We can then ensure that all of these classes are exported: ```python 1view.export( 2 labels_path="/path/to/labels.json", 3 dataset_type=fo.types.COCODetectionDataset, 4 classes=dataset.default_classes, 5) ``` Learn more about [default\_classes](https://voxel51.com/docs/fiftyone/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.default_classes) in the FiftyOne Docs. ## **Cloud-backed media with FiftyOne Teams** With FiftyOne Teams, you can import media files directly from S3 buckets and public or private clouds. Likewise, you can export datasets to the cloud! This can come in handy when dealing with very large datasets. Learn more about [FiftyOne Teams](https://voxel51.com/fiftyone-teams/). ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,200+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,300+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,100+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 210+ repositories - 49+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## What’s next? - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [exporting](https://voxel51.com/blog/tag/exporting) [FAQ](https://voxel51.com/blog/tag/faq) [importing](https://voxel51.com/blog/tag/importing) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/d63c2ef9fed7cf00ea9af1c6dd3f2ee4d657c598-1200x685.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Nov 4, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Nov 5, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-nov-4-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/dc8a2e7a894316856af5a109ae43f8787959f179-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Embeddings Tips and Tricks – Mar 31, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 31, 2023](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/3f54d0a45faa06a04b5d0244dd7c092603150cf0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Mar 10, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 11, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-223-lllmstxt|> ## 2022 Highlights for FiftyOne [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) The Greatest Hits of 2022: FiftyOne & Voxel51 Jan 10, 2023 • 9 min read Article content In this article [TL;DR](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#e718b51a420a) [Our Commitment to Open Source and Community](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#0ce361efc39d) [Celebrating Community & Customer Success](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#eb1ed7255339) [New Products to Unlock Unprecedented Data Transparency](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#645061d7d4b4) [Top Blogs You Won’t Want to Miss from 2022](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#6b1f38aed9f0) [We Are Hiring — Join Us!](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#91ae219973f7) [What’s Next?](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#fe86dd723828) In this article [TL;DR](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#e718b51a420a) [Our Commitment to Open Source and Community](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#0ce361efc39d) [Celebrating Community & Customer Success](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#eb1ed7255339) [New Products to Unlock Unprecedented Data Transparency](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#645061d7d4b4) [Top Blogs You Won’t Want to Miss from 2022](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#6b1f38aed9f0) [We Are Hiring — Join Us!](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#91ae219973f7) [What’s Next?](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#fe86dd723828) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Now that a new year has started, it’s a great opportunity to look back at a selection of the greatest moments from last year for [FiftyOne](https://voxel51.com/docs/fiftyone/), the open source computer vision toolset, and [Voxel51](https://voxel51.com/), the company behind FiftyOne. Read on to see some of our biggest moments from 2022 and catch up on interesting news and topics you may have missed! ![](https://cdn.sanity.io/images/h6toihm1/production/2be41b07bd86d7efc5916442ca5b94aa5115234d-1200x674.png?auto=format&dpr=2&fit=max&q=75&w=1200) ## TL;DR - [Our Commitment to Open Source & Community](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#community) - [Celebrating Community & Customer Success](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#success) - [New Products to Unlock Unprecedented Data Transparency](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#products) - [Top Blogs You Won’t Want to Miss](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#top-blogs) - [We’re Hiring — Join Us!](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51#hiring) ## Our Commitment to Open Source and Community At Voxel51, open source, transparency, and giving back to the computer vision community are what we are all about. Whether it’s developing the open source FiftyOne computer vision toolset that helps tens of thousands of engineers and scientists with their ML workflows, sponsoring Meetups to help members boost their computer vision knowledge, or giving to charitable causes on behalf of the community, Voxel51 is committed to bringing transparency and clarity to the world’s data. ### The FiftyOne Community Keeps Growing! 2022 was a good year for the continued growth of the FiftyOne community. We are so thankful for all the community momentum, participation, and contributions. Here are the end-of-the-year 2022 highlights: - Nearly half a million (and counting) FiftyOne installs - 2,375+ [GitHub stars](https://github.com/voxel51/fiftyone/stargazers) - 2,250+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - 1,250+ [Community Slack members](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 215+ repositories - 50+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) In addition, another big highlight of the year: we were recognized as [one of the fastest growing open source startups](https://runacap.com/ross-index/q2-2022/) by Runa Capital! ### FiftyOne Community Rewards One of our favorite highlights last year was recognizing and rewarding community efforts and achievements. In November ’22 we launched the open source [FiftyOne Community Rewards](https://medium.com/voxel51/announcing-open-source-fiftyone-community-rewards-5c30d78df487) program. **In just two months, we rewarded 20+ community members around the world** for helping others in Slack, filing GitHub issues, speaking at Computer Vision Meetups, and sharing their success stories. If you would like to claim some FiftyOne swag for yourself, reach out to us on the [FiftyOne Community Rewards page](https://www.voxel51.com/fiftyone-computer-vision-success-story-submission/) to get the conversation started. ![](https://cdn.sanity.io/images/h6toihm1/production/e3ed352e497d0e500ff4a1484b8422b3c9bef5cb-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) ### Computer Vision Meetups In November ’22 we held the first ever Computer Vision Meetup! Voxel51 sponsors a network of a dozen [Computer Vision Meetups around the world](https://www.meetup.com/pro/computer-vision-meetups/). These Meetups are geared towards data scientists, machine learning engineers, and open source enthusiasts who want to expand their knowledge of computer vision and complementary technologies. We put an emphasis on open source software, and speakers who are computer vision practitioners or academics doing research in the field. Here are the highlights from 2022 and a glimpse at what’s to come in 2023. ### 2022 Computer Vision Meetup Highlights **November ‘22** - Talk #1 — Sri Anumakonda // Autonomous Vehicles with End2end - Talk #2 — Tarik Hammadou, NVIDIA // Retail Supply Chain & Computer Vision - [Get the recap](https://medium.com/voxel51/recapping-the-computer-vision-meetup-november-2022-a2392afd7366) with playbacks, slides, Q&A recap, and more **December ‘22** - Talk #1 — Kris Kitani, Carnegie Mellon University // Wearable Vision Sensors - Talk #2 — Anna Petrovicheva, CEO at CVAT.ai and CTO at OpenCV.ai // The Future of Data Annotation: Trends, Problems & Solutions - Talk #3 — Kacper Łukawski, Qdrant // Using Similarity Learning to Improve Data Quality - [Get the recap](https://medium.com/voxel51/recapping-the-computer-vision-meetup-december-2022-9464004df693#a6d4) with playbacks, slides, Q&A recap, and more ### Upcoming 2023 Computer Vision Meetups **January 12, 2023 (this week!)** - Hyperparameter Scheduling for Computer Vision — [Cameron Wolfe](https://www.linkedin.com/in/cameron-r-wolfe-04744a238/) (Alegion/Rice University) - An introduction to computer vision with Hugging Face transformers — [Julien Simon](https://www.linkedin.com/in/juliensimon/) (Hugging Face) - [Zoom Link](https://us02web.zoom.us/webinar/register/9616708727021/WN_XQIZlMP2RQuCFBwBylyRqA) **February 9, 2023** - Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO — [Paula Ramos, PhD](https://www.linkedin.com/in/paula-ramos-41097319/) (Intel) - Understanding Speech Recognition with OpenAI’s Whisper Model — [Vishal Rajput](https://www.linkedin.com/in/vishal-rajput-999164122/) (AI-Vision Engineer) - [Zoom Link](https://us02web.zoom.us/webinar/register/9216708727643/WN_P8UHtAZGQWOx_A2HM1dcPA) **March 9, 2023** - Lighting up Images in the Deep Learning Era — [Soumik Rakshit](https://www.linkedin.com/in/soumikrakshit/), ML Engineer (Weights & Biases) - Training and Fine Tuning Vision Transformers Efficiently with Colossal AI — [Sumanth P](https://www.linkedin.com/in/sumanth-p-09b339173/) (ML Engineer) - [Zoom Link](https://us02web.zoom.us/webinar/register/8816708728020/WN_mTdNXxTSR-e7bDG5EkH1XQ) Hope to see you at the upcoming Meetup events! ![](https://cdn.sanity.io/images/h6toihm1/production/0d65f67f35ed7c2afbfc5e6978967e492b357c43-1250x705.png?auto=format&dpr=2&fit=max&q=75&w=1250) ## Celebrating Community & Customer Success Our enterprise customers and open source community continued to soar in 2022! We heard from many of you about the successes you’re having with [FiftyOne](https://voxel51.com/fiftyone/) (the open source toolset for building high-quality computer vision datasets and machine learning models) and [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) (built on FiftyOne with additional features that enable multiple users to securely collaborate on the same datasets and models). Nothing makes us happier than hearing how the software we build helps you solve challenges and reach new heights! Here are just a few highlights from what people had to say. ### FiftyOne Open Source ![](https://cdn.sanity.io/images/h6toihm1/production/6830169cfb7c3e5351998872585a6dfc08e056d8-550x126.png?auto=format&dpr=2&fit=max&q=75&w=550) > _At Protex AI we develop a platform to monitor worker health and safety using existing CCTV infrastructure. Our core computer vision technologies are object detection, classification, and pose estimation. We use FiftyOne as a vital component in our pipeline to validate dataset annotations, intelligently subsample datasets to ensure balance, and also to visualize and debug model predictions to assess accuracy._ > > — Patrick Rowsome, Lead Computer Vision Engineer, Protex AI ![](https://cdn.sanity.io/images/h6toihm1/production/a3826946773c311c461f043c64f5ff73559a2bb0-600x314.jpg?auto=format&dpr=2&fit=max&q=75&w=600) > _We’ve been using FiftyOne for over a year and it has drastically changed the way we work. The ability to easily display and analyze our images and their metadata, including experiment results, has been a refreshing change compared to the way we’ve worked before — mainly writing our own metrics and viewers. I’ve personally used FiftyOne for a segmentation model I’ve trained — trying to analyze the results and visually see what my model outputs has been really easy and fluid thanks to FiftyOne._ > > — Ido Greenfeld, AI Team Lead, Taranis ![](https://cdn.sanity.io/images/h6toihm1/production/d4cda78ba2bc1ef8a61d55138f44433016ac7b68-400x235.png?auto=format&dpr=2&fit=max&q=75&w=400) > _FiftyOne provides a superior interface for dealing with computer vision data. Its extensive Python package lets you do almost any data transformation, perform similarity search, and easily evaluate model predictions._ > > — Rustem Galiullin, Data Scientist, G42 ### FiftyOne Teams ![](https://cdn.sanity.io/images/h6toihm1/production/c2ab32f8ed947666d06fb01aae4a2d4a364f7740-300x300.png?auto=format&dpr=2&fit=max&q=75&w=300) > _We’ve used FiftyOne Teams at ADT Commercial for a year now. It has helped us manage our huge datasets, collaborate on model evaluation, tighten our production schedule, and ultimately deliver solutions that help our customers better manage their risk. FiftyOne Teams has added tremendous value to our computer vision processes._ > > — Philippe Sawaya, Director of Artificial Intelligence, ADT Commercial ![](https://cdn.sanity.io/images/h6toihm1/production/aa4dec5c5cb70f1fddd81aaa1ae746f436126103-600x206.png?auto=format&dpr=2&fit=max&q=75&w=600) > _We have seen performance improvements in our models directly due to using FiftyOne Teams for dataset management. FiftyOne Teams has greatly improved the visibility of our datasets across our entire R&D team, and it has made it extremely easy for multiple team members to access and collaborate on datasets._ > > — Ivan Ralašić, CTO and co-founder of Forsight \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop > _We use FiftyOne Teams to organize, select, display, and share our data which has led to better collaboration with and understanding of our large volume of data. FiftyOne Teams enables us to gain insights such as identifying and understanding data problems early, hypothesis validation, and dataset management overall. This has led to better solution engineering and better testing for the products and services we deliver to our customers._ > > — Lanny Lin, Senior Director of AI and Data Science, Vivint ### Share Your Story! We’d love to hear about your successes! Simply reach out to us on the [FiftyOne Community Rewards page](https://www.voxel51.com/fiftyone-computer-vision-success-story-submission/), or work with your Voxel51 customer success manager to get the process started. ## New Products to Unlock Unprecedented Data Transparency At Voxel51, our mission is to bring transparency and clarity to the world’s data. Why are we doing this? We believe that the biggest challenge you have in getting your computer vision models into production today is not the model architecture, it’s the quality of your data. Therefore, we came together as a company [4+ years ago](https://medium.com/voxel51/its-our-birthday-voxel51-turns-four-298ccdf1937d) to make it easy to improve the quality of your data (and therefore models) and to bring data-centric machine learning to the world. With that, here are some product highlights from 2022 to help you gain unprecedented transparency into your computer vision data (images, videos, 3D, and geolocation). ### FiftyOne Open Source We released [15 versions](https://voxel51.com/docs/fiftyone/release-notes.html) of open source FiftyOne in 2022! Although each release delivered a number of powerful community-driven features, here are some of the top highlights: - [Labelbox](https://voxel51.com/docs/fiftyone/integrations/labelbox.html) was added as an official annotation integration (adding to integrations with CVAT and Label Studio) - With support from the [ActivityNet team](http://activity-net.org/download.html), FiftyOne became a recommended tool for downloading, visualizing, and evaluating on the ActivityNet dataset - FiftyOne gained support for [grouped datasets](https://medium.com/voxel51/announcing-fiftyone-0-17-0-with-grouped-datasets-3d-geolocation-and-custom-plugins-339600ab73a1), along with support for 3D data and geolocation data - FiftyOne added support for a plugin system that you can use to customize and extend the App’s behavior ( [custom plugins](https://github.com/voxel51/fiftyone/blob/develop/app/packages/plugins/README.md)) - We released the ability to [declare custom attributes](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#dynamic-attributes) on your label fields (or, in general, any embedded field) and filter by them in the FiftyOne App - Plus, with the help and support of the broader FiftyOne community, we added many more enhancements to the FiftyOne App and core library, new [documentation](https://voxel51.com/docs/fiftyone/), new [integrations](https://voxel51.com/docs/fiftyone/integrations/), new datasets in the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/), and new models in the [FiftyOne Model Zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/) Open source software like this doesn’t happen without the support of an awesome community, so thank you to everyone who helped make all of this possible! ### FiftyOne Teams Originally introduced to a group of organizations in a private beta in late 2021, in September 2022 we publicly announced the launch of [FiftyOne Teams](https://c212.net/c/link/?t=0&l=en&o=3654175-1&h=2163366275&u=https%3A%2F%2Fvoxel51.com%2Ffiftyone-teams%2F&a=FiftyOne+Teams). Built on top of open source FiftyOne, FiftyOne Teams inherits all the goodness of the open source project and adds collaborative features for teams, including cloud-backed media, dataset permissions, versioning, sharing, and more. Our enterprise customers span a variety of industries — automotive, robotics, security, retail, healthcare, agriculture, and more. We absolutely love seeing all the ways our customers are using FiftyOne Teams to power some of today’s most remarkable machine learning and artificial intelligence. Thank you to all our customers for your feedback and support! If you’re new to FiftyOne Teams and would like to find out more about it, simply [reach out](https://voxel51.com/#teams-form) and we’ll set up a consultation and technical demo. ## Top Blogs You Won’t Want to Miss from 2022 Here’s a look back over our most popular and influential blog posts of 2022, in case you missed them or would like to revisit the topics: - [Why 2022 was the most exciting year in computer vision history (so far)](https://medium.com/voxel51/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far-7a4ab8693b27): From AI-generated art to offsides detection in the World Cup, computer vision made a mark in 2022. Read the post to get last year’s computer vision highlights. - [Tunnel vision in computer vision: can ChatGPT see?](https://medium.com/voxel51/tunnel-vision-in-computer-vision-can-chatgpt-see-e6ef037c535): We pushed ChatGPT to its limits to see what it “knows” about computer vision. Check out the findings in this post. - [Why FiftyOne is the pandas of computer vision](https://medium.com/voxel51/why-fiftyone-is-the-pandas-of-computer-vision-87618c3f1c3): The pandas DataFrame & FiftyOne Dataset share similar functionalities. Learn how to use FiftyOne for computer vision data, like pandas for tabular data. - [Forsight Finds a Centralized Dataset Management Solution in FiftyOne Teams](https://medium.com/voxel51/forsight-finds-a-centralized-dataset-management-solution-in-fiftyone-teams-1220b445c8dc): Learn how Forsight uses FiftyOne Teams to manage 1.5TB of visual data, saving R&D time and increasing model performance, while being cost effective. - [Announcing Our $12.5M Series A Funding to Bring Transparency and Clarity to the World’s Data](https://medium.com/voxel51/announcing-our-12-5m-series-a-funding-to-bring-transparency-and-clarity-to-the-worlds-data-79a7d2b8bd0c): Read about our Series A Funding and how we plan to accelerate the next phase of our growth in bringing data-centric machine learning to the world. (You can also learn more in the news coverage from [TechCrunch](https://techcrunch.com/2022/09/21/voxel51-lands-funds-for-its-platform-to-manage-unstructured-data/) and [VentureBeat](https://venturebeat.com/ai/improving-accuracy-of-computer-vision-models-voxel51-raises-12-5m/).) ## We Are Hiring — Join Us! Voxel51 is growing quickly and we are looking for people to grow with us. Join our team to be a part of the ground level innovation in data-centric ML. Voxel51 team members are fueled by learning, adapt quickly to face new challenges, and aim to shatter the status quo. A sample of our [open, remote positions](https://voxel51.com/jobs/): - Account Executive - Graphic Designer - Machine Learning Customer Success Engineer - Machine Learning Developer Evangelist - Product Manager - Software Engineer - Tech Lead - VP of Product ## What’s Next? - If you like what you see on GitHub, [give the FiftyOne project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [community rewards](https://voxel51.com/blog/tag/community-rewards) [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne Teams](https://voxel51.com/blog/tag/fiftyone-teams) [OSS community](https://voxel51.com/blog/tag/oss-community) [success story](https://voxel51.com/blog/tag/success-story) [Year in Review](https://voxel51.com/blog/tag/year-in-review) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/286bca6ba83c8386a9b92750a9249e2e8aba7d8a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – June ‘23\\ \\ Product & News\\ \\ • \\ \\ Jun 2, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/22ddad88a69b3782a530e73d0ec66f5debd8d516-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Forsight Finds a Centralized Dataset Management Solution in FiftyOne Teams\\ \\ Product & News\\ \\ • \\ \\ Nov 2, 2022](https://voxel51.com/blog/forsight-finds-a-centralized-dataset-management-solution-in-fiftyone-teams) [![](https://cdn.sanity.io/images/h6toihm1/production/17422a76c76f14096dce21e43da51945f498a811-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne & FiftyOne Teams\\ \\ Event Recaps\\ \\ • \\ \\ Oct 8, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-224-lllmstxt|> ## FiftyOne Labels Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Labels Tips and Tricks — Jan 06, 2023 Jan 7, 2023 • 4 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#e296670d9752) [A primer on labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#1750d4cc4ddf) [Only download desired labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#ba6293040206) [Manage class names by mapping labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#9fcf73c0018b) [Add custom attributes to labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#0394352fe317) [Customize rendering of labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#8f8c3270b0f3) [Update dataset by merging labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#a85e5ab80854) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#86929cc13974) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#60e5a21727ae) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#e296670d9752) [A primer on labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#1750d4cc4ddf) [Only download desired labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#ba6293040206) [Manage class names by mapping labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#9fcf73c0018b) [Add custom attributes to labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#0394352fe317) [Customize rendering of labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#8f8c3270b0f3) [Update dataset by merging labels](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#a85e5ab80854) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#86929cc13974) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-labels-tips-and-tricks-jan-06-2023#60e5a21727ae) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we give practical pointers for using FiftyOne on topics inspired by discussions in the open source community. This week we’ll cover [labels](https://voxel51.com/docs/fiftyone/user_guide/basics.html#labels). ![](https://cdn.sanity.io/images/h6toihm1/production/527635fdc673c898fe52de3e8127582f6d440c9a-1200x673.png?auto=format&dpr=2&fit=max&q=75&w=1200) ## **Wait, What’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **A primer on labels** In FiftyOne, labels store semantic information about the sample, such as ground annotations or model predictions. FiftyOne provides a `Label` subclass for many common tasks, including [detection](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#object-detection), [classification](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#classification), [segmentation](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#semantic-segmentation), and [keypoints](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#keypoints). Using FiftyOne’s `Label` types enables you to visualize your labels in the FiftyOne App, and there are also a bunch of methods designed specifically to facilitate working with labels. Continue reading for some tips and tricks to help you master labels in FiftyOne! ## **Only download desired labels** The [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html) contains dozens of the most common computer vision datasets, allowing you to easily load these datasets into FiftyOne with a single line of code. Some of these datasets include multiple types of labels. MS COCO, for instance, supports both detections and segmentations. If you are working on a particular task that only uses some of the available label types, you can specify these details when loading the dataset by passing in the `label_types` keyword and an accompanying list, making the loading (and downloading of the dataset) faster. If we only need the segmentation masks for the COCO dataset, we can load in this data with ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset( 5 "coco-2014", 6 label_types=["segmentations"], 7) ``` To see which label types are available for a dataset, check out the section detailing that dataset in the FiftyOne Dataset Zoo documentation. Learn more about the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html) in the FiftyOne Docs. ## **Manage class names by mapping labels** Suppose you’ve downloaded a dataset that has ground truth detections for birds, cats, and dogs, but you want to test out a model that is trained to detect the much broader class `animal`. Once you’ve added your model’s predictions to the dataset, you need to rename the `bird`, `cat`, and `dog`, and ground truth classes to `animal` before you can use FiftyOne’s [evaluation API](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html). Rather than loop through all detections on all samples, you can use FiftyOne’s `map_labels()` method to create a new view with these labels: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# load in dataset 5dataset = foz.load_zoo_dataset("quickstart") 6 7 8ANIMALS = ["bird", "cat", "dog"] 9 10# Replace all animal detection's labels with "animal" 11mapping = {k: "animal" for k in ANIMALS} 12animals_view = dataset.map_labels("predictions", mapping) ``` This approach is not only faster and more efficient than iterating over all samples, it also has the advantage that it preserves the original class names on the dataset. These labels just exist on this view. We get the best of both worlds. Learn more about [map\_labels()](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.map_labels) in the FiftyOne Docs. ## **Add custom attributes to labels** Once you have imported or loaded a labeled dataset into FiftyOne, you can add whatever custom attributes you like! As an example, the `Detections` label class comes with a bounding box in the `bounding_box` field, specified in the `[top-left-x, top-left-y, width, height]` format. If we want to add a custom `bbox_area` attribute to this label representing the area of the bounding box, we can do so as follows: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# load in dataset 5dataset = foz.load_zoo_dataset("quickstart") 6 7# get sample to which we will add attribute 8 9# add attribute to the `Detections` labels in predictions field 10detections = sample.predictions.detections 11for detection in detections: 12 bounding_box = detection["bounding_box"] 13 detection["bbox_area"] = bounding_box[2]*bounding_box[3] 14sample.predictions.detections = detections 15sample.save() ``` Learn more about [custom attributes](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#object-detection) in the FiftyOne Docs. ## **Customize rendering of labels** When you load a labeled dataset into the FiftyOne App, you’ll notice that by default, the labels appear in lowercase text, bounding boxes for detections are all the same color, and confidence scores are displayed. But did you know that these details can be changed? You can control how labels are rendered on images using the annotations utils, `fiftyone.utils.annotations`. To change the line width of bounding boxes and free up each bounding box to have its own color, you can set an annotations config and pass that into the `draw_labeled_images()` method: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3import fiftyone.utils.annotations as foua 4 5# load in dataset 6dataset = foz.load_zoo_dataset("quickstart") 7# Pick a sample 8sample = dataset.first() 9 10config = foua.DrawConfig( 11 { 12 "bbox_linewidth": 5, 13 "per_object_label_colors": True, 14 } 15) 16 17# The path to write the annotated image 18outpath = "/path/for/image-annotated.jpg" 19 20 21# Render the annotated image 22foua.draw_labeled_image(sample, outpath, config=config) ``` Learn more about FiftyOne’s [Annotation API](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.annotations.html#module-fiftyone.utils.annotations) in the FiftyOne Docs. ## **Update dataset by merging labels** As a tool made to curate and improve dataset quality, FiftyOne integrates seamlessly with labeling services like [CVAT](https://voxel51.com/docs/fiftyone/integrations/cvat.html), [Labelbox](https://voxel51.com/docs/fiftyone/integrations/labelbox.html), [Label Studio](https://voxel51.com/docs/fiftyone/integrations/labelstudio.html). A crucial part of many common computer vision workflows is identifying and tagging errors in labeling, sending these samples out for re-annotation, and updating the dataset with the improved data. FiftyOne makes it easy to iterate on your datasets by merging updated labels into an existing label field using the `merge_labels()` method. Suppose that we have sent out a batch of detections to Labelbox for edits to the “ground truth”, using a `label_schema`: ```python 1view = dataset.match_tags("reannotate") 2 3label_schema = { 4 "ground_truth_edits": { 5 "type": "detections", 6 "classes": dataset.distinct("ground_truth.detections.label"), 7 } 8} 9 10anno_key = "fix_labels" 11results = view.annotate( 12 anno_key, 13 label_schema=label_schema, 14 backend="labelbox", 15) ``` Here, `view` is a `DatasetView` consisting of the images we have tagged for re-annotation, and our updated “ground truth” labels for the samples in this view are stored in the temporary `ground_truth_edits` label field. We can merge these revisions into the `ground_truth` label with a single line of code: ```python 1view.merge_labels("ground_truth_edits", "ground_truth") ``` Resulting in an improved dataset! Learn more about [merge\_labels()](https://voxel51.com/docs/fiftyone/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.merge_labels) and [annotating datasets with Labelbox](https://voxel51.com/docs/fiftyone/tutorials/labelbox_annotation.html#) in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,250+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,300+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,400+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 215+ repositories - 52+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## What’s next? - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [FAQ](https://voxel51.com/blog/tag/faq) [labels](https://voxel51.com/blog/tag/labels) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4ac1a727dc192a21563cde51b6e345f620e09376-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4d37703d72d4b83a85bda19eb1999d5247915fc0-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 02, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 3, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) FiftyOne Computer Vision Labels Tips and Tricks — Jan 06, 2023 - Voxel51 <|firecrawl-page-225-lllmstxt|> ## Berkeley Deep Drive Dataset [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) Exploring the Berkeley Deep Drive Autonomous Vehicle Dataset Jan 11, 2023 • 6 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#fd531e41ede8) [About the Berkeley Deep Drive dataset](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#6e2aa06c79e4) [Dataset quick facts](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#bc65e5a27745) [Step 1: Download the dataset](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#7d46c961d3b4) [Step 2: Install FiftyOne](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#7aeaec51d38f) [Step 3: Import the dataset](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#aa9aee0200d1) [Object detection](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#1a112b775cb9) [Frame attributes](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#28479f82cb04) [Drivable area](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#5811e9a2f919) [Lanes](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#e4cc36f8f581) [Final example](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#ddcfcd28b2ad) [Start working with the dataset](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#d5bcbeb30be8) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#fd531e41ede8) [About the Berkeley Deep Drive dataset](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#6e2aa06c79e4) [Dataset quick facts](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#bc65e5a27745) [Step 1: Download the dataset](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#7d46c961d3b4) [Step 2: Install FiftyOne](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#7aeaec51d38f) [Step 3: Import the dataset](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#aa9aee0200d1) [Object detection](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#1a112b775cb9) [Frame attributes](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#28479f82cb04) [Drivable area](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#5811e9a2f919) [Lanes](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#e4cc36f8f581) [Final example](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#ddcfcd28b2ad) [Start working with the dataset](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset#d5bcbeb30be8) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the latest installment of our ongoing blog series where we highlight a dataset from the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html)! FiftyOne provides a Dataset Zoo that contains a collection of common datasets that you can download and load into FiftyOne via a few simple commands. In this post, we explore the [Berkeley Deep Drive](https://bdd-data.berkeley.edu/) dataset. ![](https://cdn.sanity.io/images/h6toihm1/production/33d08c7b16ab5bfa0e4c5a4f936be4592b8e0a90-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop The FiftyOne Dataset Zoo comprises more than 30 datasets, with new datasets being added all the time! They cover a variety of computer vision use cases including: - Video - Location - Point-cloud - Action-recognition - Classification - Detection - Segmentation - Relationships ![](https://cdn.sanity.io/images/h6toihm1/production/c00044b446cf53b5d80426fb6837f80bd2563d32-1400x707.png?auto=format&dpr=2&fit=max&q=75&w=1400) ## **About the Berkeley Deep Drive dataset** The Berkeley Deep Drive (BDD) dataset is one of the largest and most diverse video datasets for autonomous vehicles and heterogeneous multitask learning. The [BDD100K](https://www.bdd100k.com/) dataset contains 100,000 video clips collected from more than 50,000 rides covering New York, San Francisco Bay Area, and other regions. The dataset contains diverse scene types such as city streets, residential areas, and highways. Furthermore, the videos were recorded in diverse weather conditions at different times of the day. The videos are split into training (70K), validation (10K), and test (20K) sets. Each video is 40 seconds long at 720p resolution and a frame rate of 30fps. The frame at the 10th second of each video is annotated for image classification, detection, and segmentation tasks. The version of the dataset we’ll be working with contains the 100K images extracted from the videos as described above, together with the image classification, detection, and segmentation labels. ## **Dataset quick facts** - **Research Paper:** [BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning](https://arxiv.org/abs/1805.04687) - **Authors:** Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, Trevor Darrell - **GitHub:** [README and Code](https://github.com/bdd100k/bdd100k) - **Docs:** [BDD100K Documentation](https://doc.bdd100k.com/) - **Dataset Source:** [BDD100K data and annotations](https://bdd-data.berkeley.edu/login.html) - **Dataset Size:** 7.10 GB - **License:** BSD 3-Clause License - **Last Release:** September 21, 2020 - **FiftyOne Dataset Name:** bdd100k - **Tags:** image, multilabel, automotive, manual - **Supported Splits:** train, validation, test - **Zoo Dataset class:** [BDD100KDataset](https://voxel51.com/docs/fiftyone/api/fiftyone.zoo.datasets.base.html#fiftyone.zoo.datasets.base.BDD100KDataset) ## **Step 1: Download the dataset** In order to load the BDD100K dataset into FiftyOne, you must [download the source data manually](https://bdd-data.berkeley.edu/login.html) with your directory organized in the following manner: ![](https://cdn.sanity.io/images/h6toihm1/production/7ffc45bfce618863422f5af5312d1be4e3518e9a-297x186.png?auto=format&dpr=2&fit=max&q=75&w=297) ## **Step 2: Install FiftyOne** If you don’t already have FiftyOne installed on your laptop, it takes just a few minutes! For example on MacOs: - Verify your version of Python - Create and activate a virtual environment - Install IPython (optionall) - Upgrade your Setuptools - Install FiftyOne ![](https://cdn.sanity.io/images/h6toihm1/production/7cc1ee53acbfa0ea5d4564e0463204a9abab3c76-1728x1080.gif?auto=format&dpr=2&fit=max&q=75&w=1600) Learn more about how to [get up and running with FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html) in the Docs. ## **Step 3: Import the dataset** Now that you have the dataset downloaded and FiftyOne installed, let’s import the dataset into FiftyOne and launch the FiftyOne App. This should take just a few minutes and a few more lines of code. _Make sure to swap in the correct path to your downloaded source files._ ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# The path to the source files that you manually downloaded 5source_dir = "/path/to/dir-with-bdd100k-files" 6 7dataset = foz.load_zoo_dataset( 8 "bdd100k", 9 split="validation", 10 source_dir=source_dir, 11) 12 13session = fo.launch_app(dataset) ``` Your output should look similar to: Preparing split 'validation' in '/path/to/dir-with-bdd100k-files/fiftyone/bdd100k/validation' Preparing training images... Preparing training labels... Preparing validation images... Preparing validation labels... Preparing test images... Parsing dataset metadata Found 10000 samples Dataset info written to '/path/to/dir-with-bdd100k-files/fiftyone/bdd100k/info.json' Loading 'bdd100k' split 'validation' 100% \|███\| 10000/10000 \[3.2m elapsed, 0s remaining, 62.9 samples/s\] Dataset 'bdd100k-validation' created App launched. Point your web browser to http://localhost:5151 The last line in the code snippet we submitted will launch the FiftyOne App in your default browser. You should see the following initial view of the dataset in the FiftyOne App: ![](https://cdn.sanity.io/images/h6toihm1/production/0532ce92f3c67fc6eaaa2a3ef200b92c5ba75071-1400x800.png?auto=format&dpr=2&fit=max&q=75&w=1400) Ok, let’s start exploring the Berkeley Deep Drive dataset! ## **Object detection** For object detection, there are 10 classes to be evaluated. They include: 1: person 2: rider 3: car 4: truck 5: bus 6: train 7: motor (motorcycle) 8: bike (bicycle) 9: traffic light 10: traffic sign **Note:** The field `category_id` range starts at 1 instead of 0. Their distribution within the dataset can be seen in the FiftyOne App under _detections > label._ ![](https://cdn.sanity.io/images/h6toihm1/production/906199a19a1e817279bc01bc026218ddeee887af-283x483.png?auto=format&dpr=2&fit=max&q=75&w=283) We can limit the detections view to specific labels by selecting _detections > label > truck_ for example: ![](https://cdn.sanity.io/images/h6toihm1/production/a75179421ce5da022f46c15884448b1def604458-845x558.png?auto=format&dpr=2&fit=max&q=75&w=845) ## **Frame attributes** The BDD100K dataset has frame attributes including `weather`, `scene`, and `timeofday`. ### **Weather** As the name implies, these labels help us understand what the general weather conditions are in the sample. weather: rainy, snowy, clear, overcast, undefined, partly cloudy, foggy For example _weather > label > snowy_: ![](https://cdn.sanity.io/images/h6toihm1/production/b519db6b6c94e417aa54fe1ada2b762e7a1bb01b-822x460.png?auto=format&dpr=2&fit=max&q=75&w=822) ### **Scene** These labels help us understand the general vehicle-centric environment the objects are moving though. scene: tunnel, residential, parking lot, undefined, city street, gas stations, highway For example _scene > label > tunnel_: ![](https://cdn.sanity.io/images/h6toihm1/production/3a0cc738c08698290a7f4b288bfab0961b353798-806x419.png?auto=format&dpr=2&fit=max&q=75&w=806) ### **Time of day** These labels help us understand at what time of day (which tells us something about the lighting conditions) the sample was taken. timeofday: daytime, night, dawn/dusk, undefined For example _timeofday > label > dawn/dusk_: ![](https://cdn.sanity.io/images/h6toihm1/production/26ff4a1f4dc8a028846487ed557f44d821ab74f4-801x335.png?auto=format&dpr=2&fit=max&q=75&w=801) ## **Drivable area** As you can imagine, with this dataset roads and pathways are going to be key elements of just about any model you develop. Linearly drawn polylines help us trace road and pathway structures connected at individual vertices. For example _polylines > label > drivable area_: ![](https://cdn.sanity.io/images/h6toihm1/production/860d2f39ec95b4c648eba35e65bdedd17af904c1-989x539.png?auto=format&dpr=2&fit=max&q=75&w=989) Zooming in on a specific sample: ![](https://cdn.sanity.io/images/h6toihm1/production/b89d1b33c5f00ebbd014847ecf57efd45274fc31-751x486.png?auto=format&dpr=2&fit=max&q=75&w=751) ## **Lanes** Dashed and solid roadway lanes that influence the behavior of vehicles, pedestrians, and bicycles can be viewed at _polylines > label > lane_: ![](https://cdn.sanity.io/images/h6toihm1/production/829b5d06e999356e0330da6dec7f30f5a9017968-992x581.png?auto=format&dpr=2&fit=max&q=75&w=992) Zooming in on a specific sample: ![](https://cdn.sanity.io/images/h6toihm1/production/7192dc27254000e01c61d0fe990f3ec7af475ffe-1400x588.png?auto=format&dpr=2&fit=max&q=75&w=1400) ## **Final example** Let’s say we want to limit our view to just samples where cars are present on city streets during the daytime while it is raining. Simply make the selections in the sidebar filter to view the 241 samples that meet this criteria. ![](https://cdn.sanity.io/images/h6toihm1/production/6f84d1cb3675b2fc3f69d1725490d65939d6c498-696x636.png?auto=format&dpr=2&fit=max&q=75&w=696) ## **Start working with the dataset** Now that you have a general idea of what the dataset contains, you can start using FiftyOne to perform a variety tasks including: ### **Create dataset views** Think of dataset views as pipeline operations that are applied to a dataset to extract a subset of the dataset whose samples and fields are filtered, sorted, shuffled, etc. Learn more about [how to create dataset views](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) in the Docs. ### **Create aggregations** If you are interested in computing aggregate statistics about datasets, such as label counts, distributions, and ranges, check out the [aggregations](https://voxel51.com/docs/fiftyone/user_guide/using_aggregations.html) section of the Docs. ### **Create interactive plots** FiftyOne provides a powerful plotting framework that contains a variety of interactive plotting methods that enable you to visualize your datasets and uncover patterns that are not apparent from inspecting either the raw media files or aggregate statistics. Learn more about the [plotting capabilities](https://voxel51.com/docs/fiftyone/user_guide/plots.html) in FiftyOne in the Docs. ![](https://cdn.sanity.io/images/h6toihm1/production/b324c405c104251cd4d48ba0022a8b02312a8429-1919x1051.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ### **Annotate datasets** FiftyOne provides a powerful annotation API that makes it easy to add or edit labels on your datasets or specific views into them. Learn more about [annotating datasets](https://voxel51.com/docs/fiftyone/user_guide/annotation.html) in the Docs. ### **Evaluate models** FiftyOne provides a variety of built-in methods for evaluating your model predictions, including regressions, classifications, detections, polygons, instance, and semantic segmentations, on both image and video datasets. Learn more about [evaluating models](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html) in the Docs. ### **FiftyOne Brain** Finally, the FiftyOne Brain provides powerful machine learning techniques you can apply to your workflows including: - [Visualize embeddings to reveal patterns and clusters](https://voxel51.com/docs/fiftyone/user_guide/brain.html#brain-embeddings-visualization) - [Find and work with visually similar data](https://voxel51.com/docs/fiftyone/user_guide/brain.html#brain-similarity) - [Determine the uniqueness of the data for optimal training](https://voxel51.com/docs/fiftyone/user_guide/brain.html#brain-image-uniqueness) - [Find annotation mistakes](https://voxel51.com/docs/fiftyone/user_guide/brain.html#brain-label-mistakes) - [Calculate how easy or difficult it is for your model to understand any given sample](https://voxel51.com/docs/fiftyone/user_guide/brain.html#brain-sample-hardness) [BDD100K](https://voxel51.com/blog/tag/bdd100k) [Berkeley Deep Drive](https://voxel51.com/blog/tag/berkeley-deep-drive) [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/129c3574861e6c307549e14106753af13ecfa2bf-1308x1044.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How to Download ActivityNet and Evaluate Video Understanding Models\\ \\ Datasets\\ \\ • \\ \\ Feb 8, 2022](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models) [![](https://cdn.sanity.io/images/h6toihm1/production/7d8209de12d18f950c73a8d2e335f4a83bc6ce71-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the UCF101 Dataset: A Large-Scale, YouTube-Based Action Recognition Dataset\\ \\ Datasets\\ \\ • \\ \\ Mar 1, 2023](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/40381f5f37fa5fcd70eddca2f63b6710568f5d2c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual Kinship Recognition with the Families in the Wild Computer Vision Dataset\\ \\ Datasets\\ \\ • \\ \\ Dec 7, 2022](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-226-lllmstxt|> ## Natural Language Image Search [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Vector Search](https://voxel51.com/blog/category/vector-search) Finding Images with Words Jan 11, 2023 • 8 min read Article content In this article [Getting set up](https://voxel51.com/blog/finding-images-with-words#63ea02e20c89) [Generating CLIP embeddings](https://voxel51.com/blog/finding-images-with-words#53dd27e8b17d) [Creating the vector index](https://voxel51.com/blog/finding-images-with-words#faab4fefd1be) [Querying the dataset](https://voxel51.com/blog/finding-images-with-words#d28a96dbe6f2) [Putting the pieces together](https://voxel51.com/blog/finding-images-with-words#23548ab812f8) [What can you use this for?](https://voxel51.com/blog/finding-images-with-words#4d661cf9a46d) [Conclusion](https://voxel51.com/blog/finding-images-with-words#c9d1f64511d4) [Join the FiftyOne community!](https://voxel51.com/blog/finding-images-with-words#73918c7fb0c5) In this article [Getting set up](https://voxel51.com/blog/finding-images-with-words#63ea02e20c89) [Generating CLIP embeddings](https://voxel51.com/blog/finding-images-with-words#53dd27e8b17d) [Creating the vector index](https://voxel51.com/blog/finding-images-with-words#faab4fefd1be) [Querying the dataset](https://voxel51.com/blog/finding-images-with-words#d28a96dbe6f2) [Putting the pieces together](https://voxel51.com/blog/finding-images-with-words#23548ab812f8) [What can you use this for?](https://voxel51.com/blog/finding-images-with-words#4d661cf9a46d) [Conclusion](https://voxel51.com/blog/finding-images-with-words#c9d1f64511d4) [Join the FiftyOne community!](https://voxel51.com/blog/finding-images-with-words#73918c7fb0c5) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _Add natural language image search to your computer vision pipelines with FiftyOne, Pinecone, and CLIP_ One of the coolest things to happen in machine learning over the past few years is the dramatic improvement of multi-modal AI , and the ensuing cross-pollination between computer vision and natural language processing. At the center of this movement is [OpenAI’s CLIP model](https://openai.com/blog/clip/), which uses a contrastive learning technique to embed multimedia content — in this case language and images — into the same latent space. The quality of multi-modal models like CLIP has led to advances in zero-shot image classification, knowledge transfer, synthetic data generation, and [semantic search](https://towardsdatascience.com/beyond-tags-and-entering-the-semantic-search-era-on-images-with-openai-clip-1f7d629a9978). It is the last of these that we will be focusing on today! While you may have seen standalone vector search tools, libraries, or demonstrations, today we’re going to show you how to _incorporate natural language image search directly into your computer vision workflows_. To do this, we will be making use of the open source computer vision toolkit [FiftyOne](https://github.com/voxel51/fiftyone), vector database [Pinecone](https://www.pinecone.io/), and OpenAI’s CLIP model. ![](https://cdn.sanity.io/images/h6toihm1/production/c332c478d66b51893447f19eb71d84a940b94a09-1200x677.png?auto=format&dpr=2&fit=max&q=75&w=1200) In this article, we’ll cover: - Loading in your data and embedding model - Generating CLIP embeddings - Creating a vector index - Querying your vector index with text prompts - What you can use this for! Read on to learn how to move beyond labels and towards more general understanding of your computer vision data! ## **Getting set up** The first thing we need to do is install the relevant packages. ```python 1# install FiftyOne 2pip install fiftyone 3 4# install Pinecone 5pip install -U pinecone-client ``` Once FiftyOne is installed, we are good to go, and we can import the library. We’ll also import the `fiftyone.zoo` submodule so we can quickly load in a subset of the MS COCO dataset from the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html), as well as a PyTorch implementation of the CLIP model from the [FiftyOne Model Zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html): ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("coco-2017", split="validation") 5model = foz.load_zoo_model("clip-vit-base32-torch") ``` If you’d like, you can instead download OpenAI’s CLIP model directly from source by following instructions [here](https://github.com/openai/CLIP). Before setting up Pinecone, we can take a look at our data in the FiftyOne App, which you can instantiate in your browser, within a Jupyter notebook, or as a standalone desktop application: ```python 1session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/8b6a3313bd35337134d39289fe017f527d96d654-1400x660.png?auto=format&dpr=2&fit=max&q=75&w=1400) To get started working with Pinecone, you need to set up an account [here](https://www.pinecone.io/), if you don’t have one already, and copy an API key from [here](https://app.pinecone.io/organizations). Import Pinecone and pass in an API key as follows: ```python 1import pinecone 2 3pinecone.init(api_key="API-KEY", environment="us-west1-gcp") ``` Finally, we’ll import the remaining packages that we will need, and will configuring our PyTorch data type: ```python 1import numpy as np 2from pkg_resources import packaging 3import torch 4 5device = "cuda" if torch.cuda.is_available() else "cpu" 6 7if packaging.version.parse( 8 torch.__version__ 9) < packaging.version.parse("1.8.0"): 10 dtype = torch.long 11else: 12 dtype = torch.int ``` ## **Generating CLIP embeddings** To perform searches on our images based on text prompts, we need to generate embeddings for both the images in our dataset, and any text prompts we might want to search against. With FiftyOne’s `compute_embeddings()` method, we can generate embeddings for all of the images in our dataset in one fell swoop, storing them in an `embedding` field: ```python 1dataset.compute_embeddings( 2 model, 3 embeddings_field="embedding" 4) ``` This might take a few minutes, as computing embeddings from deep neural models is in general a relatively intensive task. If you plan to generate embeddings for many samples, best practice is to pre-compute them ahead of time, and if you intend to use these embeddings more than once, I suggest you persist this data by setting ```python 1dataset.persistent = True ``` If we’d like, we can also pair these embeddings with a dimensionality reduction technique to visualize the embeddings. Using the [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html), we can do this with the `compute_visualization()` method: ```python 1import fiftyone.brain as fob 2from fiftyone import ViewField as F 3 4# perform dimensionality reduction using t-SNE 5results = fob.compute_visualization( 6 dataset, 7 embeddings = "embedding", 8 method = "tsne" 9) 10 11# visualize results, labeling by number of objects in image 12results.visualize(labels=F("ground_truth.detections").length()) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/0e189f7c97e6def4bce33b9b57f82e989a826f70-1400x411.png?auto=format&dpr=2&fit=max&q=75&w=1400) Even in this very basic visualization, we can see that the images with tons of objects tend to cluster together. There’s much more that can be gleaned from visualizations like this, but that is not the focus of this article. To compute the embedding for a text prompt, we need to first tokenize the text, then generate a feature vector for the input text which standardizes format, and finally encode this feature vector. The following function takes in a text prompt and the FiftyOne-wrapped PyTorch CLIP model we loaded, and performs the logic of generating the corresponding embedding vector: ```python 1def get_text_embedding(prompt, clip_model): 2 tokenizer = clip_model._tokenizer 3 4 # standard start-of-text token 5 sot_token = tokenizer.encoder["<|startoftext|>"] 6 7 # standard end-of-text token 8 eot_token = tokenizer.encoder["<|endoftext|>"] 9 10 prompt_tokens = tokenizer.encode(prompt) 11 all_tokens = [[sot_token] + prompt_tokens + [eot_token]] 12 13 text_features = torch.zeros( 14 len(all_tokens), 15 clip_model.config.context_length, 16 dtype=dtype, 17 device=device, 18 ) 19 20 # insert tokens into feature vector 21 text_features[0, : len(all_tokens[0])] = torch.tensor(all_tokens) 22 23 # encode text 24 embedding = clip_model._model.encode_text(text_features).to(device) 25 26 # convert to list for Pinecone 27 return embedding.tolist() ``` To generate the embedding for the prompt “a picture of a giraffe”, we could do so with ```python 1prompt = "a picture of a giraffe" 2query_vector = get_text_embedding(prompt, model) ``` ## **Creating the vector index** If you have a paid Pinecone account, then you can work with multiple _vector indices_ at once. However, for this article we’ll be assuming only a free Pinecone account, which limits you to one vector index. In this case, you should run the following code block to delete any existing vector index associated with your account before creating a new one. ```python 1indices = pinecone.list_indexes() 2if len(indices) > 0: 3 pinecone.delete_index(indices[0]) ``` After doing so, you can create a new vector index: ```python 1index_name = "my-index" 2pinecone.create_index( 3 index_name, 4 dimension=512, 5 metric="cosine", 6 pod_type="p1" 7) 8index = pinecone.Index(index_name) ``` Here we have named the index “my-index” for illustrative purposes, but you can use whatever name you would like. We’ve passed in `dimension=512` because that is the dimension of the embedding vectors generated by CLIP, and we have chosen to use a `"cosine"` metric — vectors will be indexed according to their [cosine similarity](https://en.wikipedia.org/wiki/Cosine_similarity) to the query vector. Of course, in general the use of metrics like cosine similarity only makes sense if the embedding vectors have been properly normalized. The `pod_type="p1"` refers to the hardware running the Pinecone service; [p1 pods](https://docs.pinecone.io/docs/indexes#p1-pods) support up to one million vectors. Now we populate the vector database with the embedding vectors from our dataset, [upserting](https://docs.pinecone.io/reference/upsert/) 100 vectors at a time: ```python 1# convert numpy arrays to lists for pinecone 2embeddings = [arr.tolist() for arr in dataset.values("embedding")] 3ids = dataset.values("id") 4 5# create tuples of (id, embedding) for each sample 6index_vectors = list(zip(ids, embeddings)) 7 8def upsert_vectors(index, vectors): 9 num_vectors = len(vectors) 10 num_vectors_per_step = 100 11 num_steps = int(np.ceil(num_vectors/num_vectors_per_step)) 12 for i in range(num_steps): 13 min_ind = num_vectors_per_step * i 14 max_ind = min(num_vectors_per_step * (i+1), num_vectors) 15 index.upsert(index_vectors[min_ind:max_ind]) 16 17upsert_vectors(index, index_vectors) ``` One subtle note is that we are using FiftyOne’s `"id"` value to identify these embedding vectors. We do this because all `"id"` identifiers are unique in FiftyOne, and we can index or subset the samples in our dataset by passing in a list of these values, as we will do shortly. ## **Querying the dataset** Now that we have a vector index and a function for computing the embedding vector for an input text prompt, we are ready to query on our data. To start, let’s say we want to find the 10 gloomiest images. We can do this by first generating the query vector: ```python 1prompt = "a gloomy day" 2query_vector = get_text_embedding(prompt, model) ``` Then querying our vector index for the 10 most similar embedding vectors `top_k=10` to this query vector: ```python 1top_k_samples = index.query( 2 vector=query_vector, 3 top_k=10, 4 include_values=False 5)['matches'] ``` This returns a list of (ten) results, sorted by \`score\` on the cosine similarity metric. ![](https://cdn.sanity.io/images/h6toihm1/production/801d47a879d07cb50a2a3bb39a7f0d8dc0f0e29d-1400x659.png?auto=format&dpr=2&fit=max&q=75&w=1400) Because cosine is a _similarity_ metric and not a distance metric, the most similar results have the highest scores. If you used the “euclidean” _distance_ metric instead, the most similar results would be those with scores closest to zero. Finally, we can use these results to generate a view of the associated samples in FiftyOne. If we wanted to, we could save and store the `score` values generated by this query in a field on the samples. For this illustration however, we’ll keep it simple: ```python 1# get ids of gloomiest samples 2top_k_ids = [res['id'] for res in top_k_samples] 3 4# view these samples, ordered by “gloominess” 5view = dataset.select(top_k_ids, ordered=True) 6session.view = view.view() ``` ![](https://cdn.sanity.io/images/h6toihm1/production/15edab2ca984d851a98e74f4aada2d3aae2c7567-1400x620.png?auto=format&dpr=2&fit=max&q=75&w=1400) We could also query for more complicated things, like “a person holding a baseball bat”: ```python 1prompt = "a person holding a baseball bat" 2query_vector = get_text_embedding(prompt, model) 3top_k_samples = index.query( 4 vector=query_vector, 5 top_k=10, 6 include_values=False 7)['matches'] 8 9# get ids of samples that most resemble a person holding a baseball bat 10top_k_ids = [res['id'] for res in top_k_samples] 11 12# view these samples, ordered by similarity 13view = dataset.select(top_k_ids, ordered=True) 14session.view = view.view() ``` ![](https://cdn.sanity.io/images/h6toihm1/production/7b3425bea25ad1ac82ad8848dd984bf5177f3b90-1600x648.png?auto=format&dpr=2&fit=max&q=75&w=1600) While the results aren’t _perfect_, it’s pretty amazing that you can write an arbitrary natural language query, run it on any arbitrary dataset, and achieve this level of performance and understanding. Just take a second and think about this. The fact that we don’t have to gather a dataset and train a classifier for this task is very cool, and very powerful! Of course, it’s worth noting that results from using CLIP or another foundational multi-modal ML model out of the box like this will in all likelihood not beat a classifier trained or fine-tuned on your specific data for your specific task. But that is not the point. This is an incredible tool to use for _completely unsupervised_ and _unstructured_ exploration of a new dataset! At this point, we could tag these samples in Python for [labeling or annotation](https://voxel51.com/docs/fiftyone/user_guide/annotation.html): ```python 1view.tag_samples("possibly holding baseball bat") ``` Or you can use FiftyOne’s [in-App tagging features](https://voxel51.com/docs/fiftyone/user_guide/app.html#tags-and-tagging). ## Putting the pieces together Now that we have seen the full workflow for querying our data with a search vector, let’s put it all together and wrap it in a function. We’ll call it `sort_by_semantic_similarity()`, drawing inspiration from the [FiftyOne Brain method](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.sort_by_similarity) `sort_by_similarity()`. Let’s make it flexible enough to return either all samples, sorted by semantic similarity, or just the _k_ most similar. _Note: the `top_k` argument for Pinecone queries is capped at 10000. We'll also give an optional `score_field` argument where if `score_field!= None`, we will store the `score` values returned by the index query on the samples._ ```python 1def sort_by_semantic_similarity( 2 dataset, 3 index, 4 prompt, 5 k=None, 6 score_field=None 7): 8 9 query_vector = get_text_embedding(prompt, model) 10 if k is not None: 11 top_k=k 12 else: 13 top_k = int(min(10000, dataset.count())) 14 15 16 result_samples = index.query( 17 vector=query_vector, 18 top_k=top_k, 19 include_values=False 20 )['matches'] 21 22 sample_ids = [res['id'] for res in result_samples] 23 view = dataset.select(sample_ids, ordered=True) 24 25 if score_field is not None: 26 scores = [res['score'] for res in result_samples] 27 dataset.add_sample_field(score_field, fo.FloatField) 28 view.set_values(score_field, scores) 29 dataset.save() 30 31 return view ``` Let’s try it out: ```python 1prompt = "a child playing with a dog" 2 3view = sort_by_semantic_similarity( 4 dataset, 5 index, 6 prompt, 7 k = 30, 8 score_field = "cosine-dist" 9) 10 11session.view = view.view() ``` ![](https://cdn.sanity.io/images/h6toihm1/production/1e8d5a7a1907ad8c576e5569ddd56ece4a4dce84-1400x658.png?auto=format&dpr=2&fit=max&q=75&w=1400) ## What can you use this for? The ability to semantically search through your computer vision datasets can be immensely useful. You can use these capabilities to [pre-annotate your data](https://voxel51.com/docs/fiftyone/tutorials/image_embeddings.html#Pre-annotation-of-samples), or tag samples to [send back to your labeling service provider](https://voxel51.com/docs/fiftyone/tutorials/cvat_annotation.html) for re-annotation. By setting a cutoff, you can also take advantage of the raw `score` values returned by the vector index query for [zero-shot image classification](https://www.pinecone.io/learn/zero-shot-image-classification-clip/) or multi-output classification, or use it as a performance baseline while prototyping. Perhaps most importantly, performing natural language image search using an off-the-shelf foundation model like CLIP can be a great way to perform ad hoc exploration of datasets in an unstructured way. You can query the data with any prompt you can possibly imagine, without the need to perform any model training or fine-tuning of any kind in order to get started. This can be useful in situations where you do not require an exhaustive or exact list of every sample that matches a given criteria. Maybe you are just interested in retrieving a few representative examples of a particular query. Or maybe you are interested in understanding the range of samples in your dataset, beyond aggregate statistics. If you need more systematic results, you’ll likely need to have the data annotated, train or fine-tune a model, or otherwise pass your data on for further processing and analysis. But using CLIP (or another multi-modal embedding model) can help you begin to bootstrap these processes! ## Conclusion Whether you are building predictive models, looking for trends, or assessing the quality of your computer vision data, semantic search is a tool that every machine learning engineer or researcher should have in their arsenal. I hope this article has given you what you need to use this tool in your current and future computer vision workflows! If you want to dive deeper, here are some things for you to play around with: - **Metrics**: cosine vs euclidean… - **Embedding models**: CLIP is not the only one! - **Vector search engines**: Pinecone is just one of many options - **Prompt-crafting**: how does the phrasing of the text prompt affect the result? ## Join the FiftyOne community! \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,250+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,400+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,500+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 220+ repositories - 54+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [CLIP](https://voxel51.com/blog/tag/clip) [multi-modal AI](https://voxel51.com/blog/tag/multi-modal-ai) [natural language processing](https://voxel51.com/blog/tag/natural-language-processing) [NLP](https://voxel51.com/blog/tag/nlp) [OpenAI](https://voxel51.com/blog/tag/openai) [Pinecone](https://voxel51.com/blog/tag/pinecone) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/2867ac2853fae5362ca6bd2d208358dc94556344-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Google Search Experience for Computer Vision Data\\ \\ Tutorials, Vector Search\\ \\ • \\ \\ Mar 22, 2023](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data) [![](https://cdn.sanity.io/images/h6toihm1/production/19eaad3d85784642bb3629627ea6081d1b5c56bc-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Computer Vision Interface for Vector Search\\ \\ Product & News, Vector Search\\ \\ • \\ \\ Jul 12, 2023](https://voxel51.com/blog/the-computer-vision-interface-for-vector-search) [![](https://cdn.sanity.io/images/h6toihm1/production/0f73f70855c4228982bd1a17b9f81c8737a224d8-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Announcing FiftyOne 0.20 with Natural Language Search, Vector Database Integrations, and Point Cloud-Only Datasets\\ \\ Product & News, Vector Search\\ \\ • \\ \\ Mar 22, 2023](https://voxel51.com/blog/announcing-fiftyone-0-20) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-227-lllmstxt|> ## Nearest Neighbor Search Tutorial [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tutorials](https://voxel51.com/blog/category/tutorials), [Vector Search](https://voxel51.com/blog/category/vector-search) Nearest Neighbor Embeddings Search with Qdrant and FiftyOne Jul 22, 2022 • 7 min read Article content In this article [Installation](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#12567a49cd00) [Processing pipeline](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#f530fa3111f8) [Loading the dataset](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#eed0e369b711) [Generating embeddings](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#10765f90f3d6) [Loading embeddings into Qdrant](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#56bbf96e7fc0) [Nearest neighbor classification](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#37b04efb71af) [Evaluation in FiftyOne](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#d9a2cc1b8bbb) [Try it yourself!](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#a32134da0aac) [Summary](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#043675ef9278) In this article [Installation](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#12567a49cd00) [Processing pipeline](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#f530fa3111f8) [Loading the dataset](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#eed0e369b711) [Generating embeddings](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#10765f90f3d6) [Loading embeddings into Qdrant](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#56bbf96e7fc0) [Nearest neighbor classification](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#37b04efb71af) [Evaluation in FiftyOne](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#d9a2cc1b8bbb) [Try it yourself!](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#a32134da0aac) [Summary](https://voxel51.com/blog/nearest-neighbor-embeddings-search-with-qdrant-and-fiftyone#043675ef9278) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _Using FiftyOne and Qdrant to generate classifications on your dataset utilizing model embeddings_ ![](https://cdn.sanity.io/images/h6toihm1/production/6305c0a4564f522f732d9ec86320bf584e282515-1200x673.png?auto=format&dpr=2&fit=max&q=75&w=1200) Neural network embeddings are a low-dimensional representation of input data that give rise to a variety of applications. Embeddings have some interesting capabilities, as they are able to capture the semantics of the data points. This is especially useful for unstructured data like images and videos, so you can not only encode pixel similarities but also some more complex relationships. ![](https://cdn.sanity.io/images/h6toihm1/production/a44c697b29bdf72e406116682bf462c174d547f4-1022x524.png?auto=format&dpr=2&fit=max&q=75&w=1022) Performing searches over these embeddings gives rise to a lot of use cases like classification, building up the recommendation systems, or even anomaly detection. One of the primary benefits of performing a nearest neighbor search on embeddings to accomplish these tasks is that there is no need to create a custom network for every new problem, you can often just use pre-trained models. In fact, it is possible to use the embeddings generated by some publicly available models without any further finetuning. While there are a lot of powerful use cases that involve embeddings, there are a number of challenges in workflows performing searches over embeddings. Specifically, performing a nearest neighbor search on a large dataset and then being able to effectively act on the results of the search, for example performing workflows like auto-labeling of data, are both technical and tooling challenges. To that end, **Qdrant and FiftyOne can help make these workflows effortless**. [Qdrant](https://qdrant.tech/) is an open-source vector database designed to perform an approximate nearest neighbor search (ANN) on dense neural embeddings which is necessary for any production-ready system that is expected to scale to large amounts of data. [FiftyOne](https://fiftyone.ai/) is an open-source dataset curation and model evaluation tool that allows you to effectively manage and visualize your dataset, generate embeddings, and improve your model results. In this article, we’re going to load the MNIST dataset into FiftyOne and perform the classification based on ANN, so the data points will be classified by selecting the most common ground truth label among the K nearest points from our training dataset. In other words, for each test example, we’re going to select its **K** nearest neighbors, using a chosen distance function, and then just select the best label by voting. All that search in the vector space will be done with Qdrant, to speed things up. We will then evaluate the results of this classification in FiftyOne. ## Installation If you want to start using the semantic search with Qdrant, you need to run an instance of it, as this tool works in a client-server manner. The easiest way to do this is to use an official Docker image and start Qdrant with just a single command: ```python 1docker run -p “6333:6333” -p “6334:6334” -d qdrant/qdrant ``` After running the command we’ll have the Qdrant server running, with HTTP API exposed at port 6333 and gRPC interface at 6334. We will also need to install a few Python packages. We’re going to use FiftyOne to visualize the data, along with their ground truth labels and the ones predicted by our embeddings similarity model. The embeddings will be created by MobileNet v2, available in torchvision. Of course, we need to communicate to Qdrant server somehow as well, and since we’re going to use Python, `qdrant_client` is a preferred way of doing that. ```python 1pip install fiftyone 2pip install torchvision 3pip install qdrant_client ``` ## Processing pipeline - Loading the dataset - Generating embeddings - Loading embeddings into Qdrant - Nearest neighbor classification - Evaluation in FiftyOne ## Loading the dataset There are several steps we need to take to get things running smoothly. First of all, we need to load the [MNIST dataset](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#mnist) and extract the train examples from it, as we’re going to use them in our search operations. To make everything even faster, we’re not going to use all the examples, but just 2500 samples. We can use the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html) to load the subset of MNIST we want in just one line of code. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# Load the data 5dataset = foz.load_zoo_dataset("mnist", max_samples=2500) 6 7# Get all training samples 8train_view = dataset.match_tags(tags=["train"]) ``` Let’s start by taking a look at the dataset in the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html). ```python 1# Visualize the dataset in FiftyOne 2session = fo.launch_app(train_view) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/ce54051dcb779ded769cdfca3a298f6535eb294f-964x800.png?auto=format&dpr=2&fit=max&q=75&w=964) ## Generating embeddings The next step is to generate embeddings on the samples in the dataset. This can always be done outside of FiftyOne, with your own custom models. However, FiftyOne also provides various different models in the [FiftyOne Model Zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html) that can be used right out of the box to generate embeddings. In this example, we use [MobileNetv2](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/models.html#mobilenet-v2-imagenet-torch) trained on ImageNet to compute an embedding for each image. ```python 1# Compute embeddings 2model = foz.load_zoo_model("mobilenet-v2-imagenet-torch") 3 4train_embeddings = train_view.compute_embeddings(model) ``` ## Loading embeddings into Qdrant Qdrant allows storing not only vectors but also some corresponding attributes — each data point has a related vector and optionally a JSON payload attached to it. We want to use this to pass in the ground truth label to make sure we can make our prediction later on. ```python 1ground_truth_labels = train_view.values("ground_truth.label") 2train_payload = [\ 3 {"ground_truth": gt} for gt in ground_truth_labels\ 4] ``` Having the embedding created, we can simply start communicating with the Qdrant server. An instance of `QdrantClient` is then helpful, as it encloses all the required methods. Let’s connect and create a collection of points, simply called `“mnist”`. The vector size is dependent on the model output, so if we want to experiment with a different model another day, then we will just need to import a different one, but the rest will be kept the same. Eventually, after making sure the collection exists, we can send all the vectors along with their payloads containing their true labels. ```python 1import qdrant_client as qc 2from qdrant_client.http.models import Distance, VectorParams 3 4# Load the train embeddings into Qdrant 5def create_and_upload_collection( 6 embeddings, payload, collection_name="mnist" 7): 8 client = qc.QdrantClient(host="localhost") 9 client.recreate_collection( 10 collection_name=collection_name, 11 vectors_config=VectorParams( 12 size=embeddings.shape[1], 13 distance=Distance.COSINE, 14 ) 15 ) 16 client.upload_collection( 17 collection_name=collection_name, 18 vectors=embeddings, 19 payload=payload, 20 ) 21 return client 22 23client = create_and_upload_collection(train_embeddings, train_payload) ``` ## Nearest neighbor classification Now to perform inference on the dataset. We can create the embeddings for our test dataset, but just ignore the ground truth and try to find it out using ANN, then compare if both match. Let’s take one step at a time and start with creating the embeddings. ```python 1# Assign the labels to test embeddings by selecting 2# the most common label among the neighbours of each sample 3test_view = dataset.match_tags(tags=["test"]) 4test_embeddings = test_view.compute_embeddings(model) ``` Time for some magic. Let’s simply iterate through the test dataset’s samples and their corresponding embeddings, and use the search operation to find the 15 closest embeddings from the training set. We’ll also need to select the payloads, as they contain the ground truth labels which are required to find the most common label in the neighborhood of a particular point. Python’s `Counter` class will be helpful to avoid any boilerplate code. The most common label will be stored as an `“ann_prediction”` on each test sample in FiftyOne. This is encompassed in the function below which takes an embedding vector as input, uses the Qdrant search capability to find the nearest neighbors to the test embedding, generates a class prediction, and returns a FiftyOne [Classification](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#classification) object that we can store in our [FiftyOne dataset](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html). ```python 1import collections 2from tqdm import tqdm 3 4def generate_fiftyone_classification( 5 embedding, collection_name="mnist" 6): 7 search_results = client.search( 8 collection_name=collection_name, 9 query_vector=embedding, 10 with_payload=True, 11 top=15, 12 ) 13 # Count the occurrences of each class and select the most common label 14 # with the confidence estimated as the number of occurrences of 15 # the most common label divided by a total number of results. 16 counter = collections.Counter( 17 [point.payload["ground_truth"] for point in search_results] 18 ) 19 predicted_class, occurences_num = counter.most_common(1)[0] 20 confidence = occurences_num / sum(counter.values()) 21 prediction = fo.Classification( 22 label=predicted_class, confidence=confidence 23 ) 24 return prediction 25 26predictions = [] 27 28# Call Qdrant to find the closest data points 29for embedding in tqdm(test_embeddings): 30 prediction = generate_fiftyone_classification(embedding) 31 predictions.append(prediction) 32 33test_view.set_values("ann_prediction", predictions) ``` By the way, we estimated the confidence by calculating the fraction of samples belonging to the most common label. That gives us an intuition of how sure we were while predicting the label for each case and can be used in FiftyOne to easily spot confusing examples. ## Evaluation in FiftyOne It’s high time for some results! Let’s start by visualizing how this classifier has performed. We can easily launch the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html) to view the ground truth, predictions, and images themselves. ```python 1session = fo.launch_app(test_view) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/31e98f5966c0e9e2d57e537bc8077284fa4364bc-965x797.png?auto=format&dpr=2&fit=max&q=75&w=965) FiftyOne provides a variety of built-in [methods for evaluating your model](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html) predictions, including regressions, classifications, detections, polygons, instance and semantic segmentations, on both image and video datasets. In two lines of code, we can compute and print an evaluation report of our [classifier](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#classifications). ```python 1# Evaluate the ANN predictions, with respect to the values in ground_truth 2results = test_view.evaluate_classifications( 3 "ann_prediction", gt_field="ground_truth", eval_key="eval_simple" 4) 5 6# Display the classification metrics 7results.print_report() ``` precision recall f1-score support 0 - zero 0.87 0.98 0.92 219 1 - one 0.94 0.98 0.96 287 2 - two 0.87 0.72 0.79 276 3 - three 0.81 0.87 0.84 254 4 - four 0.84 0.92 0.88 275 5 - five 0.76 0.77 0.77 221 6 - six 0.94 0.91 0.93 225 7 - seven 0.83 0.81 0.82 257 8 - eight 0.95 0.91 0.93 242 9 - nine 0.94 0.87 0.90 244 accuracy 0.87 2500 macro avg 0.88 0.87 0.87 2500 weighted avg 0.88 0.87 0.87 2500 After performing the evaluation in FiftyOne, we can use the `results` object to generate an [interactive confusion matrix](https://voxel51.com/docs/fiftyone/user_guide/plots.html#confusion-matrices) allowing us to click on cells and automatically update the App to show the corresponding samples. ```python 1plot = results.plot_confusion_matrix() 2plot.show() ``` \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Let’s dig in a bit further. We can use the [sophisticated query language](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#filtering) of FiftyOne to easily find all predictions that did not match the ground truth, yet were predicted with high confidence. These will generally be the most confusing samples for the dataset and the ones from which we can gather the most insight. ```python 1from fiftyone import ViewField as F 2 3# Display FiftyOne app, but include only the wrong predictions that 4# were predicted with high confidence 5false_view = ( 6 test_view 7 .match(F("eval_simple") == False) 8 .filter_labels("ann_prediction", F("confidence") > 0.7) 9) 10session.view = false_view ``` ![](https://cdn.sanity.io/images/h6toihm1/production/68c1dc623918efcc66268b1628b4ea8cff5eb1bd-963x785.png?auto=format&dpr=2&fit=max&q=75&w=963) These are the most confusing samples for the model and, as you can see, they are fairly irregular compared to other images in the dataset. A next step we could take to improve the performance of the model could be to use FiftyOne to curate additional samples similar to these. From there, those samples can then be annotated through the integrations between FiftyOne and tools like [CVAT](https://voxel51.com/docs/fiftyone/integrations/cvat.html) and [Labelbox](https://voxel51.com/docs/fiftyone/integrations/labelbox.html). Additionally, we could use some more vectors for training or just perform a fine-tuning of the model with similarity learning, for example using the triplet loss. But right now this example of using FiftyOne and Qdrant for vector similarity classification is working pretty well already. And that’s it! As simple as that, we created an ANN classification model using FiftyOne with Qdrant as an embeddings backend, so finding the similarity between vectors can stop being a bottleneck as it would in the case of a traditional k-NN. ## Try it yourself! [Click here for the notebook](https://colab.research.google.com/github/voxel51/fiftyone-examples/blob/feature/qdrant-recipe/examples/Qdrant_FiftyOne_Recipe.ipynb) containing the source code of what you saw in this. Additionally, it includes a realistic use case of this process to perform pre-annotation of night and day attributes on the BDD100K road-scene dataset. ## Summary FiftyOne and Qdrant can be used together to efficiently perform a nearest neighbor search on embeddings and act on the results on your image and video datasets. The beauty of this process lies in its flexibility and repeatability. You can easily load additional ground truth labels for new fields into both FiftyOne and Qdrant and repeat this pre-annotation process using the existing embeddings. This can quickly cut down on annotation costs and result in higher-quality datasets, faster. _This blog post was made in collaboration between the teams at [Qdrant](https://qdrant.tech/) and [Voxel51](https://voxel51.com/) and is co-authored by [Kacper Łukawski](https://www.linkedin.com/in/kacperlukawski)._ [classification](https://voxel51.com/blog/tag/classification) [nearest neighbors](https://voxel51.com/blog/tag/nearest-neighbors) [Qdrant](https://voxel51.com/blog/tag/qdrant) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/2867ac2853fae5362ca6bd2d208358dc94556344-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Google Search Experience for Computer Vision Data\\ \\ Tutorials, Vector Search\\ \\ • \\ \\ Mar 22, 2023](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data) [![](https://cdn.sanity.io/images/h6toihm1/production/19eaad3d85784642bb3629627ea6081d1b5c56bc-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Computer Vision Interface for Vector Search\\ \\ Product & News, Vector Search\\ \\ • \\ \\ Jul 12, 2023](https://voxel51.com/blog/the-computer-vision-interface-for-vector-search) [![](https://cdn.sanity.io/images/h6toihm1/production/cfa6067062cae206570d98a5e688951723545822-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why FiftyOne is the pandas of Computer Vision\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ Nov 23, 2022](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-228-lllmstxt|> ## Building High-Quality Datasets [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Meetup Recap: How to Build High-Quality Machine Learning Datasets and Computer Vision Models Jul 21, 2022 • 4 min read Article content In this article [Video Replay](https://voxel51.com/blog/meetup-recap-how-to-build-high-quality-machine-learning-datasets-and-computer-vision-models#c1a77d3151f2) [Presentation Highlights](https://voxel51.com/blog/meetup-recap-how-to-build-high-quality-machine-learning-datasets-and-computer-vision-models#dfa58c455268) In this article [Video Replay](https://voxel51.com/blog/meetup-recap-how-to-build-high-quality-machine-learning-datasets-and-computer-vision-models#c1a77d3151f2) [Presentation Highlights](https://voxel51.com/blog/meetup-recap-how-to-build-high-quality-machine-learning-datasets-and-computer-vision-models#dfa58c455268) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/a625547e8b6b712e9a9bd5b1300cf9686c11cf20-1200x673.png?auto=format&dpr=2&fit=max&q=75&w=1200) Brian Moore, Co-Founder and CTO of Voxel51, recently presented at the [Virtual MLOps and Kubeflow Meetup](https://www.meetup.com/athens-data-science-machine-learning-mlops-kubeflow/events/285995537/) to share how Voxel51 helps computer vision and machine learning engineers and scientists train better models with measurably better data. In his talk, Brian covers what data-centric ML is and why it’s important, and shows a live demo of FiftyOne, the open-source tool for building high-quality datasets and computer vision models. In this blog post, we provide the [playback recording](https://youtu.be/5E_iC9cFOEU), [slides](https://www.slideshare.net/MichelleBrinich1/building-highquality-datasets-computer-vision-models-with-fiftyone-252242385), and a recap of highlights from the presentation. If you have additional questions about data-centric machine learning, [FiftyOne](https://fiftyone.ai/), FiftyOne Teams, or other computer vision topics, join our [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-gtpmm76o-9AjvzNPBOzevBySKzt02gg) to ask and get answers or follow along with the discussion. ## Video Replay To dive into the meetup presentation, check out this recording, and/or continue reading the highlights below: https://www.youtube.com/watch?v=5E\_iC9cFOEU ## Presentation Highlights ### Introducing Voxel51 Brian opens the meetup presentation with a brief look back at the origins of Voxel51, which was conceived when he met Jason Corso at the University of Michigan. The idea came about to fill the gap at that time in machine learning tooling to take computer vision models into production. Thus Voxel51 and the open source FiftyOne project were born. ### Introducing data-centric machine learning At the core of production-ready models is data-centric machine learning. What is data-centric ML? Brian explains that the biggest challenge in getting a computer vision model into production today is not the model architecture, because there are plenty of architectures you can use and great tools to help you train them. Rather, the biggest challenge is how to improve the quality of your data. Visual datasets today are huge, now reaching hundreds of millions of samples, and you don’t have time to sift through them all to catch any errors. Maybe you can get to 80% accuracy in your dataset pretty easily, or with some additional work you can even get to 90%, but that’s not nearly enough because it can lead to huge issues on the backend of the system due to issues like biased predictions or real-world edge cases that just won’t work for a product you’re releasing to the world. When you have a model trained on poor quality data, this can also lead to a significant decrease in the performance of that model because the data that you were feeding it was not good. So how can you improve the quality of your visual datasets with the goal of getting to higher performance models? ### Where FiftyOne fits in That’s where open source FiftyOne comes in — it helps you integrate with the way that you get data annotated and the way you train your models in order to achieve higher performance models through better data. ![](https://cdn.sanity.io/images/h6toihm1/production/9bceb99b460104346521abd05fcecc3fea612369-1774x732.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### FiftyOne in action Brian then shows a live demo (starting at ~12:00 in the playback video) of FiftyOne that walks you through: - How to install the latest stable version of FiftyOne via pip - How to load in a dataset, including: – Common datasets like COCO and ActivityNet using the FiftyOne Dataset Zoo – Datasets using standard data formats like the COCO format – Custom datasets in your own format – Image datasets, as well as video datasets - How to visualize the dataset in the GUI or code, including: – How to filter and view specific data of interest – How to flexibly interact with your data through the GUI and/or through code - How to import the FiftyOne Brain to look for data by visual similarity, uniqueness, computing your own embeddings, and more - How to work with FiftyOne in an interactive Python shell - How to work with FiftyOne in Jupyter notebooks - How to get hands-on with your data in Jupyter notebooks, including how to run an experiment using a dataset of handwritten digits to find annotation mistakes and automatically or semi-automatically annotate data sets - Another example of how to get hands-on with your data in Jupyter notebooks, including how to use model embeddings from the Model Zoo together with visualization capabilities in an experiment using the BDD100K to find outliers and annotation mistakes ### FiftyOne: resources and next steps Brian shares some resources and next steps to help you get started with and contribute to the open source FiftyOne project: Light reading: - [Overview blog post](https://medium.com/voxel51/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation-73c85b8406e1) - [Installation guide](https://voxel51.com/docs/fiftyone/getting_started/install.html) - [Documentation](http://fiftyone.ai/) - [Tutorials](https://voxel51.com/docs/fiftyone/tutorials/index.html) Next steps: - Like the project? [Give us a star on GitHub](https://github.com/voxel51/fiftyone) - Want to get involved? [Join our Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-gtpmm76o-9AjvzNPBOzevBySKzt02gg) ### A shoutout for FiftyOne Teams For those of you wondering how we make money, Brian explains that we sell a version of FiftyOne called FiftyOne Teams. Teams is a version of open source FiftyOne which is designed for organizations that want to use FiftyOne as the one source of truth for their data. It’s a SaaS deployment of FiftyOne with a centralized database. You can have multiple workflows using Python in parallel to load in data both locally and in the cloud. You can visualize your data sets through the web portal without even using Python, making it more suitable for non-technical workflows. If you’re interested in learning more about Teams, simply [fill out this form](https://share.hsforms.com/1h6pTZ6jxTtSeIx_fhEjYBw2ykyk) and we’ll be in touch. [data-centric computer vision](https://voxel51.com/blog/tag/data-centric-computer-vision) [data-centric machine learning](https://voxel51.com/blog/tag/data-centric-machine-learning) [data-centric ML](https://voxel51.com/blog/tag/data-centric-ml) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne Teams](https://voxel51.com/blog/tag/fiftyone-teams) [MLOps](https://voxel51.com/blog/tag/mlops) [MLOps Meetup](https://voxel51.com/blog/tag/mlops-meetup) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/17422a76c76f14096dce21e43da51945f498a811-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne & FiftyOne Teams\\ \\ Event Recaps\\ \\ • \\ \\ Oct 8, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams) [![](https://cdn.sanity.io/images/h6toihm1/production/61b9a72c1362d209b3cb768c4ddc7f84cdf22452-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Automatically Set Up a New ML Project, Pain Free\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Feb 8, 2023](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free) [![](https://cdn.sanity.io/images/h6toihm1/production/5068fe2d5a454e641a9ad3cc910a9b84d1dec0a6-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ NeurIPS 2023 and the State of AI Research\\ \\ Computer Vision, Event Recaps, Product & News\\ \\ • \\ \\ Dec 8, 2023](https://voxel51.com/blog/neurips-2023-and-the-state-of-ai-research) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-229-lllmstxt|> ## Kinetics Dataset Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) The Kinetics Dataset: Train and Evaluate Video Classification Models Apr 13, 2022 • 6 min read Article content In this article [Setup](https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models#bf3f9f3e33ee) [Downloading Kinetics](https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models#df2f2091b89c) [Training and Evaluating a Model](https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models#f8ad12210d9e) [Summary](https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models#45c657318449) In this article [Setup](https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models#bf3f9f3e33ee) [Downloading Kinetics](https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models#df2f2091b89c) [Training and Evaluating a Model](https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models#f8ad12210d9e) [Summary](https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models#45c657318449) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _A guide to using the open source tool FiftyOne to download the Kinetics dataset and evaluate video understanding models_ ![](https://cdn.sanity.io/images/h6toihm1/production/c4a04d43252eacdfaa8acbfafe00210d4e05f92c-1090x754.png?auto=format&dpr=2&fit=max&q=75&w=1090) After the success of [image classification dataset challenges](https://proceedings.neurips.cc/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf) and the rise of deep learning, tackling video was an obvious next step. Just like the classification of images, the task of video classification is the most straightforward start on the path to general video understanding models. As for the specific labels that are being classified, the computer vision research community has gravitated toward classifying human actions in videos. One of the earliest human action recognition video datasets, even before deep learning took off, was [the KTH dataset](https://ieeexplore.ieee.org/document/1334462) from 2004. Action recognition datasets have come a long way since then, some focusing on clips from [Hollywood movies](https://ieeexplore.ieee.org/document/6126543), while others [focusing on sports](https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/42455.pdf). In 2017, DeepMind released one of the largest and most impactful human action recognition datasets yet, [Kinetics](https://arxiv.org/pdf/1705.06950.pdf). As of the writing of this post, four versions of the Kinetics dataset have been released: [400](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#kinetics-400), [600](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#kinetics-600), [700](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#kinetics-700), and [700–2020](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#kinetics-700-2020). The version number indicates the number of action classes. Additionally, each version adds new videos to replace those that have been deleted from YouTube over time. **This post walks through the [integration of Kinetics](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#kinetics-700-2020) into the open-source dataset curation and model analysis tool, [FiftyOne](https://fiftyone.ai/).** This integration includes a sophisticated way to download the dataset, as well as examples of how to evaluate and improve models trained on the dataset. Downloading Kinetics is now as easy as: import fiftyone.zoo as foz dataset = foz.load\_zoo\_dataset("kinetics-600") ## Setup To run the examples in this post, you need to [install FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html): pip install fiftyone You will also need to install [Pytube](https://pytube.io/en/latest/) which is used by FiftyOne to download videos from YouTube: pip install pytube ## Downloading Kinetics Until recently, the only way to access the Kinetics dataset was to download each video directly from their sources on YouTube. This resulted in [numerous issues](https://towardsdatascience.com/downloading-the-kinetics-dataset-for-human-action-recognition-in-deep-learning-500c3d50f776) including videos having been deleted, YouTube throttling downloads, and inefficiencies in clipping videos. The Common Visual Data Foundation (CVDF) has collaborated with the Kinetics dataset maintainers to [host all versions of the dataset on AWS](https://github.com/cvdfoundation/kinetics-dataset) for the general public to download. It should be noted that the CVDF-hosted version does not include all samples present in the original dataset, only those that were available on YouTube at the time that the CVDF version was created. The CVDF has made it much easier to gain access to the full dataset. However, you still need to handle the challenges of visualizing, wrangling, and subsetting the dataset to meet your needs. In some cases, you don’t want to have to download the entire dataset, to begin with. This is where the integration of [Kinetics into the FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#kinetics-700-2020) comes in. With just one line of Python code, you can now specify the version, the split, and the classes that you want and then visualize it in the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html) with just another line of code. import fiftyone as fo import fiftyone.zoo as foz dataset = foz.load\_zoo\_dataset( "kinetics-700-2020", split="validation", classes=\["grooming cat", "grooming dog"\], max\_samples=10, ) session = fo.launch\_app(dataset) ![](https://cdn.sanity.io/images/h6toihm1/production/d86601de87522ac10033a19edd12607d167dce33-800x600.gif?auto=format&dpr=2&fit=max&q=75&w=800) ## Training and Evaluating a Model After having downloaded Kinetics, you can now start using it to train action recognition models. Since the dataset is already in FiftyOne, it is easy to use libraries like [PyTorch](https://towardsdatascience.com/stop-wasting-time-with-pytorch-datasets-17cac2c22fa8) or [PyTorch Lightning Flash](https://voxel51.com/docs/fiftyone/integrations/lightning_flash.html) to train a model directly on the dataset. pip install lightning-flash lightning-flash\[video\] torchvision pytorchvideo import torch from flash import Trainer from flash.video import VideoClassificationData, VideoClassifier import fiftyone as fo import fiftyone.zoo as foz classes = \[\ \ "swimming backstroke",\ \ "swimming breast stroke",\ \ "swimming butterfly stroke",\ \ "swimming front crawl",\ \ \] \# Load Kinetics dataset = foz.load\_zoo\_dataset( "kinetics-700-2020", splits=\["train", "validation"\], classes=classes, max\_samples=50, shuffle=True, ) \# Replace spaces in class names with underscore labels = dataset.distinct("ground\_truth.label") labels\_map = {l: l.replace(" ", "\_") for l in labels} dataset = dataset.map\_labels("ground\_truth", labels\_map).clone() \# Create views for dataset splits train\_view = dataset.match\_tags("train") val\_view = dataset.match\_tags("validation") \# Create the Flash Datamodule datamodule = VideoClassificationData.from\_fiftyone( train\_dataset=train\_view, val\_dataset=val\_view, predict\_dataset=val\_view, label\_field="ground\_truth", batch\_size=1, clip\_sampler="uniform", clip\_duration=1, decode\_audio=False, ) \# Build the model model = VideoClassifier( backbone="x3d\_xs", labels=datamodule.labels, pretrained=True, ) trainer = Trainer( max\_epochs=10, limit\_train\_batches=5, gpus=torch.cuda.device\_count(), ) \# Finetune the model trainer.finetune(model, datamodule=datamodule, strategy="freeze") After your model is trained, you can then generate predictions on the validation and test splits and use FiftyOne to [evaluate the performance](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#classifications) of the model. from itertools import chain from flash.core.classification import FiftyOneLabelsOutput def get\_fo\_label\_preds(samples, datamodule, trainer): # Return a list of predictions in fo.Detection format predictions = trainer.predict( model, datamodule=datamodule, output=FiftyOneLabelsOutput(return\_filepath=False, labels=datamodule.labels), ) predictions = list(chain.from\_iterable(predictions)) # flatten batches return predictions predictions = get\_fo\_label\_preds(val\_view, datamodule, trainer) \# Add predictions to FiftyOne dataset val\_view.set\_values( "predictions", predictions ) session = fo.launch\_app(val\_view) results = val\_view.evaluate\_classifications( "ground\_truth", "predictions", eval\_key="eval", ) The results of the evaluation can be used for things like plotting [confusion matrices](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#confusion-matrices) and [precision-recall curves](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#map-and-pr-curves). pip install ipywidgets underscore\_classes = \[c.replace(" ", "\_") for c in classes\] plot = results.plot\_confusion\_matrix(classes=underscore\_classes) plot.show() ![](https://cdn.sanity.io/images/h6toihm1/production/7fb6dd12d2a3f9eb2cb2fe68145da08d6d607711-700x500.png?auto=format&dpr=2&fit=max&q=75&w=700) As you can see, since we only finetuned the model on a few dozen samples, it is overfitting to the backstroke and butterfly stroke classes. This implies that we should download additional samples of the other two classes and continue training. Analyzing the model to find the best and worst-performing samples can shed light on the best ways to improve your model’s performance. from fiftyone import ViewField as F eval\_view = val\_view.filter\_labels( "predictions", (F("confidence") > 0.6) & (F("eval") == False) ) session.view = eval\_view The following shows one of the top examples in this evaluation view of highly confident but incorrectly predicted samples. ![](https://cdn.sanity.io/images/h6toihm1/production/9c233cbc39b40882c4adc8e6ff9ce886955c729c-800x600.gif?auto=format&dpr=2&fit=max&q=75&w=800) There are multiple issues that we can see with this sample. First, the footage is first-person which is rare in this dataset. If we want to predict on first-person videos, then more should be added to the training set. Second, there are examples of both breaststroke and backstroke in the video so it would be difficult to assign a label. Third, the ground truth label is front crawl which does not appear at all in the dataset. Using FiftyOne to get hands-on and analyze specific samples can lead to results like these highlighting ways that you can improve the dataset itself. Since Kinetics is a very large dataset, we could easily download additional videos to supplement problematic samples that we may want to exclude from training. Improvements to your dataset can lead to easier gains in model performance than working on improving the model architecture itself. ## Summary The [integration of the Kinetics dataset into FiftyOne](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#kinetics-700-2020) makes it easier than ever to be able to download exactly the subset of Kinetics that you want or even the dataset in its entirety. Additionally, FiftyOne allows for in-depth evaluation and analysis of video models leading to better datasets and higher performing models. [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Kinetics](https://voxel51.com/blog/tag/kinetics) [video classification models](https://voxel51.com/blog/tag/video-classification-models) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/33d08c7b16ab5bfa0e4c5a4f936be4592b8e0a90-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the Berkeley Deep Drive Autonomous Vehicle Dataset\\ \\ Datasets\\ \\ • \\ \\ Jan 11, 2023](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/129c3574861e6c307549e14106753af13ecfa2bf-1308x1044.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How to Download ActivityNet and Evaluate Video Understanding Models\\ \\ Datasets\\ \\ • \\ \\ Feb 8, 2022](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models) [![](https://cdn.sanity.io/images/h6toihm1/production/cc00a6adfa618214f9bdde4d12df92d7e636781d-1400x1112.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The COCO Dataset: Best Practices for Downloading, Visualization, and Evaluation\\ \\ Datasets\\ \\ • \\ \\ Jun 30, 2021](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-230-lllmstxt|> ## Train Your Dragon Detector [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tutorials](https://voxel51.com/blog/category/tutorials) How to Train Your Dragon (Detector) Feb 4, 2022 • 12 min read Article content In this article [Follow Along in Colab](https://voxel51.com/blog/how-to-train-your-dragon-detector#6d7140dcf6bd) [The Tale of the Data and the Model FAIR](https://voxel51.com/blog/how-to-train-your-dragon-detector#090ed357159e) [Integrating ClearML Experiment Tracking](https://voxel51.com/blog/how-to-train-your-dragon-detector#fa90e5bb4437) [2 magical lines](https://voxel51.com/blog/how-to-train-your-dragon-detector#66dba788416b) [Adding Manual Information](https://voxel51.com/blog/how-to-train-your-dragon-detector#6403ff489468) [Adding scalars to the mix](https://voxel51.com/blog/how-to-train-your-dragon-detector#ada1eaf66444) [Debug samples as the proverbial cherry on top](https://voxel51.com/blog/how-to-train-your-dragon-detector#5cc56a11bcf7) [Digging in with FiftyOne](https://voxel51.com/blog/how-to-train-your-dragon-detector#94b552339870) [Loading predictions into FiftyOne](https://voxel51.com/blog/how-to-train-your-dragon-detector#53b10b766dda) [Evaluating and Filtering Results](https://voxel51.com/blog/how-to-train-your-dragon-detector#b5cf8018f741) [What can our failures teach us?](https://voxel51.com/blog/how-to-train-your-dragon-detector#4306043e1698) [Next Steps](https://voxel51.com/blog/how-to-train-your-dragon-detector#f7a67ea697eb) [Summary](https://voxel51.com/blog/how-to-train-your-dragon-detector#aae6924030df) In this article [Follow Along in Colab](https://voxel51.com/blog/how-to-train-your-dragon-detector#6d7140dcf6bd) [The Tale of the Data and the Model FAIR](https://voxel51.com/blog/how-to-train-your-dragon-detector#090ed357159e) [Integrating ClearML Experiment Tracking](https://voxel51.com/blog/how-to-train-your-dragon-detector#fa90e5bb4437) [2 magical lines](https://voxel51.com/blog/how-to-train-your-dragon-detector#66dba788416b) [Adding Manual Information](https://voxel51.com/blog/how-to-train-your-dragon-detector#6403ff489468) [Adding scalars to the mix](https://voxel51.com/blog/how-to-train-your-dragon-detector#ada1eaf66444) [Debug samples as the proverbial cherry on top](https://voxel51.com/blog/how-to-train-your-dragon-detector#5cc56a11bcf7) [Digging in with FiftyOne](https://voxel51.com/blog/how-to-train-your-dragon-detector#94b552339870) [Loading predictions into FiftyOne](https://voxel51.com/blog/how-to-train-your-dragon-detector#53b10b766dda) [Evaluating and Filtering Results](https://voxel51.com/blog/how-to-train-your-dragon-detector#b5cf8018f741) [What can our failures teach us?](https://voxel51.com/blog/how-to-train-your-dragon-detector#4306043e1698) [Next Steps](https://voxel51.com/blog/how-to-train-your-dragon-detector#f7a67ea697eb) [Summary](https://voxel51.com/blog/how-to-train-your-dragon-detector#aae6924030df) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _A general guide to building high-quality deep learning datasets and models using ClearML and FiftyOne_ \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Often thought to be the stuff of legends, we aim to shed light on a mythical entity. Dragons? No… An effective and repeatable data-centric machine learning workflow, of course! It seems that many machine learning researchers and engineers these days are focused on developing the optimal model architecture for their tasks. In reality, the most surefire way to develop a high-performing model is to meticulously understand, track, and improve your datasets and experimental results. In this post, we will walk through the process of developing a computer vision model and dataset in a repeatable and effective way utilizing [ClearML](https://clear.ml/) and [FiftyOne](http://fiftyone.ai/). Specifically, we will be training the object detection model [DETR](https://github.com/facebookresearch/detr) on a dataset of dragon images, though the general workflow presented is extensible to nearly any computer vision and machine learning task. [FiftyOne](https://fiftyone.ai/) is an open-source tool for building high-quality datasets and computer vision models with a [powerful API](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) and [intuitive App](https://voxel51.com/docs/fiftyone/user_guide/app.html) letting you quickly understand the quality of your dataset, find your model’s failure modes, and improve your datasets and models. On the other hand, [ClearML](https://clear.ml/) is an open-source platform that automates and simplifies developing and managing machine learning solutions through an end-to-end MLOps suite allowing you to focus on developing your ML code and automation, while ClearML ensures your work is reproducible and scalable. ClearML and FiftyOne go hand-in-hand with one another in your machine learning workflows. The combination of flexible, hands-on visualization and analysis of data and model results of FiftyOne combined with the experimental result tracking of ClearML produces a system that lets you quickly explore and improve your datasets while also persisting all of the changes and progress made to achieving a high-performing model. ## Follow Along in Colab You can follow along with this entire post directly in your browser through [this Google Colab notebook](https://colab.research.google.com/github/voxel51/fiftyone-examples/blob/cml/examples/training_clearml_detector.ipynb)! ## The Tale of the Data and the Model FAIR To keep things interesting, we created a whole new dataset especially for the occasion: dragons! We gathered 115 images of dragons and annotated them. Interestingly they are both cartoon-style dragons and more ‘realistic’ dragons. To start things out, let’s [download the dataset from here](https://github.com/thepycoder/dragon_data). For the detector, we used [Meta (Facebook) research’s DETR](https://github.com/facebookresearch/detr), an object detection network based on the popular transformers architecture. ## Integrating ClearML Experiment Tracking ClearML experiment tracking works out of the box with most model training frameworks, including facebook’s own detectron2. Most things or imports are still tracked automagically using only the 2 magic lines (which we will cover below), but we also wanted to show how you can track specific metrics or images manually if required. DETRs original codebase keeps track of training metrics in a `.txt` file and does not integrate with any additional tools like tensorboard. So let's change that by adding only a few lines to make ClearML track our training runs automatically. This way, we can keep making the model better in an efficient way. To get started with ClearML, go to the [community server](https://app.community.clear.ml/) and get a ClearML account, or set up your own ClearML server as described [here](https://clear.ml/docs/latest/docs/deploying_clearml/clearml_server). This server will keep track of and consolidate all your code, models, data, and output. After you have the server set up, it’s just a few commands before we can get cracking! First of all, get the ClearML python package. pip install clearml Now we need to let your computer know where your server is. clearml-init This command will set up the connection from your local PC to the ClearML server. It will ask you for your ClearML configuration. You can get this info from the Profile page in your ClearML server. ## 2 magical lines With only 2 lines of python code, we can already start tracking a lot! We add the following 2 lines of code to the top of the [main.py](http://main.py/) file in the DETR repository. Every time we call this training script, ClearML will create a new task and log as much as it can. \# Initialise a clearML task and its corresponding logger from clearml import Task task = Task.init(project\_name='dragon\_detector', task\_name=f'DETR') These 2 lines will already track a lot of information, they will track: - Artifacts (like saved models and checkpoints) - Source code and packages information - Configuration and hyperparameters - Other info (runtime, hardware specs, etc.) - Console output ## Adding Manual Information When using a framework like detectron2, these 2 lines are all you need. But this time, we’re training DETR using raw PyTorch and no additional tools like tensorboard, so we have to do a little more work ourselves. For our own information to the experiment, we need a logger, which can be made as such: logger = task.get\_logger() And now, we’re ready to add whatever information we desire to the experiment! ## Adding scalars to the mix Scalars are any type of value that we want to track and plot later down the line. In most cases, these will be training output values such as losses and accuracy metrics. The original DETR implementation already logs the scalars we want to a `.txt` file, so all we have to do, is capture these parameters and add them to the ClearML task using the `logger.report_scalar` function. \# This code was already in DETR and keeps track of training metrics test\_stats, coco\_evaluator = evaluate( model, criterion, postprocessors, data\_loader\_val, base\_ds, device, args.output\_dir ) log\_stats = {\*\*{f'train\_{k}': v for k, v in train\_stats.items()}, \*\*{f'test\_{k}': v for k, v in test\_stats.items()}, 'epoch': epoch, 'n\_parameters': n\_parameters} \# We add these lines to capture those training metrics in clearML \# Add all metrics except for coco\_eval\_bbox, \# since that is a bbox and not a scalar. for key, value in log\_stats.items(): if 'coco\_eval\_bbox' in key: continue logger.report\_scalar(title=key, series=key, value=value, iteration=epoch) \# This is where DETR normally writes these values to a txt file if args.output\_dir and utils.is\_main\_process(): with (output\_dir / "log.txt").open("a") as f: f.write(json.dumps(log\_stats) + "n") ## Debug samples as the proverbial cherry on top Debug samples are images that can be logged in much the same way as scalars. ClearML will pick up a `matplotlib imshow` image using the `logger.report_matplotlib_figure` function. In this case, we added a function that runs the model on the validation set at the end of training and log the annotated images to the experiment, to provide a quick glance at how well the model performs. def plot\_image\_results(pil\_img, prob, boxes, classes, logger, img\_name): # colors for visualization colors = \[\[0.000, 0.447, 0.741\], \[0.850, 0.325, 0.098\], \[0.929, 0.694, 0.125\],\ \ \[0.494, 0.184, 0.556\], \[0.466, 0.674, 0.188\], \[0.301, 0.745, 0.933\]\] colors = colors \* 100 figure = plt.figure(figsize=(16,10)) plt.imshow(pil\_img) ax = plt.gca() for p, (xmin, ymin, xmax, ymax), c in zip(prob, boxes.tolist(), colors): ax.add\_patch(plt.Rectangle((xmin, ymin), xmax - xmin, ymax - ymin, fill=False, color=c, linewidth=3)) cl = p.argmax() text = f'{classes\[cl\]}: {p\[cl\]:0.2f}' ax.text(xmin, ymin, text, fontsize=15, bbox=dict(facecolor='yellow', alpha=0.5)) plt.axis('off') logger.report\_matplotlib\_figure('Evaluation Results', img\_name, figure, iteration=None, report\_image=True, report\_interactive=False) Now we can easily train DETR in whatever way we want and be sure we captured all the relevant metrics. We can also always recreate our best experiments thanks to our comprehensive logging. If we want to dig deeper and debug our model’s performance meticulously, we can analyze it using FiftyOne! ## Digging in with FiftyOne To start, let’s [install FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html): pip install fiftyone The first step is to [load the dataset into FiftyOne](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html) so we can take a look at it in the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html). [Loading any custom dataset](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html#custom-formats) into FiftyOne is as simple as writing a Python loop. However, since this dataset is already in [COCO format](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/datasets.html#cocodetectiondataset), we can load the splits with just one line of code. import fiftyone as fo dataset\_name = "dragons" dataset\_dir = "/path/to/dataset" \# Load the training dataset into FiftyOne and tag all samples with "train" dataset = fo.Dataset.from\_dir( dataset\_type=fo.types.COCODetectionDataset, data\_path= os.path.join(dataset\_dir, "train"), labels\_path= os.path.join(dataset\_dir, "annotations/train.json"), name=dataset\_name, tags="train", ) \# Add the validation data and tag the samples with "val" dataset.add\_dir( dataset\_type=fo.types.COCODetectionDataset, data\_path= os.path.join(dataset\_dir, "val"), labels\_path= os.path.join(dataset\_dir, "annotations/val.json"), tags="val", ) The `launch_app()` method launches the App directly in the output of this cell and also returns a Session instance, which you can subsequently use to interact programmatically with the App. session = fo.launch\_app(dataset) ## Loading predictions into FiftyOne Similar to how you load ground truth labels into a FiftyOne Dataset, [loading model predictions](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html#model-predictions) is as easy as writing a Python loop. import fiftyone as fo dataset = fo.load\_dataset("dragons") img\_paths = \["/path/to/img1.png", ...\] \# Ex. custom prediction format: \[bbox, label, confidence\] predictions = \[\[\[0.1,0.2,0.3,0.5\], "car", 0.921\], ...\] for img\_path, img\_preds in zip(img\_paths, predictions): sample = dataset\[img\_path\] dets = \[\] for bbox, label, conf in img\_preds: dets.append( fo.Detection( bounding\_box=bbox, label=label, confidence=confidence, ) ) sample\["predictions"\] = fo.Detections(detections=dets) sample.save() \# View predictions in the App session = fo.launch\_app(dataset) Once the predictions are in FiftyOne, we can easily export them in [COCO format](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/datasets.html#cocodetectiondataset) into a JSON file on disk. \# Export predictions from FiftyOne dataset to disk in COCO-formatted JSON dataset.export( label\_field="predictions", label\_path="/path/to/coco\_predictions.json", dataset\_type=fo.types.COCODetectionDataset, ) ## Evaluating and Filtering Results Now that the model predictions are loaded, we can dig in and analyze the results. [FiftyOne provides methods for evaluating classification, detection, and segmentation](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html) models. While these methods can be used to compute dataset-wide metrics like so many other tools, the primary benefit is that this evaluation also populates instance-level results on the dataset like tagging individual true or false positive predictions. This allows you to not only understand how the model performs on the dataset as a whole but also specific instances in which the model performs well or poorly which is the best way to build intuition about the type of data you should use to retrain the model. Visualizing predictions in FiftyOne shows many low confidence incorrect predictions. This indicates that we should find an appropriate confidence threshold to limit the predictions in the dataset. One way to find a threshold value for detection confidence is to calculate the number of true and false positives that exist currently and find the point at which there are an equal number of both. Let’s call the [evalute\_detections()](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#detections) method to use COCO-style object detection evaluation to compute if each ground truth and prediction is either a true positive, false positive, or false negative. eval\_key = "full\_dataset\_eval" results = dataset.evaluate\_detections( "predictions", gt\_field="ground\_truth", eval\_key="full\_data\_eval", ) The FiftyOne API provides a [powerful query language](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#) that can be used to [filter and slice datasets](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#filtering) letting you look at the specific view in which you are interested. It also provides [dataset-wide aggregation functions](https://voxel51.com/docs/fiftyone/user_guide/using_aggregations.html) that let you easily access content from your datasets such as label values, counts, distributions, and ranges. One of these aggregations is the [histogram\_values()](https://voxel51.com/docs/fiftyone/user_guide/using_aggregations.html#histogram-values) function that is perfect for our use case of computing the number of true and false positives for each confidence bin. from fiftyone import ViewField as F import numpy as np import matplotlib.pyplot as plt \# Compute views of only True Positives and only False Positives tp\_view = dataset.filter\_labels("predictions", F(eval\_key) == "tp") fp\_view = dataset.filter\_labels("predictions", F(eval\_key) == "fp") \# Aggregate and plot histogram values tp\_counts, tp\_edges, other = tp\_view.histogram\_values("predictions.detections.confidence", bins=50) fp\_counts, fp\_edges, other = fp\_view.histogram\_values("predictions.detections.confidence", bins=50) plt.plot(tp\_edges\[:-1\]\[::-1\], np.cumsum(tp\_counts\[::-1\])) plt.plot(fp\_edges\[:-1\]\[::-1\], np.cumsum(fp\_counts\[::-1\])) plt.show() Based on this graph, we should set our confidence threshold to around 0.5. After browsing through the samples in the dataset, a threshold of 0.5 provides enough flexibility to detect many of the dragons in the dataset without too many false positives. Now to apply this threshold and rerun evaluation to compute mAP. high\_conf\_view = dataset.filter\_labels( "predictions", F("confidence") > 0.5, ) results = high\_conf\_view.evaluate\_detections( "predictions", gt\_field="ground\_truth", eval\_key="eval", compute\_mAP=True, ) print(results.mAP()) \# 0.3417 ## What can our failures teach us? One of the primary uses of FiftyOne is the ability to easily query and explore your dataset and model predictions for any question that comes to mind. An especially useful workflow is to explore the failure modes of your model to get a sense of how to best improve it going forward. For example, let’s take a look at all of the predictions that were false positives but with high confidence, indicating that the model was fairly certain about its detection but was incorrect. These types of examples usually indicate either an ingrained issue with the model or an error in the ground truth annotations. Either need to be addressed promptly. high\_conf\_fp = high\_conf\_view.filter\_labels( "predictions", (F(eval\_key) == "fp") & (F("confidence") > 0.9), ) \# Update App session.view = high\_conf\_fp From the example above, it seems that one issue with our dataset is that we did not consistently annotate the wings of dragons. The model relatively accurately detected the dragon, but also included the wing which resulted in an IoU below the threshold used for evaluation (IoU=0.5). This detection would not necessarily be incorrect, though, so we may want to take a pass over the dataset to ensure dragon wings are consistently annotated. An easy way to reannotate this dataset is to use the integrations between FiftyOne and annotation tools like [CVAT](https://voxel51.com/docs/fiftyone/integrations/cvat.html) or [Labelbox](https://voxel51.com/docs/fiftyone/integrations/labelbox.html). Now, let’s take a look at the false negatives in the dataset, where the model did not detect a ground truth object. fn\_view = high\_conf\_view.filter\_labels( "ground\_truth", F(eval\_key) == "fn", ) \# Update App session.view = fn\_view From the example above, we see another issue of the model incorrectly localizing the bounding box, even though it did detect the presence of a dragon. The comment about reannotating the dataset to include dragon wings still holds, however, it would also be useful to add additional training data to allow the model to learn to more accurately localize the boxes. In the example above, we see that the model is frequently detecting non-dragon objects as dragons. The majority of the samples in this dataset contain only one or a few dragons isolated from other objects. Thus, the model seems to be learning to just detect all of the focal objects in the scene. The best way to resolve this would be to add more scenes with multiple types of objects to the dataset as well as expand the classes to other object types so that the model is able to learn to better differentiate between dragons and other objects. ## Next Steps Based on the observations in the previous sections, we have a plan of action for producing a higher-quality dataset and a higher-performing model. - **Update the dataset annotations taking into account the wings** - **Incorporate augmentations into the training loop** - **Add more difficult samples like crowds of objects** - **Automate and parameterize the training loop for fast retraining iterations whenever we update the data** ## Summary Creating a high-performing model requires much more than just some PyTorch code. Being able to iteratively track and analyze model performance and then use that to inform dataset improvements is necessary for a high-quality model. The combination of the model analysis capabilities of [FiftyOne](https://fiftyone.ai/) with the experiment tracking capabilities of [ClearML](https://clear.ml/) results in a system that will lead to better models, faster. _This post was made in collaboration between the teams at [ClearML](https://clear.ml/) and [Voxel51](https://voxel51.com/) and co-authored by [Victor Sonck](https://medium.com/@victor.sonck_78979)._ [ClearML](https://voxel51.com/blog/tag/clearml) [deep learning](https://voxel51.com/blog/tag/deep-learning) [experiment tracking](https://voxel51.com/blog/tag/experiment-tracking) [object detection](https://voxel51.com/blog/tag/object-detection) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/98e839c6e81bb9c4fa3ad96bf0d5d1b77ee11f6c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 1)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1) [![](https://cdn.sanity.io/images/h6toihm1/production/8cb892c18f65a83d75021c76f891400ac219dcef-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The ML Menu for Model Selection: Hugging Face, Weights & Biases, and FiftyOne\\ \\ Tutorials\\ \\ • \\ \\ May 1, 2023](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/047b21a97f6c858334f9f35ed89fa7655ebf5767-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ State-of-the-Art Object Detection with YOLO-NAS & FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ May 4, 2023](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-231-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023 Jan 14, 2023 • 5 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#de48b48e7e2f) [Splitting data in FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#80fb241a0fdc) [Coloring by tags in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#ed662e357688) [Dealing with duplicate data](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#aa7cc40d4062) [Merging datasets in COCO format](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#0d97981148de) [Hiding labels in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#9ee7e37cbb4b) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#b6313177b4f7) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#de48b48e7e2f) [Splitting data in FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#80fb241a0fdc) [Coloring by tags in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#ed662e357688) [Dealing with duplicate data](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#aa7cc40d4062) [Merging datasets in COCO format](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#0d97981148de) [Hiding labels in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#9ee7e37cbb4b) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023#b6313177b4f7) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=max&q=75&w=1200) ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **Splitting data in FiftyOne** Community Slack member Muhammad Ali asked, _“Can we separate a dataset into train, test, and validation splits in FiftyOne?”_ If you want to randomly split the data, and are okay with the splits being slightly larger or smaller than the specified percentages, then you can do so with the `random_split()` method in FiftyOne’s random utils. To create train, test, and validation splits with approximately 70%, 20%, and 10% of the data respectively, you can do the following: ```python 1import fiftyone.utils.random as four 2four.random_split( 3 dataset, 4 {"train": 0.7, "test": 0.2, "val": 0.1} 5) ``` The information about which split a sample is in can be found in the `tags` field, and you can use `match_tags()` to get a view containing only the samples in each split: ```python 1train_view = dataset.match_tags("train") 2test_view = dataset.match_tags("test") 3val_view = dataset.match_tags("val") ``` In the FiftyOne App, you can achieve the same effect by [selecting or deselecting any of these tags](https://voxel51.com/docs/fiftyone/user_guide/app.html#tags-and-tagging). If you need more precise splitting, you can compute the number of samples that should be in each split, and index into a shuffled view of the data: ```python 1nsample = dataset.count() 2ntrain = int(nsample * 0.7) 3ntest = int(nsample * 0.2) 4nval = nsample - (ntrain + ntest) 5 6shuffled_view = dataset.shuffle() 7 8train_view = shuffled_view[:ntrain] 9test_view = shuffled_view[ntrain:ntrain+ntest] 10val_view = shuffled_view[-nval:] 11 12for sample in train_view.iter_samples(autosave = True): 13 sample.tags.append("train") 14 15for sample in test_view.iter_samples(autosave = True): 16 sample.tags.append("test") 17 18for sample in val_view.iter_samples(autosave = True): 19 sample.tags.append("val") ``` Learn more about FiftyOne’s [random utils](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.random.html) in the FiftyOne Docs. ## **Coloring by tags in the FiftyOne App** Community Slack member Adrian Tofting asked, _“I have a use case where it would be helpful to color embeddings by tags in the FiftyOne App. Is this possible?”_ By combining the FiftyOne Brain’s `visualize()` method with the [ViewField](https://voxel51.com/docs/fiftyone/api/fiftyone.core.expressions.html#fiftyone.core.expressions.ViewField), you can create your own view expressions that allow you to color by a variety of properties. The syntax is both easy to use and very flexible. If you want to color points in an embedding visualization by the last tag found on each sample, you can do this with a single line of code: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.utils.random as four 4import fiftyone.zoo as foz 5from fiftyone import ViewField as F 6 7dataset = foz.load_zoo_dataset("quickstart") 8## add tags to dataset 9four.random_split(dataset, {"train": 0.7, "test": 0.2, "val": 0.1}) 10 11results = fob.compute_visualization(dataset, method = "tsne") 12 13### Color by last tag 14plot = results.visualize(labels=F("tags")[-1]) 15plot.show() ``` ![](https://cdn.sanity.io/images/h6toihm1/production/ae29b6570a2a720424c400b9fc90e86401a2740f-2356x706.png?auto=format&dpr=2&fit=max&q=75&w=1600) If you instead wanted to color by number of ground truth objects, you could use the following: ```python 1plot = results.visualize( 2 labels=F("ground_truth.detections").length() 3) 4plot.show() ``` Learn more about [visualizing with the FiftyOne Brain](https://voxel51.com/docs/fiftyone/api/fiftyone.brain.visualization.html#fiftyone.brain.visualization.VisualizationResults.visualize) in the FiftyOne Docs. ## **Dealing with duplicate data** Community Slack member George Pearse asked, _“Is there a safe way to avoid adding duplicate samples to a given dataset? What happens when you add the same collection of samples to a dataset twice with `dataset.add_samples(samples)`”_ Duplicate data can appear in many scenarios in computer vision workflows — sometimes intentionally, and other times erroneously. In your case, when you ran _`dataset.add_samples(samples)`_ the second time, the samples that were added to the dataset were identical to those added in all but one crucial way — they were given unique `id` s. This is because in FiftyOne, all samples are required to have unique sample ids. This means that if you were to use `dataset.delete_samples(samples)`, you would only be removing the original samples from the dataset, while the copied samples with new sample ids would remain. If you wanted to delete both the original and copied samples from the dataset, you could do so by first finding all samples with a given file path, and then deleting these. ```python 1for sample in samples: 2 double = dataset.match( 3 F("filepath") == sample.filepath 4 ) 5 dataset.delete_samples(double) ``` Of course, this assumes that only the added samples and their copies shared a common file path. If you want to identify and remove samples which have duplicated images stored in different locations, you can use the image deduplication procedure documented [here](https://voxel51.com/docs/fiftyone/recipes/image_deduplication.html). Learn more about the [adding](https://voxel51.com/docs/fiftyone/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.add_samples) and [deleting samples](https://voxel51.com/docs/fiftyone/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.delete_samples) in FiftyOne datasets in the FiftyOne Docs. ## **Merging datasets in COCO format** Community Slack member Dan Erez asked, _“Let’s say I have two different datasets with different classes, and both are in COCO format. Is there a way to concatenate them into a third dataset, also in COCO format, which contains all of the classes present in each?”_ Fortunately, FiftyOne supports all of the functionality necessary to perform this operation! One way to approach this would be to load in the first dataset in COCO format using FiftyOne’s [COCODetectionDatasetImporter](https://voxel51.com/docs/fiftyone/integrations/coco.html) with the `from_dir()` method: ```python 1import fiftyone as fo 2 3dataset = fo.Dataset.from_dir( 4 data_path="/path/to/images", 5 labels_path="/path/to/coco1.json", 6 dataset_type=fo.types.COCODetectionDataset, 7) ``` Then merge in the data from the second dataset using the `merge_dir()` method: ```python 1dataset.merge_dir( 2 data_path="/path/to/images", 3 labels_path="/path/to/coco2.json", 4 dataset_type=fo.types.COCODetectionDataset, 5) ``` Finally, you can use FiftyOne’s dataset export functionality to export the composite dataset in COCO format: ```python 1dataset.export( 2 labels_path="/path/for/coco3.json", 3 dataset_type=fo.types.COCODetectionDataset, 4) ``` Learn more about the [importing](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/datasets.html), [exporting](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html), and [merging datasets](https://voxel51.com/docs/fiftyone/recipes/merge_datasets.html) in the FiftyOne Docs. ## **Hiding labels in the FiftyOne App** Community Slack member Santiago Arias asked, _“Does anyone know if there is a way to omit some labels from a view in the FiftyOne App?”_ Absolutely! There are many reasons for doing this, all of which involve retaining information while making it easier to understand your computer vision data. You can do so using either `select_fields()` to explicitly specify which fields to view, or using `exclude_fields()` to specify which fields should be omitted. In both cases, reducing the number of fields displayed in the app can make your visualization, analysis, and evaluation workflows seamless. One example where this approach might be useful is when dealing with versatile datasets. [Open Images V6](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#dataset-zoo-open-images-v6), for instance, supports a variety of computer vision tasks, including detection, classification, relationships, and segmentations. To focus our present attention on only segmentations while retaining the other labels, we can employ `select_fields()`: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset( 5 "open-images-v7", 6 split="validation", 7 label_types = ["detections", "classifications", "segmentations"], 8 max_samples=50, 9 shuffle=True, 10 ) 11 12segmentation_view = dataset.select_fields("segmentations") 13 14session = fo.launch_app(dataset) 15session.view = segmentation_view.view() ``` Alternatively, if you’re testing a bunch of models on a single dataset, and you’ve added these predictions to the dataset, you may want to create a view containing some, but not all of these model predictions. In this case, you can `exclude_fields()` to exclude predictions you are not concerned with: ```python 1models = [model1, model2, model3, ...] 2for model in models: 3 dataset.apply_model( 4 model, 5 label_field=model.name, 6 confidence_thresh=0.5 7 ) 8 9model_to_exclude = model2 10 11view = dataset.exclude_fields(model2.name) 12session = fo.launch_app(dataset) 13session.view = view.view() ``` Learn more about [selecting](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.select_fields) and [excluding fields](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.exclude_fields) in the FiftyOne Docs. ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,275+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,400+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,500+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 224+ repositories - 54+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [embeddings](https://voxel51.com/blog/tag/embeddings) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [labels](https://voxel51.com/blog/tag/labels) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/4ac1a727dc192a21563cde51b6e345f620e09376-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0fc50e89593dec2117ce5c1934761cf61b24f8c4-1920x1080.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Tips and Tricks for Accelerating Computer Vision Workflows – Mar 17, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 18, 2023](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/dc8a2e7a894316856af5a109ae43f8787959f179-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Embeddings Tips and Tricks – Mar 31, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 31, 2023](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-232-lllmstxt|> ## Download ActivityNet Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) How to Download ActivityNet and Evaluate Video Understanding Models Feb 8, 2022 • 4 min read Article content In this article [Setup](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#df3156d03730) [The ActivityNet Dataset](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#61d9d1e5752f) [Downloading the Dataset](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#0fb6d61886cd) [ActivityNet Model Evaluation](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#f40433609e8d) [Adding Model Predictions to Dataset](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#6ebc5a933fd3) [Computing mAP](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#547bb3ff79be) [Analyzing Results](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#acd21da6d318) [Summary](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#139282b92968) In this article [Setup](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#df3156d03730) [The ActivityNet Dataset](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#61d9d1e5752f) [Downloading the Dataset](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#0fb6d61886cd) [ActivityNet Model Evaluation](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#f40433609e8d) [Adding Model Predictions to Dataset](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#6ebc5a933fd3) [Computing mAP](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#547bb3ff79be) [Analyzing Results](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#acd21da6d318) [Summary](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models#139282b92968) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _A guide to downloading, visualizing, and evaluating models on the ActivityNet dataset using FiftyOne_ Year over year, the [ActivityNet challenge](http://activity-net.org/challenges/2021/index.html) has pushed the boundaries of what video understanding models are capable of. Training a machine learning model to classify video clips and detect actions performed in videos is no easy task. Thankfully, the teams behind datasets like [Kinetics](https://deepmind.com/research/open-source/kinetics) and [ActivityNet](http://activity-net.org/) have provided massive amounts of video data to the computer vision community to assist in building high-quality video models. We’re [announcing a partnership with the ActivityNet team](http://activity-net.org/download.html) to make the [ActivityNet datasets natively available](https://voxel51.com/docs/fiftyone/integrations/activitynet.html) in the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html) letting you access and visualize the dataset more easily than ever before as well as evaluate and analyze models trained on the dataset. [FiftyOne](http://fiftyone.ai/) is an open-source tool designed for dataset curation and model analysis. Downloading a subset of ActivityNet is now as easy as: ## Setup To run the examples in this post, you need to [install FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html). ## The ActivityNet Dataset There are two popular versions of ActivityNet, one with 100 activity classes and a newer version with 200 activity classes. These versions contain 9682 and 19994 videos respectively across their training, testing, and validation splits with over 8k and 23k labeled activity instances. These labels are temporal activity detections which are each represented by a start and stop time for the segment as well as a label. ## Downloading the Dataset Downloading the entirety of the ActivityNet requires filling [out this form](https://docs.google.com/forms/d/e/1FAIpQLSeKaFq9ZfcmZ7W0B0PbEhfbTHY41GeEgwsa7WobJgGUhn4DTQ/viewform) to gain access to the dataset. However, many use cases involve specific subsets of ActivityNet. For example, you may be training a model specifically to detect the class “Bathing dog”. The integration between FiftyOne and ActivityNet means that the dataset can now be accessed through the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html). Additionally, it makes it easier than ever to download specific subsets of ActivityNet directly from YouTube. You can download all samples of the class “Bathing dog” from the above example like so: Other useful parameters include the ability to specify the dataset `split` to download, `max_samples` if you are only interested in subsets of the dataset, `max_duration` to define the maximum length of videos you want to download and more. If you are working with the full ActivityNet dataset, you can use the `source_dir` parameter to point to the location of the dataset on disk and easily load it into FiftyOne to visualize it and analyze your models. ## ActivityNet Model Evaluation In this section, we cover how to add your custom model predictions to a FiftyOne dataset and evaluate temporal detections following the [ActivityNet evaluation protocol](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#activitynet-style-evaluation-default-temporal). ## Adding Model Predictions to Dataset Any custom labels and metadata can be easily added to a FiftyOne dataset. In this case, we need to populate a `TemporalDetection` field named `predictions` with the model predictions for each sample. We can then visualize the predictions in the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html). ## Computing mAP The [ActivityNet evaluation protocol](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#activitynet-style-evaluation-default-temporal) used to evaluate temporal activity detection models trained on the dataset is similar to the object detection evaluation [protocol for the COCO dataset](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#coco-style-evaluation-default-spatial). Both involve computing the mean average precision of detections, either temporal or spatial, across the dataset by matching predictions with ground truth annotations. The specifics of the ActivityNet mAP computation can be [found here](https://voxel51.com/docs/fiftyone/integrations/activitynet.html#map-protocol) as it has been reimplemented in FiftyOne. After having added your predictions to the FiftyOne dataset, you can call the [`evaluate_detections()`](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#detections) method. The evaluation results object can also be used to compute PR curves and [confusion matrices](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#confusion-matrices) of your model’s performance. Additional metrics can be computed using the [DETAD tool](https://github.com/HumamAlwassel/DETAD) introduced to diagnose the performance of action detection models on ActivityNet and THUMOS14. A follow-up post will explore how to incorporate DETAD analysis into our FiftyOne workflow to analyze models trained on ActivityNet. ## Analyzing Results One of the primary benefits of the FiftyOne implementation of the ActivityNet evaluation protocol is that it stores not only dataset-wide metrics like mAP but also individual label-level results. Specifying the `eval_key` parameter when calling `evaluate_detections()` will populate fields on our dataset containing the individual sample- and label-level results of whether a predicted or ground truth temporal detection was a true positive, false positive, or false negative. Using the [powerful querying capabilities](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) of FiftyOne, we can really dig into the model results and explore specific cases of where the model performed well or poorly. This analysis quickly informs us of the type of data that the model struggles with that we need to incorporate more into the training scheme. For example, in one line we can find all predictions that had high confidence but were evaluated as false positives. From here you can use the `to_clips()` method to convert the dataset of full videos to a [view of only the clips](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#clip-views) of videos according to the annotations of a field like our `TemporalDetection` predictions. The samples in these evaluation views are videos that the model was not able to correctly understand. One way to resolve this is to train on more samples similar to those that the model failed on. If you are interested in using this workflow in your own projects, check out the `compute_similarity()` method of the [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html#brain-similarity). ## Summary [ActivityNet](http://activity-net.org/download.html) is one of the leading video datasets in the computer vision community that is included in the popular yearly ActivityNet challenge. The [integration](https://voxel51.com/docs/fiftyone/integrations/activitynet.html) between ActivityNet and [FiftyOne](http://fiftyone.ai/) makes it easier than ever to access the dataset, visualize it in the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html), and analyze your model performance to figure out how to build better activity classification and detection models. [ActivityNet](https://voxel51.com/blog/tag/activitynet) [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [model evaluation](https://voxel51.com/blog/tag/model-evaluation) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/33d08c7b16ab5bfa0e4c5a4f936be4592b8e0a90-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the Berkeley Deep Drive Autonomous Vehicle Dataset\\ \\ Datasets\\ \\ • \\ \\ Jan 11, 2023](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/7d8209de12d18f950c73a8d2e335f4a83bc6ce71-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the UCF101 Dataset: A Large-Scale, YouTube-Based Action Recognition Dataset\\ \\ Datasets\\ \\ • \\ \\ Mar 1, 2023](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/40381f5f37fa5fcd70eddca2f63b6710568f5d2c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual Kinship Recognition with the Families in the Wild Computer Vision Dataset\\ \\ Datasets\\ \\ • \\ \\ Dec 7, 2022](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-233-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision View Stages Tips and Tricks – Jan 20, 2023 Jan 21, 2023 • 3 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#05d39f451c47) [A primer on view stages](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#8cacf81b7661) [See what view stages have been applied](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#5b30f596f831) [Listing available view stages](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#668f0a77c7c6) [Add a stage to a view](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#39fbef85f962) [Create a stage with MongoDB](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#f97d720b62a9) [Check for conversions](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#1e5f2ec0c080) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#b56dc66f8b3a) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#05d39f451c47) [A primer on view stages](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#8cacf81b7661) [See what view stages have been applied](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#5b30f596f831) [Listing available view stages](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#668f0a77c7c6) [Add a stage to a view](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#39fbef85f962) [Create a stage with MongoDB](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#f97d720b62a9) [Check for conversions](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#1e5f2ec0c080) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023#b56dc66f8b3a) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we give practical pointers for using FiftyOne on topics inspired by discussions in the open source community. This week we’ll cover [view stages](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#view-stages). ## **Wait, What’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **A primer on view stages** In FiftyOne, a [ViewStage](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#view-stages) represents a stage of logic such as matching, filtering, or slicing, which can be used to determine the samples that appear in a dataset view. These stages can be chained together to produce flexible logical pipelines, enabling the programmatic isolation and evaluation of subsets of computer vision data. Every view stage is exposed via an associated method on Dataset and DatasetView instances. Continue reading for some tips and tricks to help you master view stages in FiftyOne! ## **See what view stages have been applied** Many computer vision workflows involve chaining view stages together to create several tailored views of a dataset. If you ever lose track of what logical operations have been performed to generate a DatasetView, you can get a summary of this information by printing out the dataset view. As an example, we can chain together three stages of filtering labels, matching, and converting to patches: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5dataset = foz.load_zoo_dataset("quickstart") 6dataset.evaluate_detections("predictions", eval_key="eval") 7 8complex_view = ( 9 dataset 10 .filter_labels( 11 "predictions", ( 12 (F("confidence") > 0.3) 13 & ((F("bounding_box")[2] * F("bounding_box")[3]) > 0.3) 14 & (F("eval") == "fp") 15 ) 16 ).match( 17 F("predictions.detections").length() > 2 18 ).to_patches("predictions") 19) 20 21print(complex_view) 22 ``` At the bottom of the print out, we can see the three view stages summarized in order. ![](https://cdn.sanity.io/images/h6toihm1/production/b5bcd8e2bd9e798de8225920f1d50bae75b2e5fc-2006x556.png?auto=format&dpr=2&fit=max&q=75&w=1600) Learn more about the [ToPatches](https://voxel51.com/docs/fiftyone/api/fiftyone.core.stages.html#fiftyone.core.stages.ToPatches) view stage in the FiftyOne Docs. ## **Listing available view stages** If you know you want to perform logical operations on your dataset to generate a new view, but aren’t sure exactly how to do so, use the `list_view_stages()` method to refresh your memory on all of the possibilities. You can then use the `'?'` character to print out a view stage’s intended usage as well as some examples. For example, running `dataset.concat?` returns the following: ![](https://cdn.sanity.io/images/h6toihm1/production/c3d563f78885fe7c63e8f42ee8e70a3b147f0e83-1968x1318.png?auto=format&dpr=2&fit=max&q=75&w=1600) Learn more about [available view stages](https://voxel51.com/docs/fiftyone/api/fiftyone.core.stages.html#module-fiftyone.core.stages) in the FiftyOne Docs. ## **Add a stage to a view** While every stage is exposed via an associated method on datasets and dataset views, you can also create a stage on its own. For example, instead of generating a view without predictions using the `exclude_fields()` method, we can create the corresponding `ExcludeFields` stage, and _then_ apply this logic to the dataset using the `add_stage()` method. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# load in dataset 5dataset = foz.load_zoo_dataset("quickstart") 6 7# use dataset method directly 8view = dataset.exclude_fields(“predictions”) 9 10# alternative: create and add stage 11stage = fo.ExcludeFields("predictions") 12view = dataset.add_stage(stage) ``` If we want to apply the same logic to multiple datasets or dataset views, it can be easier and less error-prone to define the stage separately, rather than reproducing identical logic in multiple places. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4stage = fo.FilterLabels( 5 "ground_truth", 6 F("label").is_in(["cat", "dog"]) 7) 8 9### apply the cat/dog filter to multiple existing views 10cat_dog_view1 = view1.add_stage(stage) 11cat_dog_view2 = view2.add_stage(stage) 12cat_dog_view3 = view3.add_stage(stage) ``` Learn more about [FilterLabels](https://voxel51.com/docs/fiftyone/api/fiftyone.core.stages.html#fiftyone.core.stages.FilterLabels) in the FiftyOne Docs. ## **Create a stage with MongoDB** If you feel more comfortable working with MongoDB’s query language, you can code up your desired logic in raw [MongoDB aggregation pipeline syntax](https://www.mongodb.com/docs/manual/core/aggregation-pipeline/) and generate a view stage from this using the `Mongo` view stage, or the associated `mongo()` method for datasets and dataset views. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# load in dataset 5dataset = foz.load_zoo_dataset("quickstart") 6 7# Stage for getting the second and third samples in the dataset 8view = dataset.mongo([{"$skip": 1}, {"$limit": 2}]) 9 10## this is equivalent to: 11view = dataset.skip(1).limit(2) ``` This is all possible because FiftyOne uses MongoDB to efficiently store and access datasets. Dive into the details of [FiftyOne’s MongoDB backend](https://voxel51.com/docs/fiftyone/user_guide/config.html#configuring-a-mongodb-connection) in the FiftyOne Docs. ## **Check for conversions** Some of FiftyOne’s view stages, including `ToFrames` and `ToPatches` generate views which differ from the datasets or views to which they are applied not only in content, but also in structure. In the case of `ToFrames`, the resulting view consisting of images is generated from an initial dataset or view consisting of videos: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5dataset = foz.load_zoo_dataset("quickstart-video") 6 7# Create a frames view for an entire video dataset 8frames = dataset.to_frames(sample_frames=True) 9 ``` ![](https://cdn.sanity.io/images/h6toihm1/production/024913c3c097f4d7b5037aeb25fa3c5466b0689b-2302x1280.png?auto=format&dpr=2&fit=max&q=75&w=1600) As another example, applying the `SelectGroupSlices` stage to a Grouped Dataset results in a view whose media type is _not_ `group`. When in question, it is always a good idea to validate the anticipated properties of your generated views by looking at the `media_type` and other attributes. Learn more about [SelectGroupSlices](https://voxel51.com/docs/fiftyone/api/fiftyone.core.stages.html#fiftyone.core.stages.SelectGroupSlices) and [Grouped Datasets](https://voxel51.com/docs/fiftyone/user_guide/groups.html) in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,275+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,400+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,500+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 222+ repositories - 54+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [MongoDB](https://voxel51.com/blog/tag/mongodb) [view stages](https://voxel51.com/blog/tag/view-stages) [ViewStage](https://voxel51.com/blog/tag/viewstage) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4ac1a727dc192a21563cde51b6e345f620e09376-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/33c550b5b632dcb569967bca8130dc1f7b9df017-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Model Evaluation Tips and Tricks – Feb 03, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Feb 4, 2023](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-234-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023 Jan 28, 2023 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#48912321fc40) [Sorting samples by filepath](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#befcfc1bbb58) [Viewing images with many small objects](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#3ccb49773ce2) [Assigning the same CVAT job to multiple users](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#e10451dca9b0) [Converting instance segmentation mask to full image mask](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#f6bee5a28ddc) [Viewing in Jupyter notebooks](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#39a9afed9f00) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#07244ff685f3) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#48912321fc40) [Sorting samples by filepath](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#befcfc1bbb58) [Viewing images with many small objects](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#3ccb49773ce2) [Assigning the same CVAT job to multiple users](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#e10451dca9b0) [Converting instance segmentation mask to full image mask](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#f6bee5a28ddc) [Viewing in Jupyter notebooks](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#39a9afed9f00) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023#07244ff685f3) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **Sorting samples by filepath** Community Slack member Oğuz Hanoğlu asked, _“Is it possible to have the samples in my dataset appear sorted by their filepath?”_ The samples in a FiftyOne `Dataset` are always enumerated in the same order in which they were added to the `Dataset`. However, if you want to create a _new_ dataset where the samples appear in the desired order, you can use the `sort_by()` method, passing in the `filepath` field. This creates a reindexed view into the original dataset. You can then use the `clone()` method to create a new dataset from this view. As the samples appear in the desired order in the view, they will be inserted into the new dataset in this order as well. Putting it all together, your workflow may look like this: Using the same idea, you could similarly sort by other properties of the samples in the dataset, such as `id`. With the `ViewField` to handle expressions, you can even sort by the number of ground truth object detections in the sample: Learn more about [sort\_by()](https://voxel51.com/docs/fiftyone/api/fiftyone.core.view.html#fiftyone.core.view.DatasetView.sort_by) and [DatasetViews](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#datasetviews) in the FiftyOne Docs. ## **Viewing images with many small objects** Community Slack member Dan Erez asked, _“Is there a way to not show the labels for the detections? I have an image with a lot of little detection bounding boxes, and it's hard to see.”_ There are a few ways you might be able to work around this. First, if you have a lot of overlapping bounding boxes, you can use FiftyOne’s IoU utils to [identify and remove duplicate objects](https://voxel51.com/docs/fiftyone/recipes/remove_duplicate_annos.html#Finding-duplicate-objects). In a similar vein, if a lot of your predictions are low confidence, you can reduce clutter by filtering for objects with confidence values about some threshold. You can also reduce clutter by setting attributes in your [FiftyOne App config](https://voxel51.com/docs/fiftyone/user_guide/config.html#configuring-the-app). For instance, you can set `show_confidence=False` and `show_label=False` to get rid of all the text above the bounding boxes. Additionally, you can set `color_by = “label” ` to have the color of each bounding box set by the label class of the object. This may help distinguishing classes for smaller objects. To make these changes in Python, you can run the following: Learn more about [configuring FiftyOne](https://voxel51.com/docs/fiftyone/user_guide/config.html#) in the FiftyOne Docs. ## **Assigning the same CVAT job to multiple users** Community Slack member Bahram Marami asked, _“If I want the same dataset to be annotated by two users with FiftyOne’s CVAT integration, how can I create jobs for two users and retrieve/load annotations individually for them?”_ Making a single call to FiftyOne’s annotation API with multiple users listed will result in each job only being assigned to one of the listed users. For instance, the snippet below will assign the CVAT job to either user 1 or user 2. If you would like for each annotation task to be performed by both users, then one way to achieve this is by creating separate fields in the dataset for each user, and publishing the jobs twice, as in the code below. Here we have cloned the ground truth field twice - once for each user, and have made separate API calls assigning the jobs for each. Learn more about [FiftyOne’s annotation API](https://voxel51.com/docs/fiftyone/user_guide/annotation.html) in the FiftyOne Docs. ## **Converting instance segmentation mask to full image mask** Community Slack member Onuralp Sezer asked, _“How do I use the instance segmentation masks in FiftyOne to generate analogous masks that span the entire image?”_ In FiftyOne, instance segmentations are represented with a bounding box and a two dimensional array. The bounding box, which is in the format `[top-left-x, top-left-y, width, height]`, specifies what portion of the image the grid lies on, and the two dimensional array specifies which pixels within that portion of the image are part of the object. Converting from this into full image scale instance segmentation masks requires using image metadata (width and height) and bounding box coordinates to place the object’s mask grid onto the appropriate pixels in the image. However, you can use FiftyOne’s utilities to handle these details for you! If all you require is semantic segmentation masks, then you can use the FiftyOne utils method `objects_to_segmentations()` out of the box. For instance, to generate a semantic segmentation mask for COCO 2014 samples, accounting for dogs and cats, the following code suffices: To generate full image instance segmentation masks with this method, you have to get a bit craftier. You can iterate through the objects in a given image, filtering for each one individually, creating a view for this single image, single object pair, and using the `objects_to_segmentations()` method. Learn more about [`objects_to_segmentations()`](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.labels.html?highlight=objects_to_segmentations#fiftyone.utils.labels.objects_to_segmentations) and [label type coercion](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html#label-type-coercion) in the FiftyOne Docs. ## **Viewing in Jupyter notebooks** Community Slack member Adrian Tofting asked, _“Does anyone know how I can make a Jupyter notebook show the FiftyOne output cell view in full height? Mine is cropped on the bottom”_ Absolutely! You can specify this directly in the `launch_app()` command when you launch the FiftyOne App in a Jupyter notebook with the `height` argument. To set the height to 600 pixels, for example, the following would work: Learn more about [running FiftyOne in a notebook](https://voxel51.com/docs/fiftyone/environments/index.html#notebooks) in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,275+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,450+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,750+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 231+ repositories - 55+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [CVAT](https://voxel51.com/blog/tag/cvat) [data annotation](https://voxel51.com/blog/tag/data-annotation) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Jupyter Notebook](https://voxel51.com/blog/tag/jupyter-notebook) [labels](https://voxel51.com/blog/tag/labels) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/a3a918e30b0553723b9392ea90763379f98480a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – April 7, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Apr 7, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/79d00d175a8098516cb2f4a7711131cbf322d01a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Finding and Correcting Mistakes – FiftyOne Tips and Tricks – Aug 18, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Aug 18, 2023](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-235-lllmstxt|> ## FiftyOne Evaluation Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Model Evaluation Tips and Tricks – Feb 03, 2023 Feb 4, 2023 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#2f8596346127) [A primer on model evaluations](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#9ff8ec64d0c8) [Task-specific evaluation methods](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#7b53f7ae027a) [Evaluations on views](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#3be394f92d22) [Plotting interactive confusion matrices](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#9d6b67096222) [Evaluating frames of a video](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#cc1c2188de68) [Managing multiple evaluations](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#81481a9f0ec9) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#1df7f44f83ef) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#2f8596346127) [A primer on model evaluations](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#9ff8ec64d0c8) [Task-specific evaluation methods](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#7b53f7ae027a) [Evaluations on views](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#3be394f92d22) [Plotting interactive confusion matrices](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#9d6b67096222) [Evaluating frames of a video](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#cc1c2188de68) [Managing multiple evaluations](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#81481a9f0ec9) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023#1df7f44f83ef) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we give practical pointers for using FiftyOne on topics inspired by discussions in the open source community. This week we’ll cover [model evaluation](https://docs.voxel51.com/user_guide/evaluation.html). ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **A primer on model evaluations** FiftyOne provides a variety of builtin methods for evaluating your model predictions, including regressions, classifications, detections, polygons, instance and semantic segmentations, on both image and video datasets. When you evaluate a model in FiftyOne, you get access to the standard [aggregate metrics](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#aggregate-metrics) such as classification reports, [confusion matrices](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#id11), and [PR curves](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#map-and-pr-curves) for your model. In addition, FiftyOne can also record fine-grained statistics like accuracy and false positive counts at the sample-level, which you can leverage via dataset views and the FiftyOne App to interactively explore the strengths and weaknesses of your models on individual data samples. FiftyOne’s model evaluation methods are conveniently exposed as methods on all `Dataset` and `DatasetView` objects, which means that you can evaluate entire datasets or specific views into them via the same syntax. Continue reading for some tips and tricks to help you master evaluations in FiftyOne! ## **Task-specific evaluation methods** In FiftyOne, the Evaluation API supports common computer vision tasks like object detection and classification with default evaluation methods that implement some of the standard routines in the field. For standard object detection, for instance, the default evaluation style is MS COCO. In most other cases, the default evaluation style is denoted `"simple"`. If the default style for a given task is what you are looking for, then there is no need to specify the `method` argument. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3dataset = foz.load_zoo_dataset("quickstart") 4results = dataset.evaluate_detections( 5 "predictions", 6 gt_field = "ground_truth" 7) ``` Alternatively, you can explicitly specify a method to use for model evaluation: ```python 1dataset.evaluate_detections( 2 "predictions", 3 gt_field = "ground_truth", 4 method = "open-images" 5) ``` Each evaluation method has an associated evaluation config, which specifies what arguments can be passed into the evaluation routine when using that style of evaluation. For [ActivityNet style evaluation](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.eval.activitynet.html#fiftyone.utils.eval.activitynet.ActivityNetEvaluationConfig), for example, you can pass in an `iou` argument specifying the IoU threshold to use, and you can pass in `compute_mAP = True` to tell the method to compute the mean average precision. To see which label types are available for a dataset, check out the section detailing that dataset in the FiftyOne Dataset Zoo documentation. Learn more about [evaluating object detections](https://voxel51.com/docs/fiftyone/tutorials/evaluate_detections.html) in the FiftyOne Docs. ## **Evaluations on views** All methods in FiftyOne’s Evaluation API that are applicable to `Dataset` instances are also exposed to `DatasetView`. This means that you can compute evaluations on subsets of your dataset obtained through filtering, matching, and chaining together any number of view stages. As an example, we can evaluate detections only on samples that are highly [unique](https://voxel51.com/docs/fiftyone/tutorials/uniqueness.html) in our dataset, and which have fewer than 10 predicted detections: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4from fiftyone import ViewField as F 5 6## compute uniqueness of each sample 7fob.compute_uniqueness(dataset) 8 9dataset = foz.load_zoo_dataset("quickstart") 10 11## create DatasetView with 50 most unique images 12unique_view = dataset.sort_by( 13 "uniqueness", 14 reverse=True 15).limit(50) 16 17## get only the unique images with fewer than 10 predicted detections 18few_pred_unique_view = unique_view.match( 19 F("predictions.detections").length() < 10 20) 21 22## evaluate detections for this view 23few_pred_unique_view.evaluate_detections( 24 "predictions", 25 gt_field="ground_truth", 26 eval_key="eval_few_unique" 27) ``` Learn more about the [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html) in the FiftyOne Docs. ## **Plotting interactive confusion matrices** For classification and detection evaluations, FiftyOne’s evaluation routines generate [confusion matrices](https://en.wikipedia.org/wiki/Confusion_matrix). You can plot these confusion matrices in FiftyOne with the `plot_confusion_matrix()` method. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("quickstart") 5 6## generate evaluation results 7results = dataset.evaluate_detections( 8 "predictions", 9 gt_field = "ground_truth" 10) 11 12## plot confusion matrix 13classes = ["person", "kite", "car", "bird"] 14plot = results.plot_confusion_matrix(classes=classes) 15plot.show() ``` Because the confusion matrix is implemented in [plotly](https://plotly.com/python/), it is interactive! To interact visually with your data via the confusion matrix, attach the plot to a session launched with the dataset: ```python 1## create a session and attach plot 2session = fo.launch_app(dataset) 3session.plots.attach(plot) ``` Clicking into a cell in the confusion matrix then changes which samples appear in the sample grid in the FiftyOne App. ![](https://cdn.sanity.io/images/h6toihm1/production/f151096ff8d9e05c867f2eb7f8e47519d77bff97-1870x856.gif?auto=format&dpr=2&fit=max&q=75&w=1600) Learn more about [interactive plotting](https://voxel51.com/docs/fiftyone/user_guide/plots.html) in the FiftyOne Docs. ## **Evaluating frames of a video** All of the evaluation methods in FiftyOne’s Evaluation API can be applied to frame-level labels in addition to sample-level labels. This means that you can evaluate video samples without needing to convert the frames of a video sample to standalone image samples. Applying FiftyOne evaluation methods to video frames also has the added benefit that useful statistics are computed at both the frame and sample levels. For instance, the following code populates the fields `eval_tp`, `eval_fp`, and `eval_fn` as summary statistics on the sample level, containing the total number of true positives, false positives, and false negatives across all frames in the sample. Additionally, on each frame, the evaluation populates an `eval` field for each detection with a value of either `tp`, `fp`, or `fn`, as well as an `eval_iou` field where appropriate. ```python 1import random 2 3import fiftyone as fo 4import fiftyone.zoo as foz 5 6dataset = foz.load_zoo_dataset( 7 "quickstart-video", 8 dataset_name="video-eval-demo" 9) 10 11## Create some test predictions 12classes = dataset.distinct("frames.detections.detections.label") 13 14def jitter(val): 15 if random.random() < 0.10: 16 return random.choice(classes) 17 18 return val 19 20predictions = [] 21for sample_gts in dataset.values("frames.detections"): 22 sample_predictions = [] 23 for frame_gts in sample_gts: 24 sample_predictions.append( 25 fo.Detections( 26 detections=[\ 27 fo.Detection(\ 28 label=jitter(gt.label),\ 29 bounding_box=gt.bounding_box,\ 30 confidence=random.random(),\ 31 )\ 32 for gt in frame_gts.detections\ 33 ] 34 ) 35 ) 36 37 predictions.append(sample_predictions) 38 39dataset.set_values("frames.predictions", predictions) 40 41dataset.evaluate_detections( 42 "frames.predictions", 43 gt_field="frames.detections", 44 eval_key="eval", 45) ``` Note that the only difference in practice is the prefix “frames” used to specify the predictions field and the ground truth field. Learn more about [video views](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#video-views) and [evaluating videos](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#evaluating-videos) in the FiftyOne Docs. ## **Managing multiple evaluations** With all of the flexibility the Evaluation API provides, you’d be well within reason to wonder what evaluation you should perform. Fortunately, FiftyOne makes it easy to perform, manage, and store the results from multiple evaluations! The results from each evaluation can be stored and accessed via an evaluation key, specified by the `eval_key` argument. This allows you to compare different evaluation methods on the same data, ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3dataset = foz.load_zoo_dataset("quickstart") 4dataset.evaluate_detections( 5 "predictions", 6 gt_field = "ground_truth", 7 method = "coco", 8 eval_key = "coco_eval" 9) 10dataset.evaluate_detections( 11 "predictions", 12 gt_field = "ground_truth", 13 method = "open-images", 14 eval_key = "oi_eval" 15) ``` evaluate predictions generated by multiple models, ```python 1dataset.evaluate_detections( 2 "model1_predictions", 3 gt_field = "ground_truth", 4 eval_key = "model1_eval" 5) 6dataset.evaluate_detections( 7 "model2_predictions", 8 gt_field = "ground_truth", 9 eval_key = "model2_eval" 10) ``` Or compare evaluations on different subsets or views of your data, such as a view with only small bounding boxes and a view with only large bounding boxes: ```python 1from fiftyone import ViewField as F 2bbox_area = ( 3 F("bounding_box")[2] * 4 F("bounding_box")[3] 5) 6 7large_boxes = bbox_area > 0.7 8small_boxes = bbox_area < 0.3 9 10# Create a view that contains only small-sized objects 11small_view = ( 12 dataset 13 .filter_labels( 14 "ground_truth", 15 small_boxes 16 ) 17) 18 19# Create a view that contains only large-sized objects 20large_view = ( 21 dataset 22 .filter_labels( 23 "ground_truth", 24 large_boxes 25 ) 26) 27 28small_view.evaluate_detections( 29 "predictions", 30 gt_field="ground_truth", 31 eval_key="eval_small", 32) 33 34large_view.evaluate_detections( 35 "predictions", 36 gt_field="ground_truth", 37 eval_key="eval_large", 38) ``` Learn more about [managing model evaluations](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#managing-evaluations) in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,300+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,500+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,900+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 241+ repositories - 55+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [Evaluation](https://voxel51.com/blog/tag/evaluation) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [model evaluation](https://voxel51.com/blog/tag/model-evaluation) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0ecb0645c4938217bcade4d3d80cf59f7b05329b-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision View Stages Tips and Tricks – Jan 20, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4ac1a727dc192a21563cde51b6e345f620e09376-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-236-lllmstxt|> ## FiftyOne Data Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks for Adding and Merging Data – Feb 17, 2023 Feb 18, 2023 • 4 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#ddb5e8b0e7c1) [A primer on adding and merging](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#0a0aca52c29a) [Encountering a sample multiple times](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#f8f5f146b820) [Add samples by directory](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#5ac8aaa3a64b) [Add from archive](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#bd8cf5183d4b) [Add model predictions](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#78d76b48d70e) [Export multiple labels with merge\_labels()](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#f40a97738bf0) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#fc36761c1765) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#ddb5e8b0e7c1) [A primer on adding and merging](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#0a0aca52c29a) [Encountering a sample multiple times](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#f8f5f146b820) [Add samples by directory](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#5ac8aaa3a64b) [Add from archive](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#bd8cf5183d4b) [Add model predictions](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#78d76b48d70e) [Export multiple labels with merge\_labels()](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#f40a97738bf0) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023#fc36761c1765) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we give practical pointers for using FiftyOne on topics inspired by discussions in the open source community. This week we’ll cover [adding and merging data](https://voxel51.com/docs/fiftyone/recipes/merge_datasets.html). ## **Wait, What’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **A primer on adding and merging** [Datasets](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#using-datasets) are the core data structure in FiftyOne, allowing you to represent your raw data, labels, and associated metadata. [Samples](https://voxel51.com/docs/fiftyone/user_guide/basics.html#samples) are the atomic elements of a `Dataset` that store all the information related to a given piece of data. When you query and manipulate a Dataset object using [dataset views](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#using-views), a DatasetView object is returned, which represents a filtered view into a subset of the underlying dataset’s contents. Many computer vision workflows involve operations that merge data from multiple sources, such as adding new samples to an existing dataset, or merging a model’s predictions into a dataset which contains ground truth labels. In FiftyOne, `Dataset` and `DatasetView` objects come with a variety of methods that make performing these add and merge operations easy. Continue reading for some tips and tricks to help you master adding and merging data in FiftyOne! ## **Encountering a sample multiple times** If you want to add a completely new collection of samples, `samples`, to a dataset, `dataset`, then you can use the `add_samples()` and `add_collection()` methods for the most part interchangeably. However, if there are samples that appear multiple times in your workflows, due to sources of randomness, for instance, then these two methods have different consequences. When `add_samples()` encounters samples that are already present in the dataset to which the method is applied, it generates a _new_ sample with a new `id`, and adds this to the dataset. On the other hand, `add_collection()` ignores the sample and moves on. In the code block below, when applied to the Quickstart Dataset with a random collection from the dataset as input, `add_collection()` leaves the dataset unchanged, whereas `add_samples()` increases the size of the dataset: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# 200 samples 5dataset = foz.load_zoo_dataset("quickstart") 6 7# randomly select 50 samples 8samples = dataset.take(50) 9 10# doesn’t change dataset size 11dataset.add_collection(samples) 12 13# 200 samples → 250 samples 14dataset.add_samples(samples) ``` Learn more about [FiftyOne’s random utils](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.random.html) in the FiftyOne Docs. ## **Add samples by directory** FiftyOne supports a variety of common computer vision data formats, making it easy to load your data into FiftyOne and accelerating your computer vision workflows. FiftyOne’s `DatasetImporter` classes allow you to import data in various formats without needing to write your own loops and I/O scripts. If you have VOC-style data stored in a single directory on disk, for instance, you can create a dataset from this data using the `from_dir()` method: ```python 1import fiftyone as fo 2 3name = "my-dataset" 4data_path = "/path/to/images" 5labels_path = "/path/to/voc-labels" 6 7# Import dataset by explicitly providing paths to the source media and labels 8dataset = fo.Dataset.from_dir( 9 dataset_type=fo.types.VOCDetectionDataset, 10 data_path=data_path, 11 labels_path=labels_path, 12 name=name, 13) ``` With the `add_dir()` method, you can extend the logic of any existing `DatasetImporter` to data that is stored in multiple directories. To add `train` and `val` data in YOLOv5 format to a single dataset, you can run the following: ```python 1import fiftyone as fo 2 3name = "my-dataset" 4dataset_dir = "/path/to/yolov5-dataset" 5 6# The splits to load 7splits = ["train", "val"] 8 9# Load the dataset, using tags to mark the samples in each split 10dataset = fo.Dataset(name) 11for split in splits: 12 dataset.add_dir( 13 dataset_dir=dataset_dir, 14 dataset_type=fo.types.YOLOv5Dataset, 15 split=split, 16 tags=split, 17) ``` This allows you to add the contents of each directory directly to the final dataset without having to instantiate temporary datasets. The `merge_dir()` can also be similarly useful! Learn more about [loading data into FiftyOne](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html) in the FiftyOne Docs. ## **Add from archive** On a related note, if you have data in a common archived format, such as `.zip`, `.tar`, or `.tar.gz` stored on disk, you can use the `add_dir()` or `merge_dir()` methods to add this data to your dataset. If the archived data has not been unpacked yet, FiftyOne will handle this extraction for you! Learn more about [from\_archive()](https://voxel51.com/docs/fiftyone/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.from_archive), [add\_archive()](https://voxel51.com/docs/fiftyone/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.add_archive), and [merge\_archive()](https://voxel51.com/docs/fiftyone/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.merge_archive) in the FiftyOne Docs. ## **Add model predictions** In machine learning workflows, it is common practice to withhold ground truth information at inference time. To accomplish this, it is often beneficial to separate the various fields of your dataset so that only certain subsets of information are available at different steps. When it comes to evaluating model performance at the end of the day, however, we would like to merge ground truth labels and predictions into a common dataset. In FiftyOne, this is possible with the `merge_samples()` method. If we have a `predictions_view` only containing predictions, and a `dataset` with all other information, we can merge the predictions into our base dataset as follows: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# Create a dataset containing only ground truth objects 5dataset = foz.load_zoo_dataset("quickstart") 6dataset = dataset.exclude_fields("predictions").clone() 7 8# Example predictions view 9predictions_view = dataset1.select_fields("predictions") 10 11# Merge the predictions 12dataset.merge_samples(predictions_view) ``` Learn more about [selecting](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.select_fields) and [excluding fields](https://voxel51.com/docs/fiftyone/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.exclude_fields) in the FiftyOne Docs. ## **Export multiple labels with `merge_labels()`** If you have multiple `Label` fields and you want to export your data using a common format, you can use the `merge_labels()` method to merge all of these label fields into one field for export. For instance, if you have three labels, `ground_truth`, `model1_predictions`, and `model2_predictions`, you can merge all of these labels as follows: ```python 1import fiftyone as fo 2 3dataset = fo.load_dataset(...) 4 5## clone label fields into temporary fields 6dataset.clone_sample_field("ground_truth", "tmp") 7dataset.clone_sample_field("model1_predictions", "tmp1") 8dataset.clone_sample_field("model2_predictions", "tmp2") 9 10## merge model1 predictions into ground_truth 11dataset.merge_labels("tmp1", "tmp") 12## merge model2 predictions into ground_truth 13dataset.merge_labels("tmp2", "tmp") 14 15## export the merged labels field 16dataset.export(..., label_field="tmp") 17 18## clean up 19dataset.delete_sample_fields(["tmp", "tmp1", "tmp2"]) ``` If you want to export the data just so that you can import it at a later time, however, then you can avoid all of this and instead make your dataset persistent! ```python 1dataset.persistent = True ``` Learn more about [labels](https://voxel51.com/docs/fiftyone/user_guide/basics.html#labels) and [dataset persistence](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#dataset-persistence) in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,350+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,500+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,100+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 246+ repositories - 56+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [adding data](https://voxel51.com/blog/tag/adding-data) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [merging data](https://voxel51.com/blog/tag/merging-data) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0ecb0645c4938217bcade4d3d80cf59f7b05329b-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision View Stages Tips and Tricks – Jan 20, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4ac1a727dc192a21563cde51b6e345f620e09376-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-237-lllmstxt|> ## FiftyOne Customization Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Tips and Tricks for Customizing your Computer Vision Workflows – Mar 03, 2023 Mar 4, 2023 • 4 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#bb0233a816bd) [A primer on configurations](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#1b5a8a8da442) [Access your configs](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#ea644397836b) [Modify your configs](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#a68d061c16aa) [Facilitate your ML workflows](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#deec1c0ffe9a) [Enable plugins](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#a25ec09f8cfd) [Configure the app on a per-dataset basis](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#3e90a518f24a) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#2170f501f765) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#bb0233a816bd) [A primer on configurations](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#1b5a8a8da442) [Access your configs](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#ea644397836b) [Modify your configs](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#a68d061c16aa) [Facilitate your ML workflows](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#deec1c0ffe9a) [Enable plugins](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#a25ec09f8cfd) [Configure the app on a per-dataset basis](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#3e90a518f24a) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023#2170f501f765) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we give practical pointers for using FiftyOne on topics inspired by discussions in the open source community. This week we’ll cover [configuring FiftyOne](https://voxel51.com/docs/fiftyone/user_guide/config.html#configuring-fiftyone) and [the FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/config.html#configuring-the-app) to craft your ideal experience working with computer vision data. ## **Wait, What’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **A primer on configurations** FiftyOne is built for flexibility, customization, and extensibility. Want to integrate with annotation tools? Check out our tight integrations with [CVAT](https://docs.voxel51.com/integrations/cvat.html), [Labelbox](https://docs.voxel51.com/integrations/labelbox.html), and [Label Studio](https://docs.voxel51.com/integrations/labelstudio.html). Want to work with benchmark datasets or evaluation routines? We’ve got you covered with native support for [COCO](https://docs.voxel51.com/integrations/coco.html), [Open Images](https://docs.voxel51.com/integrations/open_images.html), and more. FiftyOne integrates seamlessly into a variety of computer vision workflows, allowing you to tailor your experience to your specific needs. FiftyOne is made up of the FiftyOne Library, which is a software development kit (SDK), and the FiftyOne App, which is a GUI for visualizing and graphically interacting with your computer vision data. With FiftyOne configs, and FiftyOne App configs, each can be configured separately. You get control over everything from databases to ML backends, default file formats, and layout in the app. Continue reading for some tips and tricks to help you fashion FiftyOne and the FiftyOne App using configs! ## **Access your configs** The first step on the way to customizing your configs is accessing them. Fortunately, in FiftyOne your config and app config are available in Python as properties. To see the global FiftyOne config, use ```python 1import fiftyone as fo 2print(fo.config) ``` And to get the global app config, run the similar command ```python 1import fiftyone as fo 2print(fo.app_config) ``` Additionally, every session has its own app config, which can be accessed via the `config` property: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4## load in dataset 5dataset = foz.load_zoo_dataset("quickstart") 6 7## create a session that doesn't display automatically 8session = fo.launch_app(dataset, auto=False) 9 10print(session.config) ``` Learn more about [sessions](https://voxel51.com/docs/fiftyone/user_guide/app.html#sessions) in the FiftyOne Docs. ## **Modify your configs** Once you’ve identified your configs, there are multiple ways to modify them. If you want to set the default FiftyOne behavior to displaying progress bars, you can do so either by creating a JSON config file `~/.fiftyone/config.json` with the code: ```json 1{ 2 "show_progress_bars": true 3} ``` Or you can the corresponding environment variable from the command line: ```bash 1export FIFTYONE_SHOW_PROGRESS_BARS=true ``` Lastly, you can achieve the same effect by setting the config directly in your Python code: ```python 1import fiftyone as fo 2fo.config.show_progress_bars = True ``` Learn more about [order of precedence](https://voxel51.com/docs/fiftyone/user_guide/config.html#order-of-precedence) for configs in the FiftyOne Docs. ## **Facilitate your ML workflows** With FiftyOne configs, you can set up FiftyOne for optimized use in your machine learning workflows. When downloading models from the [FiftyOne Model Zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html), sometimes a model has both TensorFlow and PyTorch implementations available. By setting the `default_ml_backend` property in your FiftyOne config, you can specify the preferred ML library to use in these scenarios. If you prefer Pytorch over TensorFlow, you can make this the default: ```python 1import fiftyone as fo 2fo.config.default_ml_backend = "torch" ``` Additionally, you can specify what batch size you wish to use, if any, when applying models from the model zoo to datasets. Running this sets the default batch size to 100 samples. ```python 1import fiftyone as fo 2fo.config.default_batch_size = 100 ``` Learn more about the [API for FiftyOne’s Model Zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/api.html#model-zoo-api) in the FiftyOne Docs. ## **Enable plugins** By modifying your app config, you can enable plugins that customize and extend the behavior of FiftyOne. You can leverage built-in plugins like the [Map panel](https://docs.voxel51.com/user_guide/app.html#map-panel) or the [3D visualizer](https://docs.voxel51.com/user_guide/groups.html#using-the-3d-visualizer), or you can [enable your own application-specific plugins](https://docs.voxel51.com/plugins/index.html)! Plugin configurations are controlled via the `”plugins”` key in the app config. If you want to set the default camera position in the 3D visualizer, you can do so by modifying properties of the `”plugins.3d”` config in `~/.fiftyone/app_config.json`: ```json 1// The default values are shown below 2{ 3 "plugins": { 4 "3d": { 5 // Whether to show the 3D visualizer 6 "enabled": true, 7 8 // The initial camera position in 3D space 9 "defaultCameraPosition": {"x": 100, "y": 0, "z": 20}, 10 } 11 } 12} ``` When we use the app config to set the default camera position in the 3D visualizer, the [Quickstart Groups](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#quickstart-groups) dataset looks like this: ![](https://cdn.sanity.io/images/h6toihm1/production/423d342a5104e8ee8ae4ec1a21b795765d6bea0a-3440x1810.png?auto=format&dpr=2&fit=max&q=75&w=1600) Learn more about [developing and installing plugins](https://docs.voxel51.com/plugins/index.html) in the FiftyOne Docs. ## **Configure the app on a per-dataset basis** In addition to the global app config, each dataset has its own app config, stored in the `dataset.app_config` attribute. This allows you to configure behavior in the FiftyOne App on a per-dataset basis! Continuing on with the 3D visualizer plugin example above, let’s suppose we are working with a bunch of grouped datasets (with point clouds), and for most of them we want to use a default camera position specified in the App config JSON file, with `z = 100` so we are looking top-down. However, we have one particular dataset for which we would like the default view to be profile. We can set the default app config for this dataset in Python without changing the default for the rest of our datasets: ```python 1pos = {'x':100, 'y':0, 'z':0} 2dataset.app_config.plugins["3d"]["defaultCameraPosition"] = pos ``` Learn more about [dataset app configs](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#custom-app-config) in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,375+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,575+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,300+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 245+ repositories - 56+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne App](https://voxel51.com/blog/tag/fiftyone-app) [PyTorch](https://voxel51.com/blog/tag/pytorch) [TensorFlow](https://voxel51.com/blog/tag/tensorflow) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0fc50e89593dec2117ce5c1934761cf61b24f8c4-1920x1080.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Tips and Tricks for Accelerating Computer Vision Workflows – Mar 17, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 18, 2023](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0ecb0645c4938217bcade4d3d80cf59f7b05329b-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision View Stages Tips and Tricks – Jan 20, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-238-lllmstxt|> ## FiftyOne Tips for Workflows [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Tips and Tricks for Accelerating Computer Vision Workflows – Mar 17, 2023 Mar 18, 2023 • 5 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#e03d721c287b) [Load data into FiftyOne faster](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#fac265960b5d) [Peruse through samples faster](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#3295cf738432) [Filter samples faster](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#e5576dd3108c) [Start with most unique samples](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#fcd7a7fcab4e) [Pre-annotate with embeddings](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#8a352c39a97d) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#772961cc6dc5) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#e03d721c287b) [Load data into FiftyOne faster](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#fac265960b5d) [Peruse through samples faster](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#3295cf738432) [Filter samples faster](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#e5576dd3108c) [Start with most unique samples](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#fcd7a7fcab4e) [Pre-annotate with embeddings](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#8a352c39a97d) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023#772961cc6dc5) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we give practical pointers for using FiftyOne on topics inspired by discussions in the open source community. This week we’ll cover some tips and tricks that will help you accelerate your computer vision workflows using FiftyOne. ## **Wait, What’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **Load data into FiftyOne faster** One of the great features of PyTorch is the `DataLoader` class, which makes it easy to [efficiently load and process data](https://pytorch.org/tutorials/beginner/basics/data_tutorial.html). This becomes especially useful when dealing with large datasets, where it can be inefficient or impossible to load the entire dataset into memory at once. When you’re working with PyTorch data in FiftyOne, you can leverage the speed and simplicity of your PyTorch `DataLoader` to import your data into and analyze that data in FiftyOne. Here’s an example using the test split of the [CIFAR-10 dataset](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#cifar-10): First, we create the TorchVision dataset and instantiate a `DataLoader` for the dataset. ```python 1import torch 2import torchvision 3 4# Downloads dataset and prepares it for loading in a DataLoader 5dataset = torchvision.datasets.CIFAR10( 6 "/tmp/fiftyone/custom-parser/pytorch", 7 train=False, 8 download=True, 9 transform=torchvision.transforms.ToTensor(), 10) 11classes = dataset.classes 12data_loader = torch.utils.data.DataLoader(dataset, batch_size=1) ``` Then we create the FiftyOne `Dataset` for this data and specify the directory in which we will store the images. ```python 1import fiftyone as fo 2dataset = fo.Dataset("cifar10-samples") 3 4# The directory to use to store the individual images on disk 5dataset_dir = "/tmp/fiftyone/custom-parser/fiftyone" ``` Finally, we create a sample parser and use the PyTorch `DataLoader` with FiftyOne’s `ingest_labeled_images()` method to fill our FiftyOne `Dataset` with the PyTorch data. ```python 1sample_parser = PyTorchClassificationDatasetSampleParser(classes) 2 3dataset.ingest_labeled_images( 4 data_loader, 5 sample_parser, 6 dataset_dir=dataset_dir 7) ``` Once you have PyTorch data loaded into FiftyOne, it is also easy to train [Lightning Flash](https://github.com/Lightning-AI/lightning-flash) tasks on your FiftyOne data and add model predictions to FiftyOne! Learn more about FiftyOne’s PyTorch [Lightning Flash integration](https://voxel51.com/docs/fiftyone/integrations/lightning_flash.html) in the FiftyOne Docs. For other ways to add samples, check out the docs for [adding samples manually](https://docs.voxel51.com/user_guide/dataset_creation/index.html#loading-custom-datasets) or [using dataset importers](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#loading-datasets-from-disk) to load data into FiftyOne. ## **Peruse through samples faster** If you’re working with high-resolution media files like satellite imagery or photo-realistic AI-generated data, you might notice some slight delays in rendering when you scroll through samples in the sample grid in the FiftyOne App. This is because, by default, the app is rendering the high-resolution media files in real time. If you’re experiencing this buffering, you might get a speed boost by configuring the app to use thumbnail images! By modifying the App Config, you can enable [multiple media fields](https://voxel51.com/docs/fiftyone/user_guide/app.html#multiple-media-fields) and choose which media field is displayed in the sample grid. One common workflow is creating lower resolution, downscaled images for each sample, and configuring the app to display these in the sample grid. This way you retain the depth of information in your high-resolution media files for downstream workflows without sacrificing speed while investigating your data. When you click on a thumbnail in the sample grid, the media file in the resulting full-screen modal will be the high-resolution version! Here’s an example of how you might accomplish this effect: First, we create thumbnail images and store their paths in a `thumbnail_path` field: ```python 1import fiftyone as fo 2import fiftyone.utils.image as foui 3import fiftyone.zoo as foz 4 5dataset = foz.load_zoo_dataset("quickstart") 6 7# Generate some thumbnail images 8foui.transform_images( 9 dataset, 10 size=(-1, 32), 11 output_field="thumbnail_path", 12 output_dir="/tmp/thumbnails", 13) ``` Then we modify the [dataset’s App config](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#custom-app-config) to expose these thumbnails in the sample grid and create a session: ```python 1# Modify the dataset's App config 2dataset.app_config.media_fields = ["filepath", "thumbnail_path"] 3dataset.app_config.grid_media_field = "thumbnail_path" 4dataset.save() # must save after edits 5 6session = fo.launch_app(dataset) ``` Learn more about [configuring the FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html#configuring-the-app) in the FiftyOne Docs. ## **Filter samples faster** Another way to accelerate visually inspection of your data in the FiftyOne App is by setting the sidebar mode to `”fast”`. When filtering and matching via the sidebar in the App, the `sidebar_mode` property in the [App Config](https://voxel51.com/docs/fiftyone/user_guide/config.html#configuring-the-app) allows you to specify whether these operations should be applied to all samples in the view and all of the relevant fields, or just the samples visible in the sample grid and fields that are expanded in the filter tray. This can be a dramatic time-saver for datasets with many samples, many fields, or both. For datasets with more than 10,000 samples, the FiftyOne App’s default behavior is `”fast”`, but if you want to make this the default for all of your datasets, you can set this in the App Config with: ```python 1import fiftyone as fo 2fo.app_config.sidebar_mode = “fast” ``` Alternatively, if you want to set this only for a specific dataset, you can modify the dataset’s app config: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# load in dataset 5dataset = foz.load_zoo_dataset("quickstart") 6dataset.app_config.sidebar_mode = “fast” ``` Learn more about [using the sidebar](https://voxel51.com/docs/fiftyone/user_guide/app.html#using-the-sidebar) in the FiftyOne Docs. ## **Start with most unique samples** In many machine learning workflows, some samples matter more than others. Whether you are deciding what subsets of your dataset to send out for annotation, or identifying edge cases and failure modes of your models, it helps to focus your attention. One trick you can use to explore a broad range of samples rather than scrolling through a bunch of similar examples is to look at the most “unique” samples first. With the FiftyOne Brain, you can use embeddings to compute a score for each sample in your dataset that tells you how unique that sample is, relative to all of the other samples. ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4 5# load in dataset 6dataset = foz.load_zoo_dataset("quickstart") 7 8# compute uniqueness using default (CLIP) embedding 9fob.compute_uniqueness(dataset) ``` You can then sort your samples by uniqueness, passing in `reverse=True` so that the most unique samples (highest values in `uniqueness` field) appear first: ```python 1unique_view = dataset.sort_by("uniqueness", reverse=True) ``` Then you can save time by looking at the most unique samples in your dataset first! ```python 1session = fo.launch_app(unique_view) ``` Learn more about the [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html) and [image uniqueness](https://voxel51.com/docs/fiftyone/tutorials/uniqueness.html) in the FiftyOne Docs. ## **Pre-annotate with embeddings** Embeddings can also save you precious time by helping you pre-annotate your computer vision data. Labeling and annotation of ground truth data often represents one of the most tedious and time-intensive tasks in computer vision workflows. In addition to hours or days spent, annotation can also be quite expensive. Fortunately, for some datasets it is possible to use the structure uncovered by embeddings to do a coarse first pass through the data. With the [MNIST dataset](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#mnist), for instance, computing embeddings and then using dimensionality reduction techniques like [tSNE](https://en.wikipedia.org/wiki/T-distributed_stochastic_neighbor_embedding), you can uncover multiple mostly-distinct clusters of samples. Using the lasso to interact with this visualization and view each cluster in the FiftyOne App, it quickly becomes clear that these clusters roughly correspond to different numerals. The MNIST dataset of course already has ground truth labels, but if it didn’t, you could leverage these clusters to tag samples with pre-annotations. ![](https://cdn.sanity.io/images/h6toihm1/production/b324c405c104251cd4d48ba0022a8b02312a8429-1919x1051.gif?auto=format&dpr=2&fit=max&q=75&w=1600) Learn more about [using image embeddings](https://voxel51.com/docs/fiftyone/tutorials/image_embeddings.html#Pre-annotation-of-samples) and pre-annotating samples in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,400+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,600+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,400+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 256+ repositories - 57+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [annotation](https://voxel51.com/blog/tag/annotation) [CIFAR-10](https://voxel51.com/blog/tag/cifar-10) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [embeddings](https://voxel51.com/blog/tag/embeddings) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Lightning Flash](https://voxel51.com/blog/tag/lightning-flash) [MNIST](https://voxel51.com/blog/tag/mnist) [PyTorch](https://voxel51.com/blog/tag/pytorch) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/dc8a2e7a894316856af5a109ae43f8787959f179-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Embeddings Tips and Tricks – Mar 31, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 31, 2023](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/205569e4c6b9ed68023e0ed430acbb40a981026f-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Tips and Tricks for Customizing your Computer Vision Workflows – Mar 03, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 4, 2023](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-239-lllmstxt|> ## FiftyOne Embeddings Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Embeddings Tips and Tricks – Mar 31, 2023 Mar 31, 2023 • 5 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#8cd2889a52de) [A primer on embeddings](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#d558a2712cf3) [Compute sample properties (uniqueness, similarity, and more) from embeddings](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#9e00edd25177) [Visualize embeddings](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#4e2114629006) [Use your own embedding model](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#0e8fe638a0f3) [Compute embeddings for object patches](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#5324929779b7) [Perform zero-shot classification with CLIP](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#750c798dae68) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#1fc45d59f1ff) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#8cd2889a52de) [A primer on embeddings](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#d558a2712cf3) [Compute sample properties (uniqueness, similarity, and more) from embeddings](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#9e00edd25177) [Visualize embeddings](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#4e2114629006) [Use your own embedding model](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#0e8fe638a0f3) [Compute embeddings for object patches](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#5324929779b7) [Perform zero-shot classification with CLIP](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#750c798dae68) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023#1fc45d59f1ff) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we give practical pointers for using FiftyOne on topics inspired by discussions in the open source community. This week we’ll cover [embeddings](http://labels/). ## **Wait, What’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **A primer on embeddings** In computer vision, embeddings are used to represent visual features of images or videos as numerical vectors. These vectors, which are typically lower in dimension than the raw source media, capture underlying characteristics of the visual data and can be used for downstream tasks such as image classification, object detection, and image retrieval. In FiftyOne, embeddings are managed by the [FiftyOne Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html), a library that provides powerful machine learning techniques designed to transform how you curate your data from an art into a measurable science. Using embeddings makes it easy to analyze and visualize datasets, and accelerates your ML workflows. Continue reading for some tips and tricks to help you master embeddings in FiftyOne! ## **Compute sample properties (uniqueness, similarity, and more)** **from embedding** s If you want to harness the power of embeddings in computer vision but aren’t sure where to start, FiftyOne has you covered. Any FiftyOne Brain method that utilizes embeddings is equipped with a default general purpose embedding model. This means that in many cases you can use Brain methods out of the box to assess properties of the samples in your dataset. For instance, when training a model, unique samples will help your model to understand and represent distributions in your dataset. To aid in this process, the FiftyOne Brain provides a method to compute the “uniqueness” of images, helping you to answer questions such as “What data should I select to annotate?”. Here’s how easy it is to compute the uniqueness of images in the quickstart dataset: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4 5# load in dataset 6dataset = foz.load_zoo_dataset("quickstart") 7 8fob.compute_uniqueness(dataset) ``` In a similar vein, when constructing a dataset or training a model, have you ever wanted to find similar examples to an image or object of interest? The FiftyOne Brain makes computing visual similarity easy, too. You can compute the similarity of samples in your dataset using the default ( [MobileNetv2](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/models.html#mobilenet-v2-imagenet-torch)) embedding model and store the results in the brain key `image_sim`: ```python 1fob.compute_similarity(dataset, brain_key="image_sim") ``` You can then sort your samples by similarity or use this information to find potential duplicate images. ![](https://cdn.sanity.io/images/h6toihm1/production/1709180320224b1d50c089d7929777e6e7745074-1276x957.gif?auto=format&dpr=2&fit=max&q=75&w=1276) Learn more about the FiftyOne Brain’s [similarity interface](https://voxel51.com/docs/fiftyone/api/fiftyone.brain.similarity.html), as well as other brain methods, such as [sample hardness](https://docs.voxel51.com/user_guide/brain.html#brain-sample-hardness) and [mistakeness](https://docs.voxel51.com/user_guide/brain.html#brain-label-mistakes), in the FiftyOne Docs. ## **Visualize embeddings** Another application of embeddings in computer vision is visualization. By pairing embeddings with a dimensionality reduction technique such as [uMAP](https://umap-learn.readthedocs.io/en/latest/) or [tSNE](https://en.wikipedia.org/wiki/T-distributed_stochastic_neighbor_embedding), you can plot your data in a low dimensional space, allowing you to identify clusters and better understand the patterns in your data. In FiftyOne, all of the logic of combining embeddings, dimensionality reduction techniques, and plotting are wrapped in the easy to use interface of the FiftyOne Brain’s `compute_visualization()` method. If we want to generate default embeddings for the [Quickstart dataset](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#dataset-zoo-quickstart) and use them to visualize our dataset, employing uMAP to reduce the dimensionality of the embedding vector down to two dimensions, we can run: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4 5dataset = foz.load_zoo_dataset("quickstart") 6 7# Image embeddings 8fob.compute_visualization(dataset, brain_key="img_viz") 9 10# Object patch embeddings 11fob.compute_visualization( 12 dataset, patches_field="ground_truth", brain_key="gt_viz" 13) 14 15session = fo.launch_app(dataset) ``` Then we can visualize these results, coloring by label: ![](https://cdn.sanity.io/images/h6toihm1/production/94c168b1408e66a8d61a767d02e7ba541cd34cd6-1235x770.gif?auto=format&dpr=2&fit=max&q=75&w=1235) Learn more about the [Embeddings Panel](https://docs.voxel51.com/user_guide/app.html#embeddings-panel) in the FiftyOne Docs. ## **Use your own embedding model** If you’re working with more specialized or application-specific computer vision data, it might be beneficial to use your own, custom model to generate embeddings. Fortunately, all of the Brain methods that utilize embeddings, from `compute_uniqueness()` and `compute_similarity()` to `compute_visualization()` support this! Any FiftyOne Brain method that works with embeddings allows you to specify your own embeddings in three ways. The first two ways involve the optional `embeddings` argument, which is useful if you have already precomputed your embeddings. You can either give the precomputed embeddings explicitly as input in the form of an `num_samples x num_embedding_dims` array, or you can pass in the name of a field in the dataset containing the embeddings: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4 5# The BDD dataset must be manually downloaded. See the zoo docs for details 6source_dir = "/path/to/dir-with-bdd100k-files" 7 8# Load dataset 9dataset = foz.load_zoo_dataset( 10 "bdd100k", split="validation", source_dir=source_dir, 11) 12 13# Compute embeddings 14# You will likely want to run this on a machine with GPU, as this requires 15# running inference on 10,000 images 16model = foz.load_zoo_model("mobilenet-v2-imagenet-torch") 17 18# option 1 19embeddings = dataset.compute_embeddings(model) 20results = fob.compute_visualization(dataset, embeddings=embeddings) 21 22# option 2 23embeddings = dataset.compute_embeddings(model, embeddings_field = “my_embeddings”) 24results = fob.compute_visualization(dataset, embeddings=”my_embeddings”) ``` As an alternative, if you have not yet computed your embeddings, you can pass in a model for FiftyOne to use to compute the embeddings: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4 5# The BDD dataset must be manually downloaded. See the zoo docs for details 6source_dir = "/path/to/dir-with-bdd100k-files" 7 8# Load dataset 9dataset = foz.load_zoo_dataset( 10 "bdd100k", split="validation", source_dir=source_dir, 11) 12 13# Compute embeddings 14# You will likely want to run this on a machine with GPU, as this requires 15# running inference on 10,000 images 16model = foz.load_zoo_model("mobilenet-v2-imagenet-torch") 17 18# option 3 19results = fob.compute_visualization(dataset, model = model) ``` Any model that has logits should work, and you can check if the model supports embeddings by verifying that `print(model.has_embeddings)` returns `True`. Learn more about the [FiftyOne Model Zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html#) in the FiftyOne Docs. ## **Compute embeddings for object patches** Just as you can compute embeddings for images, FiftyOne allows you to compute embeddings for object patches, with the built-in method `compute_patch_embeddings()` functioning in analogous fashion to `compute_embeddings()`. To compute embeddings for ground truth objects, we just use `”ground_truth”` as the patches field: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# Load zoo model 5model = foz.load_zoo_model("inception-v3-imagenet-torch") 6print(model.has_embeddings) # True 7 8# Load zoo dataset 9dataset = foz.load_zoo_dataset("quickstart") 10 11dataset.compute_patch_embeddings(model, "ground_truth", embeddings_field = "gt_embed") ``` Alternatively, Brain methods like `compute_similarity()` and `compute_visualization()` work natively with patches when you specify the `patches_field` argument, telling the Brain which labels to use: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4 5# Load zoo model 6model = foz.load_zoo_model("inception-v3-imagenet-torch") 7print(model.has_embeddings) # True 8 9# Load zoo dataset 10dataset = foz.load_zoo_dataset("quickstart") 11 12# compute similarity of ground truth patches 13fob.compute_similarity(dataset, patches_field = "ground_truth") ``` Learn more about [object patches](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#object-patches) and [computing embeddings for object patches](https://docs.voxel51.com/api/fiftyone.core.patches.html#fiftyone.core.patches.PatchesView.compute_patch_embeddings) in the FiftyOne Docs. ## **Perform zero-shot classification with CLIP** Over the past few years, progress in embeddings has enabled a suite of new applications and use cases in computer vision. One of the most significant advances in this area has been the development and honing of techniques which embed multimodal data into a common embedding space. In particular, with [OpenAI’s CLIP model](https://openai.com/blog/clip/) pioneering contrastive language-image pre-training, it is possible to transfer knowledge from one domain to another. This paved the way for dramatically improved zero-shot computer vision models. The [PyTorch implementation of the CLIP Vision Transformer model](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/models.html#clip-vit-base32-torch) in FiftyOne’s Model Zoo makes it easy to leverage CLIP’s multi-modal embeddings to classify images into novel categories. Here’s an example of performing zero-shot classification on [COCO](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#coco-2017) validation data with custom class labels: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset( 5 "coco-2017", 6 split="validation", 7 dataset_name=fo.get_default_dataset_name(), 8 max_samples=50, 9 shuffle=True, 10) 11 12model = foz.load_zoo_model("clip-vit-base32-torch") 13 14dataset.apply_model(model, label_field="predictions") 15 16session = fo.launch_app(dataset) 17 18# Make zero-shot predictions with custom classes 19 20model = foz.load_zoo_model( 21 "clip-vit-base32-torch", 22 text_prompt="A photo of a", 23 classes=["person", "dog", "cat", "bird", "car", "tree", "chair"], 24) 25 26dataset.apply_model(model, label_field="predictions") 27session.refresh() ``` ![](https://cdn.sanity.io/images/h6toihm1/production/54d657afb9b554b134337aa55d84d36c52c1c9f9-2636x1436.png?auto=format&dpr=2&fit=max&q=75&w=1600) Learn more about [apply\_model()](https://docs.voxel51.com/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.apply_model) and [understanding your computer vision data with a CLIP model](https://medium.com/voxel51/finding-images-with-words-92b078314ed1) in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,475+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,700+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,600+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 266+ repositories - 58+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Berkeley Deep Drive](https://voxel51.com/blog/tag/berkeley-deep-drive) [CLIP](https://voxel51.com/blog/tag/clip) [COCO](https://voxel51.com/blog/tag/coco) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [embeddings](https://voxel51.com/blog/tag/embeddings) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Labelbox](https://voxel51.com/blog/tag/labelbox) [MNIST](https://voxel51.com/blog/tag/mnist) [MobileNet](https://voxel51.com/blog/tag/mobilenet) [Model Zoo](https://voxel51.com/blog/tag/model-zoo) [uMAP](https://voxel51.com/blog/tag/umap) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0fc50e89593dec2117ce5c1934761cf61b24f8c4-1920x1080.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Tips and Tricks for Accelerating Computer Vision Workflows – Mar 17, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 18, 2023](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0ecb0645c4938217bcade4d3d80cf59f7b05329b-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision View Stages Tips and Tricks – Jan 20, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-240-lllmstxt|> ## Curate Computer Vision Datasets [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tutorials](https://voxel51.com/blog/category/tutorials) How to Curate, Annotate, and Improve Computer Vision Datasets with FiftyOne and Labelbox Jan 14, 2022 • 4 min read Article content In this article [Follow along in Colab](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#a6751d34df4d) [Setup](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#b6179489e7f2) [Raw Data](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#664925fe3830) [Annotation](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#042bd8775944) [Next Steps](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#e2fb6b701e2f) [Additional Utilities](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#03848e7a25ea) [Summary](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#cc03fce34800) In this article [Follow along in Colab](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#a6751d34df4d) [Setup](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#b6179489e7f2) [Raw Data](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#664925fe3830) [Annotation](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#042bd8775944) [Next Steps](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#e2fb6b701e2f) [Additional Utilities](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#03848e7a25ea) [Summary](https://voxel51.com/blog/how-to-curate-annotate-and-improve-computer-vision-datasets-with-fiftyone-and-labelbox#cc03fce34800) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _A guide to using the integration between FiftyOne and Labelbox to build high-quality image and video datasets_ ![](https://cdn.sanity.io/images/h6toihm1/production/0d2bbf0f002969f6fab0099d9e60be2fff2a256e-1400x887.png?auto=format&dpr=2&fit=max&q=75&w=1400) Modern computer vision projects in the deep learning age always start with the same thing. LOTS OF DATA! For just about any task you can find countless models with open-source code ready for you to train them. The only thing that you need for your specific task is a sufficiently large, labeled dataset. In this post, we show how to use the [integration](https://voxel51.com/docs/fiftyone/integrations/labelbox.html) between the open-source dataset curation and model analysis tool, [FiftyOne](http://fiftyone.ai/), and the widely popular annotation tool, [Labelbox](https://labelbox.com/), in order to build a high-quality dataset for computer vision. ## **Follow along in Colab** You can follow along with the examples in this post directly in your browser through [this Google Colab notebook](https://colab.research.google.com/github/voxel51/fiftyone/blob/develop/docs/source/tutorials/labelbox_annotation.ipynb)! ![](https://cdn.sanity.io/images/h6toihm1/production/803b3a1c34f385e60132782a415685b459425df2-1400x922.png?auto=format&dpr=2&fit=max&q=75&w=1400) ## Setup To start, you need to [install FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html) and Labelbox. ```python 1pip install fiftyone labelbox ``` You also need to and set up a [Labelbox account](https://app.labelbox.com/). FiftyOne supports both standard Labelbox cloud accounts as well as Labelbox enterprise on-premises solutions. The easiest way to get started is to use the default Labelbox server, which simply requires creating an account and then providing your API key as shown below. ![](https://cdn.sanity.io/images/h6toihm1/production/71e4af2008940c124942f21745edb7b29c3513d5-842x285.png?auto=format&dpr=2&fit=max&q=75&w=842) ```python 1export FIFTYONE_LABELBOX_API_KEY=... ``` Alternatively, for a more permanent solution, you can store your credentials in your FiftyOne annotation config located at `~/.fiftyone/annotation_config.json`: ```python 1{ 2 "backends": { 3 "labelbox": { 4 "api_key": ..., 5 } 6 } 7} ``` ## Raw Data To start, you need to gather raw image or video data relevant to your task. The internet has a lot of places to look for free data. Assuming you have your raw data downloaded locally, you can [easily load it into FiftyOne](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html). ```python 1import fiftyone as fo 2 3dataset = fo.Dataset.from_dir( 4 dataset_dir="/path/to/dir", 5 dataset_type=fo.types.ImageDirectory, 6) ``` Another method is to use publically available datasets that may be relevant. For example, the [Open Images dataset](https://storage.googleapis.com/openimages/web/download.html) contains millions of images available for public use and can be accessed directly through the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html). ```python 1import fiftyone.zoo as foz 2 3dataset = foz.load_zoo_dataset( 4 "open-images-v6", 5 split="validation", 6 classes="Person", 7 max_samples=500, 8) ``` Either way, once your data is in FiftyOne, we can [visualize it in the FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html). ```python 1import fiftyone as fo 2 3session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/bfe01420f6506f1b9b4999ab60ef8305fdc9846f-951x778.png?auto=format&dpr=2&fit=max&q=75&w=951) FiftyOne provides a [variety of methods](https://voxel51.com/docs/fiftyone/user_guide/brain.html) that can help you understand the quality of the dataset and pick the best samples to annotate. For example, the `compute_similarity()` the method can be used to find both the most similar, and the most unique samples, ensuring that your dataset will contain an even distribution of data. ```python 1import fiftyone.brain as fob 2 3results = fob.compute_similarity(dataset, brain_key="img_sim") 4results.find_unique(10) ``` Now to select only the slice of our dataset that contains the 10 most unique samples. ```python 1unique_view = dataset.select(results.unique_ids) 2session.view = unique_view ``` ![](https://cdn.sanity.io/images/h6toihm1/production/c07c7c7a9d2417c069747a6ceaa4139763546259-951x794.png?auto=format&dpr=2&fit=max&q=75&w=951) ## Annotation The [integration between FiftyOne and Labelbox](https://voxel51.com/docs/fiftyone/integrations/labelbox.html) allows you to begin annotating your image or video data by calling a single method! ```python 1anno_key = "annotation_run_1" 2classes = ["vehicle", "animal", "plant"] 3unique_view.annotate( 4 anno_key, 5 backend="labelbox", 6 label_field="detections", 7 classes=classes, 8 label_type="detections", 9) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/91d730beb19614247949268fb0518425a3e5aa1e-1072x619.png?auto=format&dpr=2&fit=max&q=75&w=1072) The annotations can then be [loaded back into FiftyOne](https://voxel51.com/docs/fiftyone/integrations/labelbox.html#loading-annotations) in just one more line. ```python 1unique_view.load_annotations(anno_key) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/a14cb991b4d294067954fe6ff573500b52042147-957x791.png?auto=format&dpr=2&fit=max&q=75&w=957) This API provides [advanced customization options](https://voxel51.com/docs/fiftyone/integrations/labelbox.html#requesting-annotations) for your annotation tasks. For example, we can construct a sophisticated schema to define the annotations we want and even directly assign the annotators: ```python 1anno_key = "labelbox_assign_users" 2 3members = [\ 4 ("fiftyone_labelbox_user1@gmail.com", "LABELER"),\ 5 ("fiftyone_labelbox_user2@gmail.com", "REVIEWER"),\ 6 ("fiftyone_labelbox_user3@gmail.com", "TEAM_MANAGER"),\ 7] 8 9# Set up the Labelbox editor to reannotate 10# existing "detections" labels and 11# a new "keypoints" field 12label_schema = { 13 "detections_new": { 14 "type": "detections", 15 "classes": dataset.distinct("detections.detections.label"), 16 }, 17 "keypoints": { 18 "type": "keypoints", 19 "classes": ["Person"], 20 } 21} 22 23unique_view.annotate( 24 anno_key, 25 backend="labelbox", 26 label_schema=label_schema, 27 members=members, 28 launch_editor=True, 29) 30# Annotate in Labelbox 31# Download results and clean the run from FiftyOne and Labelbox 32unique_view.load_annotations(anno_key, cleanup=True) ``` ## Next Steps Now that you have a labeled dataset, you can go ahead and start training a model. FiftyOne lets you [export your data to disk in a variety of formats](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html) (ex: COCO, YOLO, etc) expected by most training pipelines. It also provides workflows for using popular model training libraries like [PyTorch](https://towardsdatascience.com/stop-wasting-time-with-pytorch-datasets-17cac2c22fa8), [PyTorch Lightning Flash](https://voxel51.com/docs/fiftyone/integrations/lightning_flash.html), and [Tensorflow](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html#tfobjectdetectiondataset). Once the model is trained, the [model predictions can be loaded](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html#model-predictions) back into FiftyOne. These [predictions can then be evaluated against the ground truth](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html) annotations to find where the model is performing well, and where it is performing poorly. This provides insight into the type of samples that need to be added to the training set, as well as any annotation errors that may exist. ```python 1# Load an existing dataset with predictions 2dataset = foz.load_zoo_dataset("quickstart") 3 4# Evaluate model predictions 5dataset.evaluate_detections( 6 "predictions", 7 gt_field="ground_truth", 8 eval_key="eval", 9) ``` We can use the [powerful querying capabilities of the FiftyOne API](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) to create a [view filtering](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#filtering) these model results by false positives with high confidence which generally indicates an error in the ground truth annotation. ```python 1from fiftyone import ViewField as F 2 3fp_view = dataset.filter_labels( 4 "predictions", 5 (F("confidence") > 0.8 & F("eval") == "fp"), 6) 7session = fo.launch_app(view=fp_view) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/e58fe3e99ad1cacbc038f0077f698a753b79fa32-959x789.png?auto=format&dpr=2&fit=max&q=75&w=959) This sample appears to be missing a ground truth annotation of skis. Let’s [tag it in FiftyOne](https://voxel51.com/docs/fiftyone/user_guide/app.html#tags-and-tagging), and send it to Labelbox for reannotation. ![](https://cdn.sanity.io/images/h6toihm1/production/f8c03248c2da5e63f7ed8c7dfad4e9684c511865-958x778.png?auto=format&dpr=2&fit=max&q=75&w=958) ```python 1view = dataset.match_tags("reannotate") 2anno_key = "fix_labels" 3label_schema = { 4 "ground_truth_edits": { 5 "type": "detections", 6 "classes": dataset.distinct("ground_truth.detections.label"), 7 } 8} 9view.annotate( 10 anno_key, 11 label_schema=label_schema, 12 backend="labelbox", 13) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/bb9b0a38f086828474d990f8b204efab07a6cf1f-1400x895.png?auto=format&dpr=2&fit=max&q=75&w=1400) ```python 1view.load_annotations(anno_key, cleanup=True) 2view.merge_labels("ground_truth_edits", "ground_truth") ``` Iterating over this process of training a model, evaluating its failure modes, and improving the dataset is the most surefire way to produce high-quality datasets and subsequently high-performing models. ## Additional Utilities You can perform [additional Labelbox-specific operations](https://voxel51.com/docs/fiftyone/integrations/labelbox.html#additional-utilities) to monitor the progress of an annotation project initiated through this integration with FiftyOne. For example, you can [view the status of an existing project](https://voxel51.com/docs/fiftyone/integrations/labelbox.html#viewing-project-status): ```python 1results = dataset.load_annotation_results(anno_key) 2results.print_status() ``` Project: FiftyOne\_quickstart ID: cktixtv70e8zm0yba501v0ltz Created at: 2021-09-13 17:46:21+00:00 Updated at: 2021-09-13 17:46:24+00:00 Members: User: user1 Role: Admin ID: ckl137jfiss1c07320dacd81l Nickname: user1 Email: USER1\_EMAIL@email.com User: user2 Role: Labeler Name: FIRSTNAME LASTNAME ID: ckl137jfiss1c07320dacd82y Email: USER2\_EMAIL@email.com Reviews: Positive: 2 Zero: 0 Negative: 1 You can also [delete projects](https://voxel51.com/docs/fiftyone/integrations/labelbox.html#deleting-projects) associated with an annotation run directly through the FiftyOne API. ```python 1results = dataset.load_annotation_results(anno_key) 2api = results.connect_to_api() 3 4print(results.project_id) 5# "bktes8fl60p4s0yba11npdjwm" 6 7api.delete_project(results.project_id, delete_datasets=True) 8 9# OR 10 11api.delete_projects([results.project_id], delete_datasets=True) 12 13# List all projects or datasets associated with your Labelbox account 14project_ids = api.list_projects() 15dataset_ids = api.list_datasets() 16 17# Delete all projects and datsets from your Labelbox account 18api.delete_projects(project_ids_to_delete) 19api.delete_datasets(dataset_ids_to_delete) ``` ## Summary No matter what computer vision projects you are working on, you will need a dataset. [FiftyOne](http://fiftyone.ai/) makes it easy to curate and dig into your dataset to understand all aspects of it, including what needs to be annotated or reannotated. In turn, the [integration with Labelbox](https://voxel51.com/docs/fiftyone/integrations/labelbox.html) makes this annotation process a breeze resulting in a dataset that will lead to higher-quality models. [annotation](https://voxel51.com/blog/tag/annotation) [data annotation](https://voxel51.com/blog/tag/data-annotation) [integrations](https://voxel51.com/blog/tag/integrations) [Labelbox](https://voxel51.com/blog/tag/labelbox) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/896cf6d1a00466cd9403a4af77cef2bddc75fe1e-1400x813.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ CVAT <> FiftyOne: Data-Centric Machine Learning with Two Open Source Tools\\ \\ Tutorials\\ \\ • \\ \\ Nov 30, 2022](https://voxel51.com/blog/cvat-fiftyone-data-centric-machine-learning-with-two-open-source-tools) [![](https://cdn.sanity.io/images/h6toihm1/production/cfa6067062cae206570d98a5e688951723545822-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why FiftyOne is the pandas of Computer Vision\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ Nov 23, 2022](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/049aec7818dc6abca36e9a736a9aa4e944a2b4ce-1200x671.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Sept 30, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Oct 1, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-30-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-241-lllmstxt|> ## Avatar: The Way of Water Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) The Making of Avatar: The Way of Water Jan 19, 2023 • 8 min read Article content In this article [How computer vision helped James Cameron wade into uncharted waters](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#867eb3a5be2f) [Stereoscopic fusion goes submersive](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#cb03e1faed2a) [Pushing performance capture to the limit](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#ee07de42f33e) [Moving the audience with motion grading](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#ddfef7286546) [Conclusion](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#fe81c7c4b998) [Wait, what’s FiftyOne?](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#e8f87c353c8a) [Join the FiftyOne community!](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#c1a0acf75ff4) [References](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#ca34520873f2) In this article [How computer vision helped James Cameron wade into uncharted waters](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#867eb3a5be2f) [Stereoscopic fusion goes submersive](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#cb03e1faed2a) [Pushing performance capture to the limit](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#ee07de42f33e) [Moving the audience with motion grading](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#ddfef7286546) [Conclusion](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#fe81c7c4b998) [Wait, what’s FiftyOne?](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#e8f87c353c8a) [Join the FiftyOne community!](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#c1a0acf75ff4) [References](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water#ca34520873f2) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## How computer vision helped James Cameron wade into uncharted waters ![](https://cdn.sanity.io/images/h6toihm1/production/e6ca14f73ab3cebac9a67e469fc0cb4f87f7b08b-960x640.jpg?auto=format&dpr=2&fit=max&q=75&w=960) Here at Voxel51 (the company behind the open source [FiftyOne](https://docs.voxel51.com/) computer vision toolset) we're computer vision practitioners and enthusiasts. As such, we're always on the lookout for new applications of computer vision. Recently, _Avatar: The Way of Water_ grabbed our attention for its stunning visuals and underwater cinematography. In this article, I’ll summarize some of the most interesting computer vision techniques the film employed to create such an immersive experience. After almost 13 years, James Cameron’s long-awaited sequel to _Avatar_ finally came out on December 16th, and it is on pace to surpass the first _Avatar_ film as the highest grossing box office production of all time. The protracted hiatus between installments was in large part due to the technical challenges director James Cameron and the _Avatar_ team took upon themselves to solve: chief among these being the difficulty of creating realistic scenes in water. Overcoming these hurdles involved the creation of new hardware, the development of novel machine learning algorithms, and humans performing physical feats they never thought possible. Computationally, the film required 18.5 petabytes of data, weeks of simulation for individual scenes, and millions of CPU hours to render the graphics [\[1\]](https://www.nytimes.com/2022/12/16/movies/avatar-2-fx-cgi.html). Here are some of the ways _Avatar: The Way of Water_ leveraged computer vision. ## Stereoscopic fusion goes submersive When one source of data is limited, why not use multiple? This is the idea behind [sensor fusion](https://en.wikipedia.org/wiki/Sensor_fusion), a technique used in computer vision to combine the strengths of multiple modalities or sources of data. Autonomous vehicles fuse together lidar, radar, and camera data, and medical imaging applications can synthesize the results from multiple anatomical scanning technologies, including CT, PET, MRI, and ultrasound. For the original _Avatar_, James Cameron and [Vince Pace](https://www.imdb.com/name/nm0655184/) developed a system with two cameras called [_Reality Camera System 1_](https://en.wikipedia.org/wiki/Fusion_Camera_System), which employed sensor fusion toward a different end: [stereoscopic vision](https://en.wikipedia.org/wiki/Stereoscopy). When a person looks at the world, they are synthesizing what the left eye is seeing with what the right eye is seeing. Our eyes can look in different directions and focus to varying extents, and the brain combines this information in real time to produce a three dimensional representation of our surroundings. ![](https://cdn.sanity.io/images/h6toihm1/production/a4bfaa3d12eb5bbce8e61cd0289620b05047beaa-740x444.jpg?auto=format&dpr=2&fit=max&q=75&w=740) With the fusion camera system, Cameron and Pace set out to bring the illusion of depth to the big screen. To accomplish this, the duo designed a system with two cameras in close proximity (inches away) but capable of being controlled independently. In contrast to the [binocular](https://en.wikipedia.org/wiki/Binocular_vision) nature of human vision, however, Reality Camera System 1 was designed with one camera mounted horizontally and the other mounted vertically, to give the filmmakers control over the degree of depth they wanted humans to experience. This invention contributed significantly to _Avatar_’s unique appearance in 3D. To take this technology under water, Cameron swapped Reality Camera System 1 for a new fusion system designed by cinematographer Pawel Achtel specifically for _Avatar: The Way of Water_ and its sequels. The new system, called [Deep X 3D](https://24x7.com.au/deepx-3d/), uses [submersible lenses](https://24x7.com.au/nikonos-lenses/) so that it can do away with the housing traditionally found in underwater cameras. In so doing, Cameron and Pawel eliminated camera-borne distortions from underwater stereoscopic fusion. ![](https://cdn.sanity.io/images/h6toihm1/production/a5bf73ccf184838e437364b7dbc7467c0ad3a7fc-1624x1092.png?auto=format&dpr=2&fit=max&q=75&w=1600) For scenes that took place both above and below the surface, the production crew had to take fusion a step further: > _“What we essentially wound up with was a volume…for underwater and a separate volume for the air. Those two volumes had to sit right on top of one another with only an inch in between…It was two completely separate methods of capture being fused together”_ > > _James Cameron_ [\[4\]](https://thewaltdisneycompany.com/how-avatar-the-way-of-water-revolutionizes-underwater-cinematography/) ## Pushing performance capture to the limit Over the past few decades, the film, sports, and video game industries have made use of technologies for capturing and recording the movements of people and objects known as [motion capture](https://filmlifestyle.com/what-is-motion-capture/), or “mocap”. The goals of motion capture are very similar to those of standard computer vision motion tracking tasks like [pose estimation](https://en.wikipedia.org/wiki/3D_pose_estimation) and [activity recognition](https://indatalabs.com/blog/human-activity-recognition). However, when used to animate digital characters in CGI scenes, motion capture typically requires that actors wear full-body suits with keypoint markers that are tracked by a stationary camera. This may not always be the case, as researchers are working towards markerless pose estimation. \[ [5](https://openaccess.thecvf.com/content_cvpr_2017/papers/Chen_3D_Human_Pose_CVPR_2017_paper.pdf), [6](https://openaccess.thecvf.com/content/CVPR2022/html/Li_MHFormer_Multi-Hypothesis_Transformer_for_3D_Human_Pose_Estimation_CVPR_2022_paper.html)\] But these techniques aren’t production ready just yet. For the first _Avatar_ film, the production crew created a motion capture “stage” on which the actors donned their suits and delivered their performances. The vast stage was [surrounded by 120 stationary cameras](https://www.studiobinder.com/blog/filming-avatar-and-avatar-2-behind-the-scenes/), recording the actors from all angles. When motion capture involves the recording of more subtle features such as hand gestures and facial expressions, it is often referred to as _performance capture_. In _Avatar_, cameras mere inches from the actors’ faces captured the detailed data used to render the faces of their Na’vi avatars. Water has a way of complicating performance capture - actually a few ways. First, the water reflects light, creating the illusion of additional markers. Second, any type of air bubble in water creates an unwanted disturbance. Notably, this includes the bubbles generated by humans when they wear scuba gear. This gear also impedes the camera’s view of the actors’ facial expressions. To solve the first of these problems, the entire tank in which filming took place was covered in floating white balls, preventing light reflections from breaching the surface. ![](https://cdn.sanity.io/images/h6toihm1/production/bf18bcf3e1b842db6201bbd1671a728672ccdab7-1200x800.jpg?auto=format&dpr=2&fit=max&q=75&w=1200) To address the other two complications, the entire cast received world class training in free diving, learning how to hold their breath for long periods under water. Kate Winslet set the cast record, [staying under for seven minutes](https://variety.com/2022/film/news/kate-winslets-filmed-avatar-2-underwater-breath-hold-record-die-1235459216/) before coming up for air. The result was a nearly perfect three dimensional performance capture stage. Using computer vision - what Cameron dubs his _Virtual camera_ \- he was able to map the motion of the actors onto the computer generated scene in real time: > _“I could see everybody where they’re supposed to be, above or below the water…they were acting to real-time direction based on what I was seeing on the Virtual Camera”_ > > James Cameron [\[4\]](https://thewaltdisneycompany.com/how-avatar-the-way-of-water-revolutionizes-underwater-cinematography/) ## Moving the audience with motion grading At a movie theater, 24 images typically flash on the screen each second. 24 ‘frames’ per second (fps) became standard for a few key practical and economic reasons: it is widely believed that 24fps is the lowest frame rate you can use and still put the smooth “motion” in “motion pictures”; and film stock wasn’t always cheap. This frame rate represented a kind of compromise between cost and quality. [\[8\]](https://www.redsharknews.com/technology-computing/item/3881-why-24-frames-per-second-is-still-the-gold-standard-for-film) Why precisely 24fps became _the_ standard rather than 23, 25, or some other number, likely boils down to luck and entrenchment (and 24 being a [neatly divisible number](https://en.wikipedia.org/wiki/Highly_composite_number))! While this “gold” standard is generally suitable for films consisting of actual images, for computer generated graphics, 24fps can still feel choppy. In real images, any natural noise that is present can help the eye gloss over gaps in motion. But no such noise exists in the pixel-perfection of CGI scenes. In order to create _Avatar: The Way of Water_’s smooth, photorealistic appearance, James Cameron insisted that many scenes - including all of the underwater scenes - were shot using the high frame rate ( [HFR](https://en.wikipedia.org/wiki/High_frame_rate)) of 48fps. But committing to HFR presented a new set of challenges. First and foremost, not all scenes in the film were shot in 48fps, leading to in-film switches between frame rates. Additionally, different screens, from televisions to movie theaters support different standards. Broadcast television, for instance, plays at about 30fps. When the frame rate is decreased from HFR to 24 or 30fps, the motion interpolation can create a ‘ [soap-opera effect](https://www.techtarget.com/whatis/definition/soap-opera-effect-motion-interpolation#:~:text=The%20soap%20opera%20effect%20is,original%20frames%20of%20a%20video.)’, where our eyes are searching for the blurring that accompanies objects in motion, to no avail. To facilitate in-film changes in frame rate and combat the soap-opera effect, _Avatar: The Way of Water_ employs [TrueCut Motion](https://www.pixelworks.com/en/truecut) by PixelWorks, which uses a suite of computer vision algorithms to exercise control over motion blurring and the apparent shaking or vibrating of images known as [judder](https://www.howtogeek.com/753131/what-is-judder-and-why-do-tvs-have-this-problem/). TrueCut Motion makes it possible to convert from 24fps to virtually any frame rate while preserving the desired cinematic quality. In effect this is a cinematic solution to frame interpolation. ## Conclusion _Avatar: The Way of Water_ is as much a technical feat as it is a work of art. Computer vision was integral to its creation, from facilitating underwater cinematography, to enabling real-time performance capture and visualization, to making it possible to overcome the high frame rate soap-opera effect in post-production. Computer vision will only play a bigger role in film and media in the coming years. If you want to better understand stereoscopic vision, check out the [KITTI 3D benchmark dataset](https://www.cvlibs.net/datasets/kitti/eval_object.php?obj_benchmark=3d), which contains left and right stereoscopic images from thousands of outdoor scenes, taken with closely positioned cameras. The open source library [FiftyOne](https://github.com/voxel51/fiftyone) makes it easy to download a [subset of this dataset](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#dataset-zoo-quickstart-groups) and get started working with the data. To learn more about how computer vision is affecting film and entertainment, along with other industries like sports, retail, autonomous vehicles, and healthcare, look out for our Industry Spotlight series, launching later this month! ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,275+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,400+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,600+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 228+ repositories - 55+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## References \[1\] [How ‘Avatar: The Way of Water’ Solved the Problem of Computer-Generated H2O](https://www.nytimes.com/2022/12/16/movies/avatar-2-fx-cgi.html) \[2\] [This is the Camera That Shot Avatar 2](https://ymcinema.com/2022/05/12/this-is-the-camera-that-shot-avatar-2/) \[3\] [Deep X 3D](https://24x7.com.au/deepx-3d/) \[4\] [How Avatar: the Way of Water Revolutionizes Underwater Cinematography](https://thewaltdisneycompany.com/how-avatar-the-way-of-water-revolutionizes-underwater-cinematography/) \[5\] [3D Human Pose Estimation = 2D Pose Estimation + Matching](https://openaccess.thecvf.com/content_cvpr_2017/papers/Chen_3D_Human_Pose_CVPR_2017_paper.pdf) \[6\] [MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation](https://openaccess.thecvf.com/content/CVPR2022/html/Li_MHFormer_Multi-Hypothesis_Transformer_for_3D_Human_Pose_Estimation_CVPR_2022_paper.html) \[7\] [Making of Avatar & Avatar 2: Behind-the-Scenes of James Cameron’s Epic](https://www.studiobinder.com/blog/filming-avatar-and-avatar-2-behind-the-scenes/) \[8\] [Why 24 frames per second is still the gold standard for film](https://www.redsharknews.com/technology-computing/item/3881-why-24-frames-per-second-is-still-the-gold-standard-for-film) [Avatar](https://voxel51.com/blog/tag/avatar) [computer vision applications](https://voxel51.com/blog/tag/computer-vision-applications) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [frame interpolation](https://voxel51.com/blog/tag/frame-interpolation) [KITTI Dataset](https://voxel51.com/blog/tag/kitti-dataset) [performance capture](https://voxel51.com/blog/tag/performance-capture) [sensor fusion](https://voxel51.com/blog/tag/sensor-fusion) [Stereoscopic vision](https://voxel51.com/blog/tag/stereoscopic-vision) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a735267ad7effa9f799f850ab7c8ffa241088710-1024x1024.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Why 2022 was the most exciting year in computer vision history (so far)\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 14, 2022](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far) [![](https://cdn.sanity.io/images/h6toihm1/production/7fabed74e8e4e741cb20a5be6a3c073bc6d71627-1400x787.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ On Notebooks and the Future of Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Feb 1, 2021](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/61b9a72c1362d209b3cb768c4ddc7f84cdf22452-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Automatically Set Up a New ML Project, Pain Free\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Feb 8, 2023](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) The Making of Avatar: The Way of Water - Voxel51 <|firecrawl-page-242-lllmstxt|> ## January 2023 Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup – January 2023 Jan 18, 2023 • 18 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#02e6245f3d65) [Computer Vision Meetup Recap at a Glance](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#88d4fde02979) [Hyperparameter Scheduling for Computer Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#9e20e7c0d8d5) [Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#87d698318648) [Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#6691f4996060) [An Intro to Computer Vision with Hugging Face Transformers](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#54d5bc6cbc12) [Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#b608d62d1015) [Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#2042d508bbc6) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#23aca35aef0f) [Upcoming Computer Vision Meetup Speakers & Schedule](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#71e55b5603e7) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#c36f0795a7d4) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#02e6245f3d65) [Computer Vision Meetup Recap at a Glance](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#88d4fde02979) [Hyperparameter Scheduling for Computer Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#9e20e7c0d8d5) [Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#87d698318648) [Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#6691f4996060) [An Intro to Computer Vision with Hugging Face Transformers](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#54d5bc6cbc12) [Q&A Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#b608d62d1015) [Additional Resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#2042d508bbc6) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#23aca35aef0f) [Upcoming Computer Vision Meetup Speakers & Schedule](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#71e55b5603e7) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#c36f0795a7d4) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Last week Voxel51 hosted the January 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/). In this blog post you’ll find the payback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. Hope to see you soon! ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes this month was [Foundation Fighting Blindness](https://www.fightingblindness.org/). We are pleased to be making a donation of $200 to them on behalf of the computer vision community! ![](https://cdn.sanity.io/images/h6toihm1/production/c1b81a240bde49092b376fffc3fdfd0717ca336b-600x278.png?auto=format&dpr=2&fit=max&q=75&w=600) ## Computer Vision Meetup Recap at a Glance ### Cameron R. Wolfe // Hyperparameter Scheduling for Computer Vision - [Video replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#hyperparameters-video) - [Presentation recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#hyperparameters-summary) - [Q&A recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#hyperparameters-q-a) - [Additional resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#hyperparameters-resources) ### Julien Simon // An Intro to Computer Vision with Hugging Face Transformers - [Video replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#hugging-face-video) - [Presentation recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#hugging-face-summary) - [Q&A recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#hugging-face-q-a) - [Additional resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#hugging-face-resources) ### Next steps - [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#locations) - [Computer Vision Meetup Speakers — February & March](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#schedule) - [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023#get-involved) ## Hyperparameter Scheduling for Computer Vision ### Video Replay https://www.youtube.com/watch?v=gwAk66W129Q ### Executive Summary [Cameron R. Wolfe](https://cameronrwolfe.me/), research scientist at Alegion and Ph.D. student at Rice University, shares some of his research on the topic of hyperparameter scheduling for computer vision. To begin, Cameron explains that hyperparameter tuning is important because deep learning can be computationally expensive. When you're training over a dataset, you go through all of the data from multiple epochs, and at the end you check whether or not your model performs well. If it doesn't, then you have to go back to stage one. This creates a loop where, if you don't select your hyperparameters properly, you’re continually retraining your network trying to get one that performs well. While many people attribute the computational expense of deep learning to big datasets and large models, the problem gets worse with incorrect hyperparameters because you have to incur the expense of training your deep neural network multiple times. In his presentation, Cameron describes how to use hyperparameter schedules, how to set them properly, and practical takeaways for hyperparameters across three main areas: learning rate, training precision, and video deep learning. #### Learning rate The learning rate part of the presentation is based on a paper Cameron wrote with Rice – [REX: Revisiting Budgeted Training with an Improved Schedule](https://arxiv.org/abs/2107.04197). The idea behind it is that when you consider different budget amounts for training a deep neural network, certain learning rate schedules work better than others. Why does budget matter? Maybe you have a computational and/or monetary budget and need to train within it; or maybe you have a deadline and can't spend too much time training your network. One of the most effective ways for working within a budget is to simply reduce your number of training epochs. Instead of training a model for 200 epochs on ImageNet, for example, you can shorten it down to 90 and you can probably get pretty good performance if you set your hyperparameters correctly. Cameron’s research looks at how to properly set the learning rate when training neural networks in a budget-aware setting. To explore different learning rate options in a budget-aware setting, Cameron decomposes learning rate schedules into two components: profile (the continuous function that models the decay of your learning rate; classic examples are – cosine, linear, exponential, and step schedules) and sampling rate (how frequently you want to update your learning rate from this profile). To set the scene, Cameron shows two classic examples – a step schedule and a linear schedule. ![](https://cdn.sanity.io/images/h6toihm1/production/ac05727b95d452ecdc560c71d50bd2a375194ac8-1854x688.png?auto=format&dpr=2&fit=max&q=75&w=1600) What Cameron’s research contributes is a new profile for learning rate decay: reverse exponential or REX, which is a learning rate schedule that keeps the learning rate high for a while and then decays it to a lower learning rate at the end of training. ![](https://cdn.sanity.io/images/h6toihm1/production/f97bfe6e926143803744af9ed121bf352c20277e-2504x746.png?auto=format&dpr=2&fit=max&q=75&w=1600) Now the question is: how do the learning rate and sampling rate pairings perform? To conduct this experiment, Cameron conducted experiments for each profile to find the optimal sampling rate across seven different domains. ![](https://cdn.sanity.io/images/h6toihm1/production/e1956ed89c4fc6cead5173e3a15e968a05c54b90-1652x822.png?auto=format&dpr=2&fit=max&q=75&w=1600) From this, the research takes the optimal profile and sampling rate pairings and sees how they compare on six different learning rate schedules. ![](https://cdn.sanity.io/images/h6toihm1/production/ca2d6119e3bcb43f49ca9fcdd014243af0c807f0-1378x740.png?auto=format&dpr=2&fit=max&q=75&w=1378) Key takeaways: - Step schedules (commonly used in computer vision) only perform well in the high-budget regime - REX performs really well across domains and budgets - Many choices exist for learning rate schedules; choose the best one based on your settings - number of epochs, domain, or problem (classification, detection, etc.) #### Training precision Because training deep neural networks is computationally expensive, another way to reduce costs would be to look at better schedules for low precision training for neural networks. Specifically, in this presentation and research he’s working on, Cameron looks at cyclical precision training (CPT) (not fixed or static training), where the precision being used for the neural network varies cyclically throughout the entire training process. The idea behind low precision training is that instead of training your neural network at the normal 32 bit precision, you can lower this to 16 bits, 8 bits, or maybe even lower, which would save compute costs. How this works is, when you do your forward pass, you quantize your activation and weights to a lower precision before performing the matrix's multiplication in the forward pass, which is faster. You can do the same thing in the backward pass, which has two matrix multiplications (one to compute your weight update and one to propagate the gradient to the previous layer), which saves twice as much compute. The question Cameron’s research addresses is: what are the best hyperparameter schedules for CPT to achieve gains (reduced cost or increased performance)? This research follows a similar approach to the REX research by decomposing the precision schedules into parts (in this case, three). ![](https://cdn.sanity.io/images/h6toihm1/production/d950bcff49d67638a833eb6e066b7322b0e1f50e-1672x832.png?auto=format&dpr=2&fit=max&q=75&w=1600) Cameron notes that the harder part is choosing repeated or triangular schedules. To better understand what is meant by repeated or triangular schedules, Cameron shows the set of 10 (repeated and reflected), which you can see below. ![](https://cdn.sanity.io/images/h6toihm1/production/0cb594b732d2a9b32a7b498e6d524ed03bf5cfdd-1574x756.png?auto=format&dpr=2&fit=max&q=75&w=1574) Cameron’s research collapses the 10 schedules into three groups (large schedules, medium schedules, and small schedules) based on the impact they have on the amount of compute savings you get. Here’s the mapping -> large = RR and RTH; medium = LR, LT, CR, CT, RTV, ETV; small = ER, ETH. (You may want to come back to these mappings later when you review the precision takeaways later in the post.) ![](https://cdn.sanity.io/images/h6toihm1/production/0faffe72219998465ea4bfd5d5b84eb8d7c300bc-2450x1296.png?auto=format&dpr=2&fit=max&q=75&w=1600) Cameron walks us through some of the highlights from the exploration: - There tends to be a correlation between the model performance and the amount of training compute that we're using. Cameron adds “if we use less compute or have more computational savings, we tend to get a little bit worse performance and vice versa. So even though we can use these alternative precision schedules, in a lot of cases we're not getting these compute benefits for free. If we want our training to go by faster, we're probably going to get a model that performs not quite as good.” - In certain cases using alternative schedules (beyond just cosine schedules) can be beneficial. For example, when training a ResNet 18 on ImageNet, the best performance actually comes with the exponential schedule. - Some settings are sensitive to lots of quantization, so you can't always train a network with really low precision; it depends on the domain. So be careful with how low the precision is with which you’re going to be training. Key takeaways for training precision are: - You can gain a lot by exploring alternative hyperparameter schedules for cyclic precision training. - Here’s Cameron’s guidance on how to choose one: - Use small schedules to minimize training costs - Use large schedules to maximize model performance - Use medium schedules to find a balance Now for a twist: none of these can actually be used practically right now because current hardware only supports low precision training at certain levels. Cameron recommends “the best thing to use for low precision training right now is just something called Apex,” which implements 16 bit precision training with minimal code changes required. For more information, check out Cameron’s article: [Quantized Training with Deep Networks](https://cameronrwolfe.substack.com/p/quantized-training-with-deep-networks-82ea7f516dc6). #### Video deep learning The last part of Cameron’s presentation is about video deep learning and how cyclical batch sizes can speed up wall-clock training time. In this section, Cameron walks us through some background found in the [SlowFast Networks for Video Recognition](https://arxiv.org/abs/1812.03982) paper. The idea is that instead of having just a single 3D CNN that we convolve over the entire input video over multiple layers, we separate our network into two different modules: the Slow Pathway, which is in charge of capturing spatial or semantic features; and the Fast Pathway that's in charge of capturing motion features. ![](https://cdn.sanity.io/images/h6toihm1/production/3c1f6f0f410ef4967da6942f8c78f2063df8cb6d-2356x1138.png?auto=format&dpr=2&fit=max&q=75&w=1600) The answer proposed in the paper, [A Multigrid Method for Efficiently Training Video Models](https://arxiv.org/abs/1912.00998), is to vary the mini-batch size according to a hyperparameter schedule. ![](https://cdn.sanity.io/images/h6toihm1/production/05c710f1556319354dbe5a8be874b1f24b6f9727-2502x1092.png?auto=format&dpr=2&fit=max&q=75&w=1600) In the image below on the right, you can see the results for training over a Kinetics dataset on a single GPU using this multigrid training method. This leads to improved efficiency – the same performance in just two days instead of nearly an entire week! ![](https://cdn.sanity.io/images/h6toihm1/production/c2c212e3767af9c82e147f73b5f23ccdc71e3dbc-2444x1114.png?auto=format&dpr=2&fit=max&q=75&w=1600) #### Conclusion Hyperparameter schedules are fundamental. You can apply them in different scenarios and they end up being useful for finding a benefit (reduced compute, increased performance, reduced training time, etc.) ## Q&A Recap Here’s a recap of the live Q&A from this presentation during the virtual Computer Vision Meetup: **Q: Any comment on REX in regards to transfer learning?** A: We didn't test that in this case, but that would be a really interesting test to run. A lot of times what I've found is that when you're fine tuning a network on a downstream dataset, if you start with too high of a learning rate, that might cause problems. So my guess would be that REX is fine for transfer learning but you have to be careful with setting the initial learning rate to make sure that it's not too high. But again, that would be pretty cool to test. **Q: In addition to problem domain and budget, do you see that just the dataset itself (size, diversity, etc.) can cause one LR to work better or worse?** A: Yes, all of these things are factors that you have to consider when choosing hyperparameters. So we can boil this down to the question of if you're training a neural network on one dataset and you switch it and try to train it on another dataset, are you going to use the same hyperparameters out of the box? Probably not. You're going to have to test a few things and see what works. When you change domains or datasets, you're going to have to do more tuning and figure out what’s best for your new dataset. **Q: When do you think we will see convex optimizers for arbitrary neural networks?** Neural networks like by nature are nonconvex. But at the same time, even though the optimization problem with neural networks is nonconvex, for a lot of optimization algorithms, all of the proofs are in convex settings or simplified settings. The literature for analyzing nonconvex convergence rates is a lot harder than the convex proof. So typically it seems like for deep learning we can take intuitions from convex optimization and see whether they'll work. And in the case of SGD, for example, it works really well. **Q: How are we initiating REX?** With REX, you choose an initial learning rate, you have a schedule by which that learning rate is varied, and then from the beginning to the end of your training, you decay the learning rate from the initial choice to kind of 10x lower than that initial choice or maybe 100x according to that schedule. So there's no initiation needed, it's just a fixed function profile by which we decay our learning rate. **Q: In addition to randomizing training sets, do you insert intentional "disturbances"?** A: No, really all we do here is pick a bunch of different domains, a bunch of different publicly available datasets, and then run training across all of these to see which learning rate schedules work best. All of the setups are pretty standard I would say; we're not doing anything extra just running training over public datasets. **Q: Low Precision training, is that also called Quantization Aware Training?** A: Quantization aware training refers to methods that make your network perform better when you make it smaller at the end so that you can deploy it onto an edge device, for example. The purpose of that is training the network in a way that when you quantize it, it will still perform well at the edge because when you convert it to a low precision representation, that uses less memory so that you can deploy it. The difference between that and what I'm doing here is that cyclical precision training (CPT) is not focused on trying to create a smaller neural network at the end, it's focused upon trying to reduce training costs in general. We're performing the training process in a low precision manner so that we can save compute costs during training. We're not necessarily trying to quantize the network to make it smaller at the end. Though, there could be a relationship between low precision training and quantization aware training such that performing low precision training like this could be a good method of quantization aware training to where these networks are easier to quantize at the end. **Is CPT theoretical (e.g. implemented entirely in software)?** A: Correct, there's no hardware support for this yet. I hope that hardware will catch up with it eventually, but right now the best thing to use is Apex because you can train it at floating point 16 precision, it’s easy to implement, and it’ll increase speed without losing any performance, though check to make sure. **Q: Are there trade offs/disadvantages for Low Training Precision?** Yes. There's a correlation between model performance and training compute. So if you're lowering the precision too much, you're going to pay the cost – your network is going to deteriorate in performance. This is relative to training with a fixed, but lower precision level. But if you adopt a different strategy that doesn't quantize quite as aggressively, you might actually improve model performance while saving compute. In the talk I shared an example of an ImageNet case compared to the baselines, where the exponential precision schedules both reduced compute compared to baselines and improved the performance. So it depends; there is a trade off if you quantize too much, but there are certain cases where you can actually get better performance and save training costs. It just depends on the scenario. ## Additional Resources Check out the additional resources on the presentation: - [Talk transcript](https://www.rev.com/transcript-editor/shared/cJ4DfGaeFipg3-7HA7_H3pOSCL696oefkEJqbfCXi3HrV3VhrMQafk07VUeiVbwEPVEZzwV4Wos-7boDe_FbtjTC9bk?loadFrom=SharedLink) - [Presentation links](https://gist.github.com/wolfecameron/1f5da9ec7a365828d4ff000148e81cab) (to papers, articles, etc.) - [All links](https://cameronrwolfe.me/links) (newsletter, Twitter, company, lab, etc.) Thank you Cameron on behalf of the entire Computer Vision Meetup community for sharing your research and helping us better understand how to approach the tuning of hyperparameters! ## An Intro to Computer Vision with Hugging Face Transformers ### Video Replay https://www.youtube.com/watch?v=yGUQSa6Emnc ### Executive Summary Julien Simon’s talk provides an overview of what it means to work with computer vision models, especially those hosted on the [Hugging Face](https://huggingface.co/) Hub. Julien starts by sharing a brief history of computer vision in deep learning (in fact, what got him into deep learning was computer vision!). For a long time, working on computer vision applications meant working with convolutional neural networks (CNNs). But, Julien notes, “the golden age of CNNs is probably behind us.” Why? Enter the Vision Transformer (ViT), originally published by Google in 2021 and described in the paper: [An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale](https://arxiv.org/abs/2010.11929). Julien summarizes how the ViT works: ViT breaks an image into patches, which are flattened and treated as a sequence of tokens. This results in several key benefits. First, Julien shares a benchmark from the research paper, specifically one that compares the vision transformer to ResNet-152, which was state of the art at the time, and it performed quite a bit better. Also in that paper, the ViT was four times less compute intensive than training ResNet-152. Additionally, transfer learning is built in, meaning you can train your model initially on one dataset, learn the embeddings, and then specialize that model for downstream tasks. Finally, ViT offers state-of-the-art-accuracy. CV Transformers are gaining in popularity, but historically have been difficult to work with. This opened up an opportunity for Hugging Face to build open source libraries, a website, and commercial tools that make it very easy to find, share, and work with state-of-the-art models, even if you're not an expert. Here’s a sampling of Hugging Face by the numbers: - 120,000+ models - 18,000+ datasets - 25+ ML models including, Keras, Scikit-Learn models, fastai, and more Zooming in on computer vision, Julien notes that Hugging Face hosts about 4,000 computer vision models covering a variety of tasks across four broad categories. ![](https://cdn.sanity.io/images/h6toihm1/production/3cd1056485bdb502bc598369134a893f11617ea1-2664x1048.png?auto=format&dpr=2&fit=max&q=75&w=1600) Julien then heads over to the Hugging Face Hub to show us how easy it is to download, train, and deploy models. He also shows the inference widget where you can predict images in the Hub, as well as links to datasets that the model has been trained on and [Spaces](https://huggingface.co/spaces), where you can quickly build, host, and share your ML apps. Live demo time! Starting at ~16:49 in the presentation, Julien shows an app he built in Spaces that predicts an image (prediction happens with just two lines of Python code!) with 10+ models and displays the results. But what about more advanced models? Julien picks three – image captioning, zero shot segmentation using a text prompt, and audio classification using spectrogram – and shows several interactive live demos on Spaces. Then he covers multi-modal transformers, and diffusers, all with live demos and code snippets. Experimenting with models has never been easier! In addition to working with inferences and models, you can also train and deploy with Hugging Face. Julien walks us through the “family picture” that includes all the goodies available from Hugging Face today. ![](https://cdn.sanity.io/images/h6toihm1/production/1d39cc3b9fe1fb0f527bd37ea9bd2cd2abe4398a-2878x1274.png?auto=format&dpr=2&fit=max&q=75&w=1600) Before opening up for Q&A, Julien shares some links to get you started: - Tasks: [https://huggingface.co/tasks](https://huggingface.co/tasks) - Training course: [https://huggingface.co/course/](https://huggingface.co/course/) - Docs: [https://huggingface.co/docs](https://huggingface.co/docs) - GitHub: [https://github.com/huggingface](https://github.com/huggingface) - Forum: [https://discuss.huggingface.co/](https://discuss.huggingface.co/) - Support: [https://huggingface.co/support](https://huggingface.co/support) ## Q&A Recap **Q: Is ViT already using production somewhere?** A: Yes, we have customers who use a vision transformer in production. If you’re wondering about potential challenges in production, I can see two. The first is fine tuning the model on your own data, but that's easy enough to do. The second challenge you might run into is inference latency. If you have really low latency requirements or if you're deploying models at the edge, then transformers may still be a little bit too big (as compared to CNNs). Although, in the classifier space, you'll find models that are just a few megabytes; they have small mobile compatible versions of vision transformers. **Q: Can we target edge devices for Hugging Face models?** A: Yes. Although the hardware requirements for some transformer models can still be overwhelming for edge devices, there's a lot of work going on there to shrink the models (ex: you can find models that are 2-3 megabytes). Reductions are happening by using smaller architectures, or through techniques like quantization, pruning, or through optimization tools, etc. In addition, if you need to consider transformer models at the edge, look for a hardware solution that will meet your needs and fit within your hardware budget. **Q: When playing with applications in Hugging Face Spaces, are user uploaded images ultimately archived, or deleted, or something else?** A: Nothing is stored in a Space; there's no data storage, there's no data usage, and none of the data you send either on Spaces or on inference widgets, etc. is stored. Unless you want your Space to store something there somewhere, there's no other code than the code you're able to see in the Space. **Q: You mentioned 4x better efficiency for ViT compared to ResNet. Is that training? If so, is that training from scratch or transfer learning from a pretrained model, or both?** A: Yes, in the research paper, they reported a 4x training speed up compared to ResNet-152, and that was training from scratch. Generally, there seems to be some consensus, based on what I read in multiple papers, that transformers are a fit when you have a ton of data. For example, if you have 10M+ images and you train from scratch, you'll get better results with transformers than CNNs. On the other hand, if you have only, let's say, a hundred thousand images to train initially, CNNs tend to do better. However, one of the big benefits of transformers is transfer learning. So, unless you have a super exotic domain, I would encourage you to start from one of those pre-trained models, evaluate them on your own data, and fine tune a little bit. **Q: Can you compare the size of transformers to equivalent CNNs?** The vision transformer (going from memory) is about 300 to 400 megabytes, which is pretty large. But like I said, you can find some downscaled versions, even down to tens of megabytes or less; small enough that you could fit on a small device. ## Additional Resources Check out these additional resources: - [Presentation slides](https://www.slideshare.net/JulienSIMON5/an-introduction-to-computer-vision-with-hugging-face) - [Talk transcript](https://www.rev.com/transcript-editor/shared/ihBSiawy0qmQMIogE0g8diV_I7nXExjGHyKB96j-w9uXhEFDtAvhYMpyMv5f1X-Xi0QGC74MC5hIWruMeHNRzGSXINE?loadFrom=SharedLink) A big thank you to Julien on behalf of the entire Computer Vision Meetup community for getting us up-to-speed with Hugging Face Transformers! ## Computer Vision Meetup Locations Computer Vision Meetup membership has grown to more than [2,500+ members](https://www.meetup.com/pro/computer-vision-meetups/) in just a few months! The goal of the meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. New Meetup Alert – We just added a Computer Vision Meetup location in Singapore! Join one of the (now) 13 Meetup locations closest to your timezone. - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) ## Upcoming Computer Vision Meetup Speakers & Schedule We have an exciting lineup of speakers already signed up for February and March. Become a member of the Meetup closest to you, then register for the Zoom for the Meetups of your choice. ### February 9 - Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO — [Paula Ramos, PhD](https://www.linkedin.com/in/paula-ramos-41097319/) (Intel) - Understanding Speech Recognition with OpenAI’s Whisper Model — [Vishal Rajput](https://www.linkedin.com/in/vishal-rajput-999164122/) (AI-Vision Engineer) - [Zoom Link](https://us02web.zoom.us/webinar/register/9216708727643/WN_P8UHtAZGQWOx_A2HM1dcPA) ### March 9 - Lighting up Images in the Deep Learning Era — [Soumik Rakshit](https://www.linkedin.com/in/soumikrakshit/), ML Engineer (Weights & Biases) - Training and Fine Tuning Vision Transformers Efficiently with Colossal AI — [Sumanth P](https://www.linkedin.com/in/sumanth-p-09b339173/) (ML Engineer) - [Zoom Link](https://us02web.zoom.us/webinar/register/8816708728020/WN_mTdNXxTSR-e7bDG5EkH1XQ) ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [Hugging Face](https://voxel51.com/blog/tag/hugging-face) [hyperparameter scheduling](https://voxel51.com/blog/tag/hyperparameter-scheduling) [hyperparameters](https://voxel51.com/blog/tag/hyperparameters) [vision transformers](https://voxel51.com/blog/tag/vision-transformers) [ViT](https://voxel51.com/blog/tag/vit) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ed751b5c0fbc3d2cb74f0b30e6418d3319564c-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — December 2022\\ \\ Event Recaps\\ \\ • \\ \\ Dec 13, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/bbb1d9add0b0b9aa12682acac795df7c2ba760a9-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — November 2022\\ \\ Event Recaps\\ \\ • \\ \\ Nov 16, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/e56436d38d294978c25356ba4c5482b53a29afb3-960x540.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup – February 2023\\ \\ Event Recaps\\ \\ • \\ \\ Feb 14, 2023](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-243-lllmstxt|> ## FiftyOne: Data Experimentation Tool [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Introducing FiftyOne: A Tool for Rapid Data & Model Experimentation Sep 12, 2020 • 6 min read Article content In this article [The Backstory](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#3126f0e63b9b) [Getting Closer To Your Data](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#c777d9cf96e3) [Introducing FiftyOne](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#894c603f6203) [Loading Data With Ease](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#242050f8f144) [Powerful Dataset Analysis](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#009ca27a0ab9) [Interactive Model Evaluation](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#f4b6cf062ddd) [Less Wrangling, More Science](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#a34540675d03) [We Believe In Open Source Software](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#695af95b8e89) In this article [The Backstory](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#3126f0e63b9b) [Getting Closer To Your Data](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#c777d9cf96e3) [Introducing FiftyOne](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#894c603f6203) [Loading Data With Ease](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#242050f8f144) [Powerful Dataset Analysis](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#009ca27a0ab9) [Interactive Model Evaluation](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#f4b6cf062ddd) [Less Wrangling, More Science](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#a34540675d03) [We Believe In Open Source Software](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation#695af95b8e89) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _A new (and open source!) tool for your machine learning toolbox_ ![](https://cdn.sanity.io/images/h6toihm1/production/3d28e0671a1080a123d766887ed1818c6534496e-1446x1100.gif?auto=format&dpr=2&fit=max&q=75&w=1446) ## The Backstory My co-founder [Jason](https://www.linkedin.com/in/jason-corso-95b66a24/) and I started [Voxel51](https://voxel51.com/) in 2017 with the vision of building tools that enable CV/ML engineers to tackle the hardest problems in computer vision. We started that journey by participating in the [NIST Public Safety Innovation Accelerator Program](https://www.nist.gov/ctl/pscr/funding-opportunities/past-funding-opportunities/psiap-2017), which was created to incubate new technologies with the potential to transform the future of public safety. Over the next two years, we set out to translate our academic research on image/video understanding — over 250 papers and 30 years of experience — into a scalable platform for developing and deploying ML models that process visual data. The platform made it easy to take a model trained on images and deploy it to process video, efficiently and at scale. We trained a variety of models ourselves on road scene videos for tasks such as vehicle recognition, road sign detection, and human activity recognition. With some early successes under our belt, we raised a round of venture capital, built a small team, and began pilots with industry partners to onboard their CV/ML teams to our platform. During this process, we learned a surprising fact: There is a serious lack of tooling available to rapidly experiment with data and models. We found ourselves building our own in-house tools to solve tasks like: - Wrangling datasets into a common format for training/evaluation - Choosing a diverse set of images to annotate - Balancing datasets across classes and visual characteristics - Validating the correctness of human annotations - Visualizing model predictions - Finding and visualizing failure modes of models Each time we encountered a new task, we wrote more custom scripts, massaged our datasets in similar-but-different ways, and generally found ourselves spending way more time wrangling data than doing actual science to improve our models. ![](https://cdn.sanity.io/images/h6toihm1/production/d65ceafedc84c1ed4f2266c583b24148dcbb507e-1400x700.png?auto=format&dpr=2&fit=max&q=75&w=1400) After cross-referencing our experience with CV/ML teams in dozens of other companies, both small and large, as well as academic groups, we learned that others were experiencing the same pain: they needed a tool that would enable them to **rapidly experiment** with their data and models. ## **Getting Closer To Your Data** Training great machine learning models requires high quality data. Academic education in CV/ML is an excellent way to develop knowledge of model architectures, tips & tricks for training models, etc. However, most research treats the dataset (often a standard dataset like ImageNet or COCO) as a static, black box. In our experience, the limiting factor of performance on real problems is not the model, but rather the dataset. The more data you have for training, the better, right? End of story. Well, sort of: > _Nothing hinders the success of machine learning systems more than poor-quality dat_ a. _\- Jason Corso, CEO of Voxel51 and Professor of CV/ML_ In [a recent blog post](https://towardsdatascience.com/i-performed-error-analysis-on-open-images-and-now-i-have-trust-issues-89080e03ba09), we found that over 1/3 of false positives of state-of-the-art models on the Open Images dataset are actually due to annotation error! Scientists are busy tweaking model architectures to squeeze a few extra points of mAP out of a dataset whose limiting factor is annotation quality. ![](https://cdn.sanity.io/images/h6toihm1/production/8b6c7aaf4b290c763c360af135bc1301e8470025-1259x947.jpg?auto=format&dpr=2&fit=max&q=75&w=1259)![](https://cdn.sanity.io/images/h6toihm1/production/1c373bf3268514b55c531e3bd6ce55deb0cdd6d1-1262x402.png?auto=format&dpr=2&fit=max&q=75&w=1262) A consistent theme arose from our customer discovery efforts: the diversity of training data and the accuracy of the labels associated with that training data are critically important to developing a high-performance ML system. _That makes sense, but don’t ML engineers already know this?_ **Not nearly as much as we thought they did.** We found that CV/ML engineers focus primarily on choosing/tweaking model architectures and tuning hyperparameters. These are all important, but guess what? In our experience and that of many of the highest performing ML teams we spoke with, the biggest breakthroughs came by focusing on the data used to train the systems. _So, why aren’t CV/ML engineers acting on this insight?_ **Lack of tooling.** Seriously. We found that, while the teams and organizations with the most ML experience have built their own tools to make it easier to visualize their data, they only built those tools because they had no other choice. And smaller teams with less computer vision expertise? Most do not have the resources to build tools in-house (they are correctly focused on getting to market as fast as possible), and many do not yet realize the importance of getting closer to their data. ## Introducing FiftyOne [We built FiftyOne](https://voxel51.com/docs/fiftyone/) to help CV/ML engineers and scientists spend dramatically less time wrangling data and more time focusing on the science of building better datasets and better models. The vignettes below give a taste of what I’m talking about. ## Loading Data With Ease FiftyOne removes the effort required to [quickly load and visualize data](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html) in a variety of common (or custom) formats. While this task is certainly doable by any ML engineer, it takes real effort to properly wrangle the various data/annotation formats out there into a standard format that they can work with. FiftyOne removes that effort and meets engineers where they’re already working: in Python. With FiftyOne, it takes two lines of code to load an image dataset with labels and display them visually in the FiftyOne App. ```python 1import fiftyone as fo 2 3# Load a dataset stored in COCO format 4dataset = fo.Dataset.from_dir( 5 dataset_dir="/path/to/coco-formatted-dataset", 6 dataset_type=fo.types.COCODetectionDataset, 7) 8 9# Explore the dataset in the App 10session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/b00d8e3cd10f18f43556e75ee86dd7f2abce5eff-1349x1009.gif?auto=format&dpr=2&fit=max&q=75&w=1349) ## Powerful Dataset Analysis FiftyOne does more than just visualize your data. It also provides [powerful dataset analysis utilities](https://voxel51.com/docs/fiftyone/user_guide/brain.html). For example, look how easy it is to find visually similar images in your dataset: ```python 1import fiftyone.brain as fob 2 3# Index the images in the dataset by visual uniqueness 4fob.compute_uniqueness(dataset) 5 6# View the least unique (most visually similar) images in the App 7session.view = dataset.sort_by("uniqueness") ``` ![](https://cdn.sanity.io/images/h6toihm1/production/94093cbedd51fc8de47650dfaff111b8b24c2890-1330x1117.gif?auto=format&dpr=2&fit=max&q=75&w=1330) ## Interactive Model Evaluation When working with image/video datasets, some tasks are best performed visually, while others are best performed programmatically. We built FiftyOne to enable seamless handoff between the App and the Python library. You can load a dataset from code, then search and filter it from the App to identify particular samples/labels in the App, then access those samples of interest back in Python. One important instance of this workflow is evaluating model predictions. Quantitative measures such as [confusion matrices](https://en.wikipedia.org/wiki/Confusion_matrix) or [mAP scores](https://medium.com/@jonathan_hui/map-mean-average-precision-for-object-detection-45c121a31173) don’t tell the whole story of a model’s performance. You need to visualize the failure modes of your model to understand what actions to take to remove the glass ceiling on the model’s performance. With FiftyOne, you can [perform complex operations](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#filtering) on your data, such as evaluating the quality of a particular class at a given confidence threshold for your model’s predictions, and then visualizing the worst performing samples in the App: ```python 1from fiftyone import ViewField as F 2 3# Creates a view into the dataset that contains samples with model predictions 4# in their `faster_rcnn` field. From those samples, only consider objects in 5# the `ground_truth` field whose label is "person", and only consider 6# predictions in the `faster_rcnn` field whose label is "person" and whose 7# confidence is at least 0.80 8person_view = ( 9 dataset 10 .exists("faster_rcnn") 11 .filter_labels("ground_truth", F("label") == "person") 12 .filter_labels("faster_rcnn", (F("label") == "person") & (F("confidence") > 0.80)) 13) 14 15# Evaluate the accuracy of the predictions in the `faster_rcnn` field with 16# respect to the ground truth predictions in the `ground_truth` field 17person_view.evaluate_detections("faster_rcnn", gt_field="ground_truth") 18 19# Visualize the results in the App, showing samples with the most false 20# positives @ IoU 0.75 first 21session.view = person_view.sort_by("fp_iou_0_75", reverse=True) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/ed0b9dc4072cfa3d1d1bd2837c0b834fccec607f-1400x775.png?auto=format&dpr=2&fit=max&q=75&w=1400) ## Less Wrangling, More Science While we’ve heard that visualizing datasets in the App and the power of the Python library alone are already useful to CV/ML engineers, we’re also building out features in FiftyOne to power common workflows that are important but-challenging to implement today: - Converting between dataset formats - Curating diverse and representative datasets - Selecting video frames for annotation and training of image models - Automatically finding label mistakes - Evaluating model predictions - Identifying and visualizing failure modes in your models Ready to jump in? Check out the easy-to-follow tutorials and recipes in [the FiftyOne Docs](https://voxel51.com/docs/fiftyone/) to learn how to execute these tasks on your datasets. Oh, and [getting started with FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html) is a breeze! ```python 1# Install FiftyOne 2pip install fiftyone 3 4# Downloads a dataset and opens it in the App for you to explore! 5fiftyone quickstart ``` ## We Believe In Open Source Software At [Voxel51](https://voxel51.com/), we’re strong believers in the [virtues of open source software](https://techcrunch.com/2019/01/12/how-open-source-software-took-over-the-world/). As a point of reference, many of the most popular tools in the ML ecosystem — TensorFlow, PyTorch, Apache Spark, and MLflow, to name a few — are open source. We believe this is not a coincidence: the best way to build a community around a developer tool is to make the project as open and transparent as possible. Making a project free and open source removes barriers to entry and allows individual developers to directly evaluate the merits of a tool, cast a vote of approval for the project with their code (and GitHub stars), and evangelize the tool within their organizations. That’s why we released FiftyOne [open source on GitHub](https://github.com/voxel51/fiftyone) with a permissive license. We’re committed to making the best tools for rapid data and model experimentation openly available for all CV/ML engineers to use. Issues, feature requests, and pull requests are welcome! Although the core FiftyOne library will always be free and open source, we know that professional organizations and other advanced users have unique requirements to adopt tools more broadly and integrate them into their production workflows. That’s why we’re pursuing an **open core model** for FiftyOne. But more on that soon… For now, we’re excited to see the amazing CV/ML-powered solutions that are built using FiftyOne! [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [product release](https://voxel51.com/blog/tag/product-release) ![](https://cdn.sanity.io/images/h6toihm1/production/8d61ff90b31d151405f9e21a33c2802509f34651-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Brian Moore Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a4c2bee9ed053c5be2a1c161e5abf758c9a12ff8-1400x923.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.18 with App Performance Improvements, Sidebar Modes, and Custom Attributes\\ \\ Product & News\\ \\ • \\ \\ Nov 15, 2022](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes) [![](https://cdn.sanity.io/images/h6toihm1/production/e94f20fa81716294c7a6caccf7256e9106cb6e89-967x800.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.17 with Grouped Datasets, 3D, Geolocation, and Custom Plugins\\ \\ Product & News\\ \\ • \\ \\ Sep 21, 2022](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins) [![](https://cdn.sanity.io/images/h6toihm1/production/8651d28f4a2978eb72cca25ef09e1f01f81847ea-2968x2042.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.19 with Spaces, In-App Embeddings Visualization, Saved Views, and More!\\ \\ Product & News\\ \\ • \\ \\ Feb 16, 2023](https://voxel51.com/blog/announcing-fiftyone-0-19) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-244-lllmstxt|> ## FiftyOne Anniversary Celebration [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) FiftyOne Turns One! Aug 12, 2021 • 1 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _The dataset curation and model analysis tool by and for the open source community_ Today is the first anniversary of open-sourcing [FiftyOne](https://fiftyone.ai/)! We are super excited and grateful to the 40,000 early adopters that are using FiftyOne to build data-centric workflows for their computer vision projects. Thank you! \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop The era of model-centric machine learning is gone. FiftyOne users have grokked this shift towards the increasing importance that data plays in machine learning. Our early adopters have enabled FiftyOne to become a leading tool for dataset curation and model analysis in computer vision; its open-source nature is facilitating easy and early adoption in myriad verticals like healthcare, surveillance, advertising, and autonomous driving and across environments like individual workstations, cloud and on-premises clusters. In each of these contexts, FiftyOne is unblocking the data quality and model bottlenecks allowing users to rapidly improve performance, with an emphasis on the critical role data plays in modern AI. The community interest, support and involvement for FiftyOne has been amazingly helpful over the last year as we have extended the data-centric capabilities to include features like Jupyter notebook support, embeddings-based analysis capabilities, tight integrations with major datasets, like [Google’s Open Images](https://towardsdatascience.com/googles-open-images-now-easier-to-download-and-evaluate-with-fiftyone-615ce0482c02) and [COCO](https://medium.com/voxel51/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation-68a3d7e97fb7), as well as tools like [Lightning Flash](https://towardsdatascience.com/open-source-tools-for-fast-computer-vision-model-building-b39755aab490), CVAT, Labelbox and Scale AI. This one year anniversary comes along with the [FiftyOne v0.12 release](https://voxel51.com/docs/fiftyone/getting_started/install.html#installing-fiftyone), which delivers a fully optimized user interface allowing datasets to scale to hundreds of thousands of samples without a performance hit. As we roll out [FiftyOne Teams](https://voxel51.com/#teams-form), the enterprise version of FiftyOne, we are intensely thankful to the open-source early adopters who have helped us grow the project. And, of course, the open-source version of FiftyOne is here to stay, and will remain free, forever, as we continue to enhance its functionality. The team at Voxel51 is looking forward to an exciting year ahead! [anniversary](https://voxel51.com/blog/tag/anniversary) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [open source](https://voxel51.com/blog/tag/open-source) [Voxel51 milestone](https://voxel51.com/blog/tag/voxel51-milestone) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/97da003bade5b163c35a6018a0f491d3df78ab29-1400x1060.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne, Six Months Post-Launch\\ \\ Product & News\\ \\ • \\ \\ Feb 4, 2021](https://voxel51.com/blog/fiftyone-six-months-post-launch) [![](https://cdn.sanity.io/images/h6toihm1/production/869a02098d1898869a250f4a5a23648c479af7da-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – April ‘23\\ \\ Product & News\\ \\ • \\ \\ Apr 6, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/338b38d41e6072dd11af86f21f5309337c52f36b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – May ‘23\\ \\ Product & News\\ \\ • \\ \\ May 5, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-245-lllmstxt|> ## COCO Dataset Best Practices [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) The COCO Dataset: Best Practices for Downloading, Visualization, and Evaluation Jun 30, 2021 • 4 min read Article content In this article [Setup](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation#aede66df77c2) [Downloading COCO](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation#966c30b7fd23) [Visualizing the dataset](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation#9a634c6ce1c7) [Evaluating models](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation#0f6e2749620d) [Summary](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation#e5aa4547e709) In this article [Setup](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation#aede66df77c2) [Downloading COCO](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation#966c30b7fd23) [Visualizing the dataset](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation#9a634c6ce1c7) [Evaluating models](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation#0f6e2749620d) [Summary](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation#e5aa4547e709) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _How to use FiftyOne’s native support for COCO to power your workflows_ ![](https://cdn.sanity.io/images/h6toihm1/production/c57e2ba433b1b203c618e70f24f88af0cdd3d58a-1400x1092.png?auto=format&dpr=2&fit=max&q=75&w=1400) [The COCO dataset](https://cocodataset.org/#home) has been one of the most popular and influential computer vision datasets since its release in 2014. It serves as a popular benchmark dataset for various areas of machine learning, including object detection, segmentation, keypoint detection, and more. Chances are that your favorite object detection architecture has pre-trained weights available from the COCO dataset. **This post describes how to [use FiftyOne](https://voxel51.com/docs/fiftyone/integrations/coco.html) to visualize and facilitate access to [COCO](https://cocodataset.org/#download) dataset resources and [evaluation](https://cocodataset.org/#detection-eval)**. With FiftyOne, you can download specific subsets of COCO, visualize the data and labels, and evaluate your models on COCO more easily and in fewer lines of code than ever. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("coco-2017", split="validation") 5session = fo.launch_app(dataset) ``` ## Setup Using FiftyOne to access and work with the COCO dataset is as simple as [installing the open-source Python package](https://voxel51.com/docs/fiftyone/getting_started/install.html): ```python 1pip install fiftyone ``` ## Downloading COCO While existing tools or your own custom scripts have likely enabled you to download COCO splits in the past, have you ever wished you could download a small sample of a dataset for higher fidelity analysis before scaling up to the full dataset? The [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html) now supports partially downloading and loading COCO directly into Python in just one command. For example, say you are working on a road scene detection task and you are only interested in samples containing vehicles, people, and traffic lights. The following code snippet will download 100 of just the relevant samples and [load them into FiftyOne](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html#basic-recipe): ```python 1import fiftyone.zoo as foz 2 3dataset = foz.load_zoo_dataset( 4 "coco-2017", 5 split="validation", 6 label_types=["detections"], 7 classes=["person", "car", "truck", "traffic light"], 8 max_samples=100, 9) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/a81d57f507b61b6f7f90e40560b3b7cd2c6fce8a-1286x1005.png?auto=format&dpr=2&fit=max&q=75&w=1286) [The command](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#coco-2017) to load COCO takes the following arguments allowing you to customize exactly the samples and labels that you are interested in: - `label_types`: a list of types of labels to load. Values are `("detections", "segmentations")`. By default, all labels are loaded but not every sample will include each label type. If `max_samples` and `label_types` are both specified, then every sample will include the specified label types. - `split` and `splits`: either a string or list of strings dictating the splits to load. Available splits are `("test", "train", "validation")`. - `classes`: a list of strings specifying required classes to load. Only samples containing at least one instance of a specified class will be downloaded. - `max_samples`: a maximum number of samples to import. By default, all samples are imported. - `shuffle`: boolean dictating whether to randomly shuffle the order in which the samples are imported. - `seed`: a random seed to use when shuffling. - `image_ids`: a list of specific image IDs to load or a filepath to a file containing that list. The IDs can be specified either as `/` or `` ## Visualizing the dataset Most ML engineers have integrated basic support for viewing small batches of images into their workflows, eg levering tools like [Tensorboard](https://www.tensorflow.org/tensorboard), but such workflows do not allow for dynamically searching your dataset for specific examples of interest. This is an important gap to address because [the best way to understand the failure modes of your model is to investigate the quality of a dataset and annotations is to visualize them and scroll through some examples](https://towardsdatascience.com/i-performed-error-analysis-on-open-images-and-now-i-have-trust-issues-89080e03ba09). The [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html) allows everyone to visualize their datasets and labels without needing to spend time and money writing their own scripts or tools for visualization. When combined with the ease of use of the FiftyOne API, it now only takes a few lines of code to get hands-on with your data. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset( 5 "coco-2017", 6 split="validation", 7 label_types=["detections", "segmentations"], 8 max_samples=250, 9) 10session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/cc00a6adfa618214f9bdde4d12df92d7e636781d-1400x1112.png?auto=format&dpr=2&fit=max&q=75&w=1400) FiftyOne also provides a powerful query language allowing you to create [custom views into your dataset](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) to answer any question you have about your data. For example, you can create a view that filters predictions of a model by confidence and sorting so that samples with the most number of ground truth objects first. ```python 1from fiftyone import ViewField as F 2 3# Only contains detections with confidence >= 0.75 4view = dataset.filter_labels("faster_rcnn", F("confidence") > 0.75) 5 6# Chaine the view to sort by number of detections in `Detections` field `ground_truth` 7view = view.sort_by(F("ground_truth.detections").length(), reverse=True) ``` ## Evaluating models ![](https://cdn.sanity.io/images/h6toihm1/production/2044df6a5b4c423c6ca4174213d26aa559669c5c-967x800.png?auto=format&dpr=2&fit=max&q=75&w=967) The [COCO API](https://github.com/cocodataset/cocoapi) has been widely adopted as the standard metric for evaluating object detections. The COCO average precision is used to compare models in nearly every object detection research paper in the last half-decade. While it’s very useful to have a single metric that can be used to compare models at a high level, in practice, there is more work that needs to be done. When developing a model that will be put to use in a real-world scenario, you need to build confidence in the model performance in a variety of situations. A single metric will not give you that information, the only way to understand exactly where your model performs well and where it performs poorly is to [look at individual samples and even individual predictions to find success and failure cases](https://voxel51.com/docs/fiftyone/tutorials/evaluate_detections.html). FiftyOne provides extensive evaluation capabilities for various tasks, but most notably, it lets you take the next step after computing COCO AP and dig into the predictions of your models. You can [add your own model predictions to a FiftyOne dataset](https://voxel51.com/docs/fiftyone/recipes/adding_detections.html) with ease. The example below shows how you can use FiftyOne’s [`evaluate_detections()`](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#detections) method to evaluate the predictions of a model from the [FiftyOne Model Zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html). ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("coco-2017", split="validation") 5 6# Load model from zoo and apply it to dataset 7model = foz.load_zoo_model("faster-rcnn-resnet50-fpn-coco-torch") 8dataset.apply_model(model, label_field="predictions") 9 10# Evaluate `predictions` w.r.t. labels in `ground_truth` field 11results = dataset.evaluate_detections( 12 "predictions", gt_field="ground_truth", eval_key="eval", compute_mAP=True, 13) 14 15# Print the COCO AP 16print(results.mAP()) 17# ex: 0.358 18 19session = fo.launch_app(dataset) 20 21# Convert to evaluation patches 22eval_patches = dataset.to_evaluation_patches("eval") 23print(eval_patches) 24 25# View patches in the App 26session.view = eval_patches ``` Then you can use the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html) to analyze individual TP/FP/FN examples and cross-reference with additional attributes like whether the annotation is a crowd. ![](https://cdn.sanity.io/images/h6toihm1/production/33a3955d14567c6ef3ee853274a5e5294077d6d5-1277x955.gif?auto=format&dpr=2&fit=max&q=75&w=1277) Once you have evaluated your model in FiftyOne, you can use the returned `results` object to [view the AP, plot precision-recall curves](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#map-and-pr-curves), and [interact with confusion matrices](https://voxel51.com/docs/fiftyone/user_guide/plots.html#confusion-matrices) to quickly find the exact samples where your model is correct and incorrect for each class. ```python 1# Plot confusion matrix 2classes = ["person", "kite", "car", "bird", "carrot", "boat", "surfboard", "airplane", "traffic_light", "chair"] 3plot = results.plot_confusion_matrix(classes=classes) 4plot.show() 5 6# Plot precision-recall curve 7plot = results.plot_pr_curve(classes=classes) 8plot.show() 9 10# (In Jupyter notebooks) 11# Connect to session 12session.plots.attach(plot) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/4ca1de061c1062151ade7666013dd5536cd6171a-1870x856.gif?auto=format&dpr=2&fit=max&q=75&w=1600) _Note: [Interactive plots](https://voxel51.com/docs/fiftyone/user_guide/plots.html) are currently only available in Jupyter notebooks but other contexts will be supported soon!_ ## Summary [COCO](https://cocodataset.org/#home) is one of the most popular and influential computer vision datasets. Now, [FiftyOne’s](https://fiftyone.ai/) native support for COCO makes it easier than ever to download specific parts of the dataset, visualize COCO and your model predictions, and evaluate your models with COCO-style evaluation. [COCO](https://voxel51.com/blog/tag/coco) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [image dataset](https://voxel51.com/blog/tag/image-dataset) [images](https://voxel51.com/blog/tag/images) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/c4a04d43252eacdfaa8acbfafe00210d4e05f92c-1090x754.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Kinetics Dataset: Train and Evaluate Video Classification Models\\ \\ Datasets\\ \\ • \\ \\ Apr 13, 2022](https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models) [![](https://cdn.sanity.io/images/h6toihm1/production/dc8a2e7a894316856af5a109ae43f8787959f179-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Embeddings Tips and Tricks – Mar 31, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 31, 2023](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0da316f9c985262ed5121d9b2e7e7a080bb5c599-1231x920.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Loading Open Images V6 and Custom Datasets with FiftyOne\\ \\ Datasets\\ \\ • \\ \\ Feb 11, 2021](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-246-lllmstxt|> ## Loading Open Images [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) Loading Open Images V6 and Custom Datasets with FiftyOne Feb 11, 2021 • 11 min read Article content In this article [Open Images V6](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#615bcd22b351) [A New Way to Download and Evaluate Open Images!](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#8af9cde8e399) [Open Images Label Formats](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#7dcb619a8216) [Downloading Data Locally](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#29d60ae64b0b) [Image-Level Labels](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#7a446e65d19e) [Detections](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#64634dbd2106) [Visual Relationships](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#419e35b970d1) [Instance Segmentation](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#4b5413ef7730) [Preprocessing](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#d442ec47d5aa) [Loading Custom Datasets into FiftyOne](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#a1da40c00bf2) [Classification Labels](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#136ff2fe7be8) [Object Detections](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#8943b7142d3c) [Visual Relationships](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#910fcc0d3c80) [Segmentations](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#748b3dbfd09b) [Creating FiftyOne Samples](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#817b9e61ee99) [Visualizing and Exploring](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#3023a30e7284) [Queries](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#8da8436b6d6e) [Converting Formats](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#59ef0e33b30e) [The General Formula for Loading Datasets](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#5eca58303056) In this article [Open Images V6](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#615bcd22b351) [A New Way to Download and Evaluate Open Images!](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#8af9cde8e399) [Open Images Label Formats](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#7dcb619a8216) [Downloading Data Locally](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#29d60ae64b0b) [Image-Level Labels](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#7a446e65d19e) [Detections](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#64634dbd2106) [Visual Relationships](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#419e35b970d1) [Instance Segmentation](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#4b5413ef7730) [Preprocessing](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#d442ec47d5aa) [Loading Custom Datasets into FiftyOne](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#a1da40c00bf2) [Classification Labels](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#136ff2fe7be8) [Object Detections](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#8943b7142d3c) [Visual Relationships](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#910fcc0d3c80) [Segmentations](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#748b3dbfd09b) [Creating FiftyOne Samples](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#817b9e61ee99) [Visualizing and Exploring](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#3023a30e7284) [Queries](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#8da8436b6d6e) [Converting Formats](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#59ef0e33b30e) [The General Formula for Loading Datasets](https://voxel51.com/blog/loading-open-images-v6-and-custom-datasets-with-fiftyone#5eca58303056) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Datasets and their annotations are often stored in very different formats. FiftyOne allows for easy loading and visualization of any image dataset and labels. ![](https://cdn.sanity.io/images/h6toihm1/production/f3cb06ba0584774c2a4010232e082c3aacfc5ef6-830x641.png?auto=format&dpr=2&fit=max&q=75&w=830) [DataFrames](https://pandas.pydata.org/pandas-docs/stable/user_guide/dsintro.html#dataframe) are a standard way of storing tabular data with various tools that exist to visualize the data in different ways. Image and video datasets, on the other hand, do not have a standard format for storing their data and annotations. Nearly every dataset that is developed creates a new schema with which to store their raw data, bounding boxes, sample-level labels, etc. I have been working on an open-source machine learning tool called FiftyOne that can help ease the pain of having to write custom loading, visualization, and conversion scripts whenever you use a new dataset. FiftyOne supports [multiple dataset formats out of the box including MS-COCO, YOLO, Pascal VOC, and more](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html#loading-datasets). However, if you have a dataset format not provided out-of-the-box, you can still easily load it into FiftyOne manually. **_Why would you want your data in [FiftyOne](https://fiftyone.ai/)?_** [FiftyOne](https://fiftyone.ai/) provides a highly functional App and API that will let you quickly visualize your dataset, generate interesting queries, find annotation mistakes, convert it to other formats, load it into a zoo of models, and more. This blog post will walk you through how to load image-level classifications, object detections, segmentations, and visual relationships into [FiftyOne](https://fiftyone.ai/), visualize them, and convert them to other formats. I’ll be using [Open Images V6](https://storage.googleapis.com/openimages/web/download.html) which was released in February 2020 as a basis for this post since it contains all of these data types. If you are only interested in loading Open Images V6, you can check it out in the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#dataset-zoo-open-images-v6) and [load it in one line of code](https://voxel51.com/docs/fiftyone/tutorials/open_images.html)! If you have your own dataset that you want to load, adjust the code in this post to parse the format that your data is stored in. ## Open Images V6 [Open Images is a dataset](https://storage.googleapis.com/openimages/web/download.html) released by Google containing over 9M images with labels spanning various tasks: - Image-level labels\* - Object bounding boxes\* - Visual relationships\* - Instance segmentation masks\* - Localized narratives _\*Loaded in this post_ These annotations were generated through a combination of machine learning algorithms followed by human verification on the test, validation, and subsets of the training splits. Versions of this dataset are also used in the [Open Images Challenges on Kaggle](https://www.kaggle.com/c/open-images-2019-object-detection). [Open Images V6](https://storage.googleapis.com/openimages/web/download.html) introduced [localized narratives](https://ai.googleblog.com/2020/02/open-images-v6-now-featuring-localized.html), which are a novel form of multimodal annotations consisting of a voiceover and mouse trace of an annotator describing an image. FiftyOne support for localized narratives is currently in the works. ## A New Way to Download and Evaluate Open Images! \[Updated May 12, 2021\] After releasing this post, we [collaborated with Google](https://storage.googleapis.com/openimages/web/2021-05-12-oid-and-fiftyone.html) to support [Open Images V6](https://storage.googleapis.com/openimages/web/download.html) directly through the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#dataset-zoo-open-images-v6). It is now as easy as this to load Open Images, data, annotations, and all: ```python 1import fiftyone.zoo as foz 2 3oi_dataset = foz.load_zoo_dataset("open-images-v6", split="validation") ``` With this implementation in FiftyOne, you can also specify any subset of Open Images with parameters like `classes`, `split`, `max_samples`, and more: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset( 5 "open-images-v6", 6 "validation", 7 label_types=["detections", "classifications"], 8 classes = ["Dog", "Cat"], 9 max_samples=1000, 10 seed=51, 11 shuffle=True, 12 dataset_name="open-images-dog-cat", 13) 14 15session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/b5098a49eec8855e610bd3c84c758ede2f65396e-1220x913.png?auto=format&dpr=2&fit=max&q=75&w=1220) Additionally, if you are training a model on Open Images, FiftyOne now supports [Open Images style evaluation](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#evaluating-detections-open-images) allowing you to produce the same [mAP metrics used in the Open Images challenges](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html#open-images-challenge). The benefit of using FiftyOne for this is that it also stores instance-level true positive, false positive, and false negative results allowing you to not rely only on aggregate dataset-wide metrics but actually get hands-on with your model results and find out how to best improve performance. ```python 1results = fo.evaluate_detections( 2 dataset, 3 pred_field="predictions", 4 gt_field="detections", 5 pos_label_field="positive_labels", 6 neg_label_field="negative_labels", 7 hierarchy = dataset.info["hierarchy"], 8 method = "open-images", 9 expand_pred_hierarchy=True, 10) ``` **_>\> For more information check out [this post](https://towardsdatascience.com/googles-open-images-now-easier-to-download-and-evaluate-with-fiftyone-615ce0482c02) or [this tutorial](https://voxel51.com/docs/fiftyone/tutorials/open_images.html)!_** ## Open Images Label Formats The previous section shows the best way to load the Open Images dataset. However, FiftyOne also lets you easily load custom datasets. The next few sections show how to load a dataset into FiftyOne from scratch. We are using Open Images as the example dataset for this since it contains a rich variety of label types. **_Note: The code in the following sections is meant to be adapted to your own datasets, it does not need to be used to [load Open Images](https://voxel51.com/docs/fiftyone/tutorials/open_images.html). Use [the examples above](https://towardsdatascience.com/googles-open-images-now-easier-to-download-and-evaluate-with-fiftyone-615ce0482c02) if you are only interested in [loading the Open Images dataset](https://voxel51.com/docs/fiftyone/tutorials/open_images.html)._** In this “Open Images Label Formats” section, we describe the format used by Google to store Open Images annotations on disk. We will use this information to write the parsers to load this dataset into FiftyOne in the next “Loading custom datasets into FiftyOne” section. ## Downloading Data Locally The AWS download links for the training split (513 GB), validation split (12 GB), and testing split (36 GB) can be [found at Open Images GitHub repository](https://github.com/cvdfoundation/open-images-dataset#download-full-dataset-with-google-storage-transfer). Annotations for the tasks that you are interested in can be [downloaded directly from the Open Images website](https://storage.googleapis.com/openimages/web/download.html). We will be using samples from the test split for this example. You can download the entire test split (36 Gb!) with the following commands: ```python 1pip install awscli 2aws s3 --no-sign-request sync s3://open-images-dataset/test ./open-images/test/ ``` **Alternatively, I will be downloading just a few images from the test split further down in this post.** We will also need to download the relevant annotation files for each task that are all found here: [https://storage.googleapis.com/openimages/web/download.html](https://storage.googleapis.com/openimages/web/download.html) ## Image-Level Labels Every image in Open Images can contain multiple image-level labels across hundreds of classes. These labels are split into two types, positive and negative. Positive labels are classes that have been verified to be in the image while negative labels are classes that are verified to not be in the image. Negative labels are useful because they are generally specified for classes that you may expect to appear in a scene but do not. For example, if there is a group of people in outfits on a field, you may expect there to be a `ball` . If there isn’t one, that would be a good negative label. ```python 1wget -P labels https://storage.googleapis.com/openimages/v5/test-annotations-human-imagelabels-boxable.csv ``` Below is a sample of the contents of this file: ImageID,Source,LabelName,Confidence 000026e7ee790996,verification,/m/0cgh4,0 000026e7ee790996,verification,/m/04hgtk,0 ... We need the class list for both labels and detections: ```python 1wget -P labels https://storage.googleapis.com/openimages/v5/class-descriptions-boxable.csv ``` Below is a sample of the contents of this file: /m/011k07,Tortoise /m/011q46kg,Container ... ![](https://cdn.sanity.io/images/h6toihm1/production/6869ae54d1a85f1d1a211b1aecf24005145821a4-1384x650.png?auto=format&dpr=2&fit=max&q=75&w=1384) ## Detections Objects are localized and labeled with the same classes as the image-level labels. Additionally, each detection contains boolean attributes indicating if the object is occluded, truncated, representing a group of other objects, inside another object, or a depiction of the object (like a cartoon). ```python 1wget -P detections https://storage.googleapis.com/openimages/v5/test-annotations-bbox.csv ``` Below is a sample of the contents of this file: ImageID,Source,LabelName,Confidence,XMin,XMax,YMin,YMax,IsOccluded,IsTruncated,IsGroupOf,IsDepiction,IsInside 000026e7ee790996,xclick,/m/07j7r,1,0.071875,0.1453125,0.20625,0.39166668,0,1,1,0,0 000026e7ee790996,xclick,/m/07j7r,1,0.4390625,0.571875,0.26458332,0.43541667,0,1,1,0,0 ... ![](https://cdn.sanity.io/images/h6toihm1/production/53ed27c9fb24505df0ddf7743a7f70c705897502-1400x786.png?auto=format&dpr=2&fit=max&q=75&w=1400) ## Visual Relationships Relationships are labeled between two object detections. Examples are if one object is wearing another. The most common relationship is `is`, indicating if an object `is` some attribute (like if a handbag `is` leather). The annotations for these relationships include the bounding boxes and labels of both objects as well as the label for the relationship. ```python 1wget -P relationships https://storage.googleapis.com/openimages/v6/oidv6-test-annotations-vrd.csv 2wget -P relationships https://storage.googleapis.com/openimages/v6/oidv6-attributes-description.csv ``` Below is a sample of the contents of the relationships file: ImageID,LabelName1,LabelName2,XMin1,XMax1,YMin1,YMax1,XMin2,XMax2,YMin2,YMax2,RelationshipLabel 9553b9608577b74b,/m/04yx4,/m/017ftj,0.023404,0.985106,0.038344,0.981595,0.238298,0.759574,0.349693,0.529141,wears 819903f6353b60b5,/m/03m3pdh,/m/0dnr7,0.096875,1.000000,0.095833,1.000000,0.096875,1.000000,0.095833,1.000000,is ... ![](https://cdn.sanity.io/images/h6toihm1/production/f40462cb2b21707d0728d7a52307fd9b0d99cc4a-1400x917.png?auto=format&dpr=2&fit=max&q=75&w=1400) ## Instance Segmentation Segmentation masks are downloaded through 16 zip files each containing the masks related to images starting with `0–9` or `A-F`. In this example, we will only be using images starting with `0`. The following command downloads just those masks, replace the `0` with `1-9` or `a-f` to download masks for other images. These segmentation annotations are stored in a separate image for each object and also include the bounding box coordinates around the segmentation and the label of the segmentation. ```python 1wget -P segmentations https://storage.googleapis.com/openimages/v5/test-masks/test-masks-0.zip 2unzip -d segmentations/masks segmentations/test-masks-0.zip 3wget -P segmentations https://storage.googleapis.com/openimages/v5/test-annotations-object-segmentation.csv ``` Below is a sample of the contents of the segmentations file: MaskPath,ImageID,LabelName,BoxID,BoxXMin,BoxXMax,BoxYMin,BoxYMax,PredictedIoU,Clicks d0ed76e0533a914d\_m01xyhv\_cffd8afa.png,d0ed76e0533a914d,/m/01xyhv,cffd8afa,0.122966,0.958409,0.389892,0.998195,0.00000, fba940ee0203b368\_m08pbxl\_13094a08.png,fba940ee0203b368,/m/08pbxl,13094a08,0.460938,0.551562,0.387500,0.572917,0.00000, ... ![](https://cdn.sanity.io/images/h6toihm1/production/623f2c582a44d23ae07b1e5632eb928f418ce0b1-1400x929.png?auto=format&dpr=2&fit=max&q=75&w=1400) ## Preprocessing We are only going to use a small subset of the dataset in this example to make it easy to follow along with. Additionally, since we want to load a lot of different types of annotations, we need to find some samples that are compatible with all of our labels. Let's load in the annotations from the `csv` files we downloaded and parse them to find a subset of images we want to use. ```python 1import csv 2 3with open("detections/test-annotations-bbox.csv") as f: 4 reader = csv.reader(f, delimiter=',') 5 det_data = [row for row in reader] 6 7with open("labels/test-annotations-human-imagelabels-boxable.csv") as f: 8 reader = csv.reader(f, delimiter=',') 9 lab_data = [row for row in reader] 10 11with open("relationships/oidv6-test-annotations-vrd.csv") as f: 12 reader = csv.reader(f, delimiter=',') 13 rel_data = [row for row in reader] 14 15with open("segmentations/test-annotations-object-segmentation.csv") as f: 16 reader = csv.reader(f, delimiter=',') 17 seg_data = [row for row in reader] 18 19# Find intersection of ImageIDs with all annotations 20det_ids = {l[0] for l in det_data[1:]} 21lab_ids = {l[0] for l in lab_data[1:]} 22rel_ids = {l[0] for l in rel_data[1:]} 23 24# We only downloaded the zip for files starting with "0" 25seg_ids = {l[1] for l in seg_data[1:] if l[1][0] == "0"} 26 27valid_ids = det_ids & lab_ids & rel_ids & seg_ids ``` We now have a list of `valid_ids` that contains all of the annotations we want to look at. Let's choose a subset of 100 of those and download the corresponding images following what is done in the [official Open Images download script](https://raw.githubusercontent.com/openimages/dataset/master/downloader.py). ```python 1pip install boto3 ``` ```python 1import boto3 2import botocore 3import os 4 5BUCKET_NAME = 'open-images-dataset' 6 7bucket = boto3.resource( 8 's3', config=botocore.config.Config( 9 signature_version=botocore.UNSIGNED)).Bucket(BUCKET_NAME) 10 11num_ids = 100 12 13os.makedirs('images/test', exist_ok=True) 14 15for image_id in list(valid_ids)[:num_ids]: 16 bucket.download_file(f'test/{image_id}.jpg', 17 os.path.join('images/test/', f'{image_id}.jpg')) ``` The last thing we need is a mapping from the class and attribute IDs to their actual names. ```python 1import csv 2 3with open("labels/class-descriptions-boxable.csv") as f: 4 reader = csv.reader(f, delimiter=',') 5 cls_data = [row for row in reader] 6 7# Map of class IDs to class names 8classes_map = {k: v for k,v in cls_data} 9 10 11with open("relationships/oidv6-attributes-description.csv") as f: 12 reader = csv.reader(f, delimiter=',') 13 attrs_data = [row for row in reader] 14 15# Map of attribute IDs to attribute names 16attrs_map = {k: v for k,v in attrs_data} ``` ## Loading Custom Datasets into FiftyOne You will first need to install [FiftyOne](https://fiftyone.ai/) through a simple pip command. It is recommended to work with [FiftyOne](https://fiftyone.ai/) in interactive Python sessions, so let's install that too. ```python 1pip install fiftyone 2pip install ipython ``` After launching `ipython` the first step is to create a [FiftyOne Dataset](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html). ```python 1import fiftyone as fo 2 3dataset = fo.Dataset("open_images_v6") ``` If you want this dataset to exist after exiting the Python session, set the `persistent` attribute to `True`. This lets us quickly load the dataset in the future. ```python 1dataset.persistent = True ``` We then need to create [FiftyOne Samples](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#samples) for each image that contain the file path to the images as well as all label information that we want to import. For each label type, we will create a corresponding object in FiftyOne and add it as a `field` to our samples. Adding image-level [classification labels](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#labels) will utilize the `fo.Classifications` class. [Detections, segmentations, and relations](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#labels) can all use the `fo.Detections` class since it supports bounding boxes, masks, and also custom attributes assigned to each detection. These custom attributes can be used for things like `IsOccluded` in the detections or the two labels that a relationship is between. The sections below outline how to create FiftyOne labels from the Open Images data we have loaded so far and then how to add them to your FiftyOne Dataset. ## Classification Labels Classification labels utilize the `fo.Classification` class. Since these are multi-label classifications, we will be using the `fo.Classifications` class to store multiple classification labels. Additionally, we want to separate out the positive and negative labels (1 and 0 confidence respectively) into different classifications fields so we can view them separately in the App. ```python 1def create_labels(lab_data, image_id): 2 pos_cls = [] 3 neg_cls = [] 4 # Get relevant data for this image 5 sample_labs = [i for i in lab_data if i[0]==image_id] 6 for sample_lab in sample_labs: 7 # sample_lab reference: [ImageID,Source,LabelName,Confidence] 8 label = classes_map[sample_lab[2]] 9 conf = float(sample_lab[3]) 10 cls = fo.Classification(label=label, confidence=conf) 11 12 if conf > 0.1: 13 pos_cls.append(cls) 14 else: 15 neg_cls.append(cls) 16 17 pos_labels = fo.Classifications(classifications=pos_cls) 18 neg_labels = fo.Classifications(classifications=neg_cls) 19 20 return pos_labels, neg_labels ``` ## Object Detections Similar to classifications, the `fo.Detections` class lets you store multiple `fo.Detection` objects in a list. We create a detection by defining the bounding box coordinates and class label of the object. We can then add any additional attributes that we want, like `IsOccluded` and `IsTruncated`. ```python 1def create_detections(det_data, image_id): 2 dets = [] 3 4 sample_dets = [i for i in det_data if i[0]==image_id] 5 for sample_det in sample_dets: 6 # sample_det reference: [ImageID,Source,LabelName,Confidence,XMin,XMax,YMin,YMax,IsOccluded,IsTruncated,IsGroupOf,IsDepiction,IsInside] 7 label = classes_map[sample_det[2]] 8 xmin = float(sample_det[4]) 9 xmax = float(sample_det[5]) 10 ymin = float(sample_det[6]) 11 ymax = float(sample_det[7]) 12 13 # Convert to [top-left-x, top-left-y, width, height] 14 bbox = [xmin, ymin, xmax-xmin, ymax-ymin] 15 16 detection = fo.Detection(bounding_box=bbox, label=label) 17 18 detection["IsOccluded"] = bool(int(sample_det[8])) 19 detection["IsTruncated"] = bool(int(sample_det[9])) 20 detection["IsGroupOf"] = bool(int(sample_det[10])) 21 detection["IsDepiction"] = bool(int(sample_det[11])) 22 detection["IsInside"] = bool(int(sample_det[12])) 23 24 dets.append(detection) 25 26 detections = fo.Detections(detections=dets) 27 28 return detections ``` ## Visual Relationships Relationships are best represented in FiftyOne through `fo.Detections` since a relationship contains a bounding box, relationship label, and object labels, all of which can be stored in a detection. We are going to have the bounding box of the relationship encompass the bounding boxes of both objects it pertains to. We add the labels of each object as additional custom fields to the detection. It should be noted, that you could easily also add the two objects that make up the relationship as individual detections, not doing so was just a design choice for this post. ```python 1def create_relationships(rel_data, image_id): 2 rels = [] 3 sample_rels = [i for i in rel_data if i[0]==image_id] 4 for sample_rel in sample_rels: 5 # sample_rel reference: [ImageID,LabelName1,LabelName2,XMin1,XMax1,YMin1,YMax1,XMin2,XMax2,YMin2,YMax2,RelationshipLabel] 6 label1 = classes_map[sample_rel[1]] 7 8 attribute = False 9 if sample_rel[2] in classes_map: 10 label2 = classes_map[sample_rel[2]] 11 else: 12 label2 = attrs_map[sample_rel[2]] 13 attribute = True 14 15 label_rel = sample_rel[-1] 16 17 xmin1 = float(sample_rel[3]) 18 xmax1 = float(sample_rel[4]) 19 ymin1 = float(sample_rel[5]) 20 ymax1 = float(sample_rel[6]) 21 22 xmin2 = float(sample_rel[7]) 23 xmax2 = float(sample_rel[8]) 24 ymin2 = float(sample_rel[9]) 25 ymax2 = float(sample_rel[10]) 26 27 xmin_int = min(xmin1, xmin2) 28 ymin_int = min(ymin1, ymin2) 29 xmax_int = max(xmax1, xmax2) 30 ymax_int = max(ymax1, ymax2) 31 32 # Convert to [top-left-x, top-left-y, width, height] 33 bbox_int = [xmin_int, ymin_int, xmax_int-xmin_int, ymax_int-ymin_int] 34 35 detection_rel = fo.Detection(bounding_box=bbox_int, label=label_rel) 36 37 detection_rel["Label1"] = label1 38 detection_rel["Label2"] = label2 39 40 rels.append(detection_rel) 41 42 relationships = fo.Detections(detections=rels) 43 44 return relationships ``` ## Segmentations We can once again use `fo.Detections` to store segmentations since a detection contains an optional `mask` argument that accepts a NumPy array and will scale it to the bounding box region. The segmentations in Open Images also contain a bounding box around the mask as well as the instance label, all of which is added to the detection objects. ```python 1def create_segmentations(seg_data, image_id): 2 segs = [] 3 sample_segs = [i for i in seg_data if i[1]==image_id] 4 for sample_seg in sample_segs: 5 # sample_seg reference: [MaskPath,ImageID,LabelName,BoxID,BoxXMin,BoxXMax,BoxYMin,BoxYMax,PredictedIoU,Clicks] 6 label = classes_map[sample_seg[2]] 7 xmin = float(sample_seg[4]) 8 xmax = float(sample_seg[5]) 9 ymin = float(sample_seg[6]) 10 ymax = float(sample_seg[7]) 11 12 # Convert to [top-left-x, top-left-y, width, height] 13 bbox = [xmin, ymin, xmax-xmin, ymax-ymin] 14 15 # Load boolean mask 16 mask_path = os.path.join("segmentations/masks", sample_seg[0]) 17 mask = cv2.imread(mask_path, cv2.IMREAD_GRAYSCALE) > 122 18 h,w = mask.shape 19 cropped_mask = mask[int(ymin*h):int(ymax*h), int(xmin*w):int(xmax*w)] 20 21 segmentation = fo.Detection(bounding_box=bbox, label=label, 22 mask=cropped_mask) 23 24 segs.append(segmentation) 25 26 segmentations = fo.Detections(detections=segs) 27 28 return segmentations ``` ## Creating FiftyOne Samples Now that we defined the functions to take in Open Images data and return FiftyOne labels, we can create samples and add these labels to them. Samples only need a `filepath` to be instantiated and we can add any FiftyOne labels to a sample. Once the sample is created, we can add it to the dataset and continue until all of our data is loaded. ```python 1# Add Samples to Dataset 2for image_id in list(valid_ids)[:num_ids]: 3 sample = fo.Sample(filepath=os.path.join("images/test/%s.jpg" % image_id)) 4 5 # Add Labels 6 pos_labels, neg_labels = create_labels(lab_data, image_id) 7 sample["positive_labels"] = pos_labels 8 sample["negative_labels"] = neg_labels 9 10 # Add Detections 11 detections = create_detections(det_data, image_id) 12 sample["detections"] = detections 13 14 # Add Segmentations 15 segmentations = create_segmentations(seg_data, image_id) 16 sample["segmentations"] = segmentations 17 18 # Add Relationships 19 relationships = create_relationships(rel_data, image_id) 20 sample["relationships"] = relationships 21 22 dataset.add_sample(sample) ``` ## Visualizing and Exploring Once we have our data loaded into a [FiftyOne](https://fiftyone.ai/) dataset, we can launch the [App](https://voxel51.com/docs/fiftyone/user_guide/app.html) and start exploring. ```python 1session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/73d029a332b12d17e56cf8c3cb052997cfc1e60f-1400x728.png?auto=format&dpr=2&fit=max&q=75&w=1400) In the [App](https://voxel51.com/docs/fiftyone/user_guide/app.html), we can select which of the label fields that we want to view, look at individual samples in an expanded view, and also view the distributions of labels. ![](https://cdn.sanity.io/images/h6toihm1/production/7eeec5fe172936fe542fb9388e034fa4cd5ced61-1232x542.png?auto=format&dpr=2&fit=max&q=75&w=1232) Opening a sample in the expanded view lets you visualize the attributes we added, like the labels of a relationship. For example, we can see that there are two `Ride` relationships between `Man` and `Horse` in the image below. ![](https://cdn.sanity.io/images/h6toihm1/production/0da316f9c985262ed5121d9b2e7e7a080bb5c599-1231x920.png?auto=format&dpr=2&fit=max&q=75&w=1231) Being able to visualize our dataset easily lets us quickly spot check the data. For example, it appears that the `Mammal` label has an inconsistent meaning between different samples. Below are two images containing humans, one has `Mammal` as a `negative_label` and the other has `Mammal` as a `positive_label`. ![](https://cdn.sanity.io/images/h6toihm1/production/1a3e42f032299cf8bc79f3cb9e9446b2174c07dc-1716x432.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Queries One of the cutting-edge features that FiftyOne provides is the ability to interact closely with your dataset [in code](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) and [in the App](https://voxel51.com/docs/fiftyone/user_guide/app.html#using-the-view-bar). This lets you write sophisticated queries that would otherwise require a large amount of scripting. For example, say that we want build a subset of Open Images containing close up images of faces. We can [create a view into the dataset](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) that will let us get all detections that contain a `Human face` with a bounding box area greater than 0.2. ```python 1from fiftyone import ViewField as F 2 3bbox_area = F("bounding_box")[2] * F("bounding_box")[3] 4 5large_face_view = ( 6 dataset 7 .filter_labels( 8 "detections", F("label")=="Human face", only_matches=True 9 ).filter_labels( 10 "detections", bbox_area > 0.2, only_matches=True 11 ) 12) 13 14session.view = large_face_view ``` ![](https://cdn.sanity.io/images/h6toihm1/production/90d20237d5c6351f859962774d15b053c41921a8-1225x735.png?auto=format&dpr=2&fit=max&q=75&w=1225) ## Converting Formats Once your data is in [FiftyOne](https://fiftyone.ai/), you can [export it in any of the formats that FiftyOne supports with just a couple of lines of code](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html). For example, if we want to export the `detections` that we added to [MS-COCO](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/datasets.html#cocodetectiondataset-import) format so that we can use the [`pycocotools`](https://github.com/cocodataset/cocoapi/tree/master/PythonAPI/pycocotools) evaluation on it, we can do so in one line of code in Python. ```python 1dataset.export( 2 export_dir="/path/to/dir", 3 dataset_type=fo.types.COCODetectionDataset, 4 label_field="detections", 5) ``` ## The General Formula for Loading Datasets The easiest way to load your data is if you follow a [standard format for your annotations](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html#loading-datasets). For example, if you just finished annotating in the open-source tool, [CVAT](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/datasets.html#cvatimagedataset), you can load it into FiftyOne as easily as: ```python 1import fiftyone as fo 2 3dataset = fo.Dataset.from_dir("path/to/data", dataset_type=fo.types.CVATImageDataset) ``` Even if your data is in a custom format, it’s easy to [manually build a FiftyOne dataset](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/samples.html#manually-building-datasets). ```python 1import fiftyone as fo 2 3# Initialize your dataset 4dataset = fo.Dataset(name) 5 6# Get a list of image paths 7images = .... 8 9# Parse your labels into samples 10for path in images: 11 sample = fo.Sample(path) 12 sample["label"] = fo.Classification(label=...) 13 14 detections = [] 15 for det in parsed_detection_labels: 16 detection = fo.Detection(label=..., bounding_box=[...]) 17 detections.append(detection) 18 19 sample["detections"] = fo.Detections(detections=detections) 20 21 dataset.add_sample(sample) 22 23# Visualize your dataset in the FiftyOne App 24session = fo.launch_app(dataset) ``` Once you’ve loaded your data, you can utilize the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html) to visualize and explore your dataset, use the [FiftyOne Model Zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html) to generate predictions on your data, [export it in various formats](https://voxel51.com/docs/fiftyone/user_guide/export_datasets.html), and [more](https://voxel51.com/docs/fiftyone/index.html)! [custom datasets](https://voxel51.com/blog/tag/custom-datasets) [image dataset](https://voxel51.com/blog/tag/image-dataset) [Open Images](https://voxel51.com/blog/tag/open-images) [Open Images v6](https://voxel51.com/blog/tag/open-images-v6) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/cc00a6adfa618214f9bdde4d12df92d7e636781d-1400x1112.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The COCO Dataset: Best Practices for Downloading, Visualization, and Evaluation\\ \\ Datasets\\ \\ • \\ \\ Jun 30, 2021](https://voxel51.com/blog/the-coco-dataset-best-practices-for-downloading-visualization-and-evaluation) [![](https://cdn.sanity.io/images/h6toihm1/production/041e567b76170fe6b3bca2b7a4ffa5bd139ee452-1244x700.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Exploring Google’s Open Images V7\\ \\ Datasets\\ \\ • \\ \\ Mar 8, 2023](https://voxel51.com/blog/exploring-google-open-images-v7) [![](https://cdn.sanity.io/images/h6toihm1/production/40381f5f37fa5fcd70eddca2f63b6710568f5d2c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual Kinship Recognition with the Families in the Wild Computer Vision Dataset\\ \\ Datasets\\ \\ • \\ \\ Dec 7, 2022](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-247-lllmstxt|> ## FiftyOne Post-Launch Update [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) FiftyOne, Six Months Post-Launch Feb 4, 2021 • 11 min read Article content In this article [Setup and Installation: Easy as Pie](https://voxel51.com/blog/fiftyone-six-months-post-launch#1d45b1f18b7b) [Analysis: Working with new Datasets](https://voxel51.com/blog/fiftyone-six-months-post-launch#756de411d566) [Access: Going to the Zoo](https://voxel51.com/blog/fiftyone-six-months-post-launch#755b48054ef5) [Video: FiftyOne now supports video datasets](https://voxel51.com/blog/fiftyone-six-months-post-launch#4c5c8a501cfb) [Sharing: Notebooks and Colab](https://voxel51.com/blog/fiftyone-six-months-post-launch#c91d81cc005b) [Wrapping Up](https://voxel51.com/blog/fiftyone-six-months-post-launch#03724b138965) In this article [Setup and Installation: Easy as Pie](https://voxel51.com/blog/fiftyone-six-months-post-launch#1d45b1f18b7b) [Analysis: Working with new Datasets](https://voxel51.com/blog/fiftyone-six-months-post-launch#756de411d566) [Access: Going to the Zoo](https://voxel51.com/blog/fiftyone-six-months-post-launch#755b48054ef5) [Video: FiftyOne now supports video datasets](https://voxel51.com/blog/fiftyone-six-months-post-launch#4c5c8a501cfb) [Sharing: Notebooks and Colab](https://voxel51.com/blog/fiftyone-six-months-post-launch#c91d81cc005b) [Wrapping Up](https://voxel51.com/blog/fiftyone-six-months-post-launch#03724b138965) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/905d611bc3c0dcd856f63a2d0af25244b909aab8-800x650.gif?auto=format&dpr=2&fit=max&q=75&w=800) [FiftyOne](https://fiftyone.ai/) [launched August 2020](https://voxel51.com/press/fiftyone-open-source-launch/), about six months ago. Although the launch was exciting for our team, as it was the culmination of many months of work, the last six months have been even more exciting. With its unique emphasis on rapid dataset analysis for better models, faster, [FiftyOne](https://fiftyone.ai/) is quickly become an everyday tool for computer vision and machine learning engineers. Even in the face of growing competition in tools like Aquarium and Nucleus, [FiftyOne’s](https://fiftyone.ai/) flexibility — it is the sole open-source offering and brings the tool to you and your data rather than forcing you to bring your data into some specific cloud-service — and capabilities — it has several unique and innovative features around dataset quality and model performance analysis — have brought it to the [forefront of the machine learning developer tools conversation](https://huyenchip.com/2020/12/30/mlops-v2.html). So, yes, definitely exciting! But, I’ve personally been lost in the dust. While [FiftyOne’s](https://fiftyone.ai/) userbase grew from a few dozen private-beta users to hundreds of returning active users, the team has been adding a rich amount of new functionality. Although we collectively plan and work through our roadmap, as the CEO and a Professor, I was more outward facing: giving talks about [open-source AI](https://tmt.knect365.com/ai-summit-san-francisco/speakers/jason-corso/) and [my views on the future of AI](https://www.youtube.com/watch?v=DzCo9DggLyY), talking with users and customers, maintaining relationships with investors, and so on. I hence thought it would be good to take a moment and review these new developments with [FiftyOne](https://fiftyone.ai/) and how it’s progressed over the last six months. Here is what I found, with an emphasis on the core features for dataset analysis, as I intentionally neglect [the Brain capabilities](https://voxel51.com/docs/fiftyone/user_guide/brain.html) for tasks like [finding label mistakes](https://voxel51.com/docs/fiftyone/tutorials/label_mistakes.html) and [unique images](https://voxel51.com/docs/fiftyone/tutorials/uniqueness.html) in a dataset, for now. - Setup and Installation: getting started with [FiftyOne](https://fiftyone.ai/) now takes less than a minute! - Analysis: working with new datasets has a rich set of new user interface features for rapidly building an intuition for a dataset through interactive querying, searching and sorting. - Access: an integrated [dataset zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html) and an innovative [model zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html) make trying out new ideas so easy. - Video: [FiftyOne](https://fiftyone.ai/) now support native video datasets and associated tasks. - Sharing: Jupyter notebooks and Colab support significantly enhance FiftyOne’s utility for computer vision and machine learning in practice. ## Setup and Installation: Easy as Pie **What’s New?** Simple `pip`-based install from the standard pypi server with no additional manual work required. When I was a core code contributor to [FiftyOne](https://fiftyone.ai/) last year, it took some serious work to get it up and running on my machine: multiple packages, various setup and configuration scripts, and the download of some packages took a good amount of time. **Today,** **here’s what I did to get it up and running.** This is from scratch, with python 3.7.5. already installed. ```python 1python3 -m venv foe 2source foe/bin/activate 3pip install -U pip setuptools wheel 4pip install ipython 5pip install fiftyone ``` This whole installation took 1 minute (over my very slow cell-phone based internet connection). The installation of `IPython` is optional but it’s my basic workflow so I kept it here for completeness; and, the upgrade of `pip`, `setuptools`, and `wheel` are also optional but I’ve learned to keep them as a best practice so that as few things as possible are built from source (instead of 1 minute, the [FiftyOne](https://fiftyone.ai/) installation takes about 10 minutes on my machine if I do not upgrade `pip`, for example, because of building some of its dependencies from source). The [FiftyOne](https://fiftyone.ai/) version installed ( `(foe)$ fiftyone constants`) is 0.7.2. And, there is a quickstart now, which did not exist the last time I used [FiftyOne](https://fiftyone.ai/)! ```python 1(foe)$ fiftyone quickstart ``` Voila! I see the quickstart dataset popup in my browser. ![](https://cdn.sanity.io/images/h6toihm1/production/97da003bade5b163c35a6018a0f491d3df78ab29-1400x1060.png?auto=format&dpr=2&fit=max&q=75&w=1400) Immediately, I’m excited by the leaps in analysis capabilities and UX that the tool has seen in the last six months. I’ll talk about a lot of these aspects below. But, one interesting surprise to me is the elegant way the interface in the browser handles disconnection from the server (in this case, when I hit `CTRL-C` from the bash shell that was running the quickstart). ![](https://cdn.sanity.io/images/h6toihm1/production/8481ef8d76c5669031a0fe4292ba3d7642a49bd8-1400x1060.png?auto=format&dpr=2&fit=max&q=75&w=1400) ## Analysis: Working with new Datasets **What’s New?** Numerous filtering and visualization capabilities that help a machine learning scientist rapidly analyze and build an intuition for a dataset. When we launched [FiftyOne](https://fiftyone.ai/), it had support for visualizing image datasets, including their annotations, as well as support for adding dynamic fields to the dataset, such as tags. The interface did allow you to easily filter based on tags at the time, but not much more. Wow, have things changed. Now, you can directly analyze the distributions over not only tags, but also labels and scalar fields captured in your dataset. The UX has a handy toggle button to expose these graphing tools. ![](https://cdn.sanity.io/images/h6toihm1/production/02d069ba0c0d433724a6bad85bf5b3c5cf483b63-1400x330.png?auto=format&dpr=2&fit=max&q=75&w=1400) Toggling that button slides out a graphing pane that blows my mind! Direct distributions over nearly any field in your dataset: the denumerable labels fields, the continuous scalar fields and various tags. Here are two example outputs from the quickstart dataset, which includes a scalar field called `uniqueness`, computed from the [Brain](https://voxel51.com/docs/fiftyone/user_guide/brain.html) functionality. ![](https://cdn.sanity.io/images/h6toihm1/production/912a649eadf1eeac66c28eefac2f9e80aaf60ccb-1400x530.png?auto=format&dpr=2&fit=max&q=75&w=1400) Under the hood, this graphing capability is enabled by the new [**Aggregations**](https://voxel51.com/docs/fiftyone/user_guide/using_aggregations.html) capabilities in [Fiftyone](https://fiftyone.ai/) that let you compute various summary statistical measures over your dataset and its properties. The importance of understanding how your labels and model outputs are distributed in aggregate on the dataset is huge. Label imbalance, for example, remains a technical challenge for training robust models that will transfer into operational production and meet the same performance measures that were observed in the lab. If the machine learning scientist does not adequately study and monitor dataset imbalance, he or she runs the great risk of delivering a tool that will underperform in practice. It is his or her responsibility to explicitly analyze dataset balance, for example. ![](https://cdn.sanity.io/images/h6toihm1/production/e2f49b045b249c5b11203ff6e80af18585f44c18-743x499.png?auto=format&dpr=2&fit=max&q=75&w=743) Beyond aggregate analysis, the last six months have seen major enhancements to analyze individual data samples and groups of samples based on various characteristics, such as the confidence on a certain model output (as I show on the left), or on more sophisticated functions based on interesting properties of the data and model output, such as object size, for example. The UX has greatly improved to allow for easy manipulation of many of these properties too with sliders and direct ability to filter images by label, as the example. When more sophisticated filtering is required than these easy-to-use elements, the app now allows power users to control all aspects of filtering, sorting, and searching datasets and dataset samples directly in the UX. These enhancements are captured in the [**view bar**](https://voxel51.com/docs/fiftyone/user_guide/app.html#using-the-view-bar). For example, as one user in [our Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-gtpmm76o-9AjvzNPBOzevBySKzt02gg) recently asked: “In the app, is there a way to search for a sample with a filepath that contains a given string?” Why yes! Place the following (appropriately customized for your dataset, of course) into the view bar’s match box: ```python 1{ 2 "$expr": { 3 "$regexMatch": { 4 "input": "$filepath", 5 "regex": "\.jpg$", 6 "options": null 7 } 8 } 9} ``` This is the same as the following direct Python expression: ```python 1view = dataset.match(F(“filepath”).re_match(“.jpg$”)) ``` The view bar unlocks huge query possibilities. But, I must caution you: it is so powerful it can overwhelm. We are considering ways to improve the UX for the view bar and its capabilities — ideas welcome! Stay tuned. ## Access: Going to the Zoo **What’s New?** A native [dataset zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html) and an innovative [model zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html) that both work out of the box to give easy access to important datasets and models. One design goal we had from the beginning with [FiftyOne](https://fiftyone.ai/) was to make the user’s work as simple as possible and meet them in their context. This design goal gave rise to a number of valuable properties of the tool at launch, such as the ability to work on a laptop on an airplane with no internet connection (the pandemic notwithstanding). However, we knew we could do better and help our users spend increasingly more of their time directly on computer vision and machine learning science, rather than tedious scripting and data wrangling. To that end, we have added not one but two Zoos to the [open-source library](https://github.com/voxel51/fiftyone). We have a [dataset zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html) and a [model zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html). At the surface, neither seems exceptionally innovative. But, when you look deeper, well, you’ll smile. The [dataset zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html), for example, currently has about two dozen datasets across the computer vision problem spectrum. It is more flexible than other dataset zoos. For example, works whether you use a TensorFlow or PyTorch backend. And, more importantly, since [FiftyOne](https://fiftyone.ai/) support so [many different dataset formats and schemas](https://voxel51.com/docs/fiftyone/user_guide/dataset_creation/index.html), the [dataset zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html) functionality can act as a Babel-like translator to help you get any of these datasets into a format you need to get your work done. ![](https://cdn.sanity.io/images/h6toihm1/production/79e6ef7058a3c135c4cfeb07a74cb9e4559e5183-1400x1152.png?auto=format&dpr=2&fit=max&q=75&w=1400) The [model zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html) **sets the bar for model sharing**. Yes, our broader community has gotten better at sharing code (e.g., [https://paperswithcode.com/](https://paperswithcode.com/)). Yet, it remains nebulous when you actually want to quickly test out an idea from a paper or try out a new model. Whereas existing model zoos present what amounts to an unconstrained bazaar of code, the [FiftyOne model zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html) is a curated collection of models that will **run out of the box** with no configuration. There are dozens of such models across the Tensorflow and Torch backends in [FiftyOne](https://fiftyone.ai/). Finally, an important reminder about the tool: since [FiftyOne is open-source](https://github.com/voxel51/fiftyone), one natural way to contribute is to add a new dataset or model to the zoo! This will help make [FiftyOne](https://fiftyone.ai/) even more valuable. Do note that given the advantages of the [FiftyOne](https://fiftyone.ai/) zoos noted above, there are more strict requirements for how to add something to them in contrast to other zoos, which, well, are more like jungles. ## Video: FiftyOne now supports video datasets **What’s New?** The section title says everything: [FiftyOne](https://fiftyone.ai/) now supports both image and video datasets natively. In some sense, it seems like video has always been the black sheep in computer vision. The number of image-based papers seems to always dwarf the number of video-based papers. Video is harder to work with: it is bigger, more complex, and exposes a new layer of potential inference problems. Yet, the number of video cameras worldwide is increasing at an alarming pace ( [nearly 45 billion video cameras today](https://www.ldv.co/blog/2017/8/8/45-billion-cameras-by-2022-fuel-business-opportunities)). Interestingly, video has taken an enhanced role in the recent literature, perhaps catapulted by the potential for self-supervision and related ideas that reduce the need for extensive and error-prone manual annotation. Well, I definitely have a bit of a proverbial _chip_ on my shoulder about this(as you may know, I [focus on video in my academic research](https://scholar.google.com/citations?hl=en&user=g9bV-_sAAAAJ)). So, you can imagine how excited I was when the [FiftyOne](https://fiftyone.ai/) dev team announced the native support for video (numerous video datasets, many video mtasks, even a state of the art [javascript-based annotated video renderer](https://github.com/voxel51/player51)) in [FiftyOne](https://fiftyone.ai/). ![](https://cdn.sanity.io/images/h6toihm1/production/719bd11960db84c9699a1ba4db89801a6397f135-1400x1015.png?auto=format&dpr=2&fit=max&q=75&w=1400) Video support is easily explored using this command: ```python 1fiftyone quickstart --video ``` ## Sharing: Notebooks and Colab **What’s New?** [FiftyOne](https://fiftyone.ai/) works in the Jupyter and Colab ecosystem while nicely maintaining the notebook metaphor for recordability and sharing. In practice, computer vision and machine learning seldom happen by lone scientists in dark, late night offices as we might have observed in [classical hackers](https://www.goodreads.com/book/show/281818.Where_Wizards_Stay_Up_Late). No, CV/ML is increasingly a team effort complete with somewhat [mature analyses over how to best compose the team](https://towardsdatascience.com/how-to-organise-your-machine-learning-teams-for-success-199f544afd20), suggesting we’re well beyond the direct appreciation of the need for a team in the first place. This is true for numerous organizations I have been a part of over the last decade, including [Voxel51](https://voxel51.com/) when we have had active machine learning application projects. And, when you have teams, you need to cooperate by sharing your work and by maintaining a good record of what you’ve done. The current _de facto_ approach to both is now through the use of [Jupyter Notebooks](https://jupyter.org/), also open-source, which make it natural to both share and record your work. Numerous services, such as [Google Colab](https://colab.research.google.com/) (which I’ve used as the backdrop for giving assignments in [my computer vision course in recent offerings](https://web.eecs.umich.edu/~jjcorso/t/)!), then make it easy to [run the notebook in the cloud](https://www.dataschool.io/cloud-services-for-jupyter-notebook/). So, **you could imagine my jaw dropped to the floor** when the team was able to add native Jupyter and Colab support to [FiftyOne](https://fiftyone.ai/). I mean — [Fiftyone](https://fiftyone.ai/) is an app — how can you inject an interactive app into a notebook? You see, my understanding of notebooks is operating under the well appreciated _notebook metaphor_: I can write things down in the notebook, I can package it up and share it with a friend, I can record my work in the notebook, I can compute things I’ve written in the notebook and then record their outputs (in the notebook), and so on. So, how could an interactive app be integrated within the notebook and yet, maintain this natural metaphor? Well, the [FiftyOne](https://fiftyone.ai/) team did it. [FiftyOne notebooks](https://voxel51.com/docs/fiftyone/environments/index.html#notebooks) indeed do let you incorporate the interactive app inside of the notebook so you can do your visual data analysis, dataset quality and model performance work like you would with an IPython session for example. But, you have the added benefit of then being able to ship off your work to share with a colleague, or come back to it and rerun it in a Colab instance. The Jupyter integration for [FiftyOne](https://fiftyone.ai/) maintains the notebook metaphor by automatically recording an image of your work for each cell in which the app was invoked; each of the cells are connected to the same underlying Python session to avoid memory overload. Then each such recorded image can be reactivated and further manipulated. You have to see it to believe, it’s seriously amazing; try the [quickstart in Colab](https://colab.research.google.com/github/voxel51/fiftyone-examples/blob/master/examples/quickstart.ipynb). Here is an example of a Jupyter notebook in action. ![](https://cdn.sanity.io/images/h6toihm1/production/6bfc986ccb78da02f3aae9b8ed25c39df9a0c852-1277x1209.gif?auto=format&dpr=2&fit=max&q=75&w=1277) This really pushes the foundational boundary of what is possible with Jupyter notebooks. We’re the first — to my knowledge — interactive app inside of a notebook that still nicely maintains the true notebook metaphor. If you’re interested in learning more about the notebook and Colab support for FiftyOne, I suggest you check out [Ben’s recent blog post](https://medium.com/voxel51/on-notebooks-and-the-future-of-computer-vision-9908da4f9c04). ## Wrapping Up Thanks for taking the time to walk through some of the cool enhancements [FiftyOne](https://fiftyone.ai/) has seen in the last six months since its initial launch. There are many more enhancements to the tool at both the core, library level and the user interface that I could not include in here for time. The [docs](https://fiftyone.ai/) are a good place to go to learn more about the capabilities that make [FiftyOne](https://fiftyone.ai/) the new standard in open-source computer vision and machine learning developer tools for dataset analysis and model performance analysis. [Add it to your toolbox today!](https://voxel51.com/docs/fiftyone/getting_started/install.html) We always love to hear from our users: what are you learning, how are you using FiftyOne in your CV/ML work, what features would help? There are two ways to connect to us: [our interactive Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-gtpmm76o-9AjvzNPBOzevBySKzt02gg) for general discussion and help, and our [GitHub repo](https://github.com/voxel51/fiftyone) for issues, forks, and stars! [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Voxel51 milestone](https://voxel51.com/blog/tag/voxel51-milestone) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/b2e0ca21c16fec18362fd532541417ccd9969a8c-1400x1127.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Turns One!\\ \\ Product & News\\ \\ • \\ \\ Aug 12, 2021](https://voxel51.com/blog/fiftyone-turns-one) [![](https://cdn.sanity.io/images/h6toihm1/production/e6ca14f73ab3cebac9a67e469fc0cb4f87f7b08b-960x640.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Making of Avatar: The Way of Water\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jan 19, 2023](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water) [![](https://cdn.sanity.io/images/h6toihm1/production/7fabed74e8e4e741cb20a5be6a3c073bc6d71627-1400x787.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ On Notebooks and the Future of Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Feb 1, 2021](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-248-lllmstxt|> ## Future of Computer Vision [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) On Notebooks and the Future of Computer Vision Feb 1, 2021 • 11 min read Article content In this article [A Scientific State of Affairs](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#4712cb574192) [Wolfram’s Walled Garden](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#b9e7bf91764a) [Python, Jupyter, and So Much More](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#1b1ab6bb445f) [FiftyOne and Jupyter](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#9837adea8282) [Digging Into COCO, with FiftyOne](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#d5a696cffe40) [Summary](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#167944d8b433) In this article [A Scientific State of Affairs](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#4712cb574192) [Wolfram’s Walled Garden](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#b9e7bf91764a) [Python, Jupyter, and So Much More](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#1b1ab6bb445f) [FiftyOne and Jupyter](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#9837adea8282) [Digging Into COCO, with FiftyOne](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#d5a696cffe40) [Summary](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision#167944d8b433) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _Why notebooks have been great for CV/ML and what can still be learned from them._ ![](https://cdn.sanity.io/images/h6toihm1/production/7fabed74e8e4e741cb20a5be6a3c073bc6d71627-1400x787.png?auto=format&dpr=2&fit=max&q=75&w=1400) [FiftyOne](http://fiftyone.ai/), an open-source productivity for tool analyzing visual datasets, includes Jupyter notebook support, a feature which I implemented and was technical lead for. I set out to write an article describing the importance of such a feature — why [automatic screenshotting](https://voxel51.com/docs/fiftyone/environments/index.html#notebooks) is great for sharing your visual findings with others, why having code and its often visual outputs in one place has been so great for CV/ML, and how using [FiftyOne](http://fiftyone.ai/) _in_ notebooks can unlock even more of the a paradigm established by notebooks for CV/ML engineers and researchers. This article hopes to do all of those things, but in its essence, it was also pulled in another direction as I learned the history of scientific notebooks. A history that begins with scientific research itself. And a history thematic of an openness that [FiftyOne](http://fiftyone.ai/) hopes to build upon. So much more than Jupyter notebooks is needed by the CV/ML community for visual research and analysis. The final section of this article is a step-by-step guide to using [FiftyOne](http://fiftyone.ai/) in notebooks to help you find problems with your visual datasets. The entire section can found in full via [Google’s Colab](https://colab.research.google.com/github/voxel51/fiftyone-examples/blob/master/examples/digging_into_coco.ipynb), as well. We will see how to confirm a common failure mode of an image detection model and identify annotation mistakes, with very few lines of code, all while visualizing results each step of the way. But before that, indulge me in an attempt to explain why so much more than Jupyter notebooks is needed by the CV/ML community for visual research and analysis. ## A Scientific State of Affairs The collaborative efficiency of the scientific paper has reached a bottleneck. In April of 2018, The Atlantic published an [article](https://www.theatlantic.com/science/archive/2018/04/the-scientific-paper-is-obsolete/556676/) declaring the obsolescence of the scientific paper as we know it. In truth, it may have merely been a willful prediction, or, in the least, shameless hyperbole for the internet to gobble up. But the history of scientific publishing it outlines is inarguable. The collaborative efficiency of the scientific paper has reached a bottleneck, nearly [400 years after its creation](https://en.wikipedia.org/wiki/Scientific_journal#:~:text=The%20history%20of%20scientific%20journals,began%20systematically%20publishing%20research%20results.). The steady and incremental progress that was made by a “buzzing mass” of scientific researchers, to paraphrase the [article](https://www.theatlantic.com/science/archive/2018/04/the-scientific-paper-is-obsolete/556676/), is now no longer suited for the times. Hundreds, perhaps even thousands, of researchers now publish in a field, not dozens. And results are often no longer computed by hand, but by computers, software, and datasets of once incomprehensible size. Repeatability has become more important now than ever because of this scale. But often the steps taken to enable fellow researchers to reproduce one’s results are inadequate. All too often the code and data provided is incomplete, if it is provided at all, and complex ideas that are formulated using rich and dynamic visualizations are described with abstract language and reductive and static diagrams. The way in which research is shared no longer matches the way in which it is done. ## Wolfram’s Walled Garden The computational complexity and scale with which modern research is done has been a boon for scientific progress and technological innovation. Juxtaposed with the stubborn form of the scientific research paper (i.e. a PDF), it has also been an acknowledged problem for decades. A solution has been around for decades, though. A solution in which research is done and shared in the same way, the same _form_ even. That solution was born as the computational “notebook” in 1988 when [Wolfram Research, founded by Steven Wolfram, released _Mathematica_](https://blog.wolfram.com/2018/06/21/weve-come-a-long-way-in-30-years-but-you-havent-seen-anything-yet/) _._ Spearheaded by [Theodore Gray](https://en.wikipedia.org/wiki/Theodore_Gray), the interface was informed by an early Apple code editor, and in part formulated with [help from none other than Steve Jobs](https://www.forbes.com/sites/ciocentral/2011/10/06/memories-of-steve-jobs/?sh=170ac7642ce1). ![](https://cdn.sanity.io/images/h6toihm1/production/48ebf2890ec5ff5cc5500f27543812463e206e22-620x352.png?auto=format&dpr=2&fit=max&q=75&w=620) For over thirty years _Mathematica_ has continued to add to the number questions it can answer for you, the number ways you can visualize data, and the amount of data at your disposal. But growth has been slow since the first decade it was introduced. Licenses are expensive, publishers don’t want to use them, and what is supported in _Mathematica_ starts and stops with Wolfram Research. It is a beautiful and incredibly capable walled garden. ## Python, Jupyter, and So Much More ![](https://cdn.sanity.io/images/h6toihm1/production/b9361842c7193145187c9a077af7825f14ca874c-1280x827.png?auto=format&dpr=2&fit=max&q=75&w=1280) As _Mathematica_ continued its march toward proprietary perfection, in early 2001 a graduate student in physics, Fernando Pérez, found himself fed up with his own ability to do research, even with _Mathematica_ at his disposal. [Per The Atlantic](https://www.theatlantic.com/science/archive/2018/04/the-scientific-paper-is-obsolete/556676/), he was enamored with the new programming language Python, and with the help of two other graduate students began a project that came to be known as IPython, the foundation of the Jupyter Project. Jupyter has triumphed over Mathematica not along the “technical dimension”, but the “social dimension”. Today, at the core of the Jupyter Project is the Jupyter Notebook. Like _Mathematica,_ Jupyter Notebook encourages scientific exploration. But unlike _Mathematica,_ it is an open-source project that anyone can contribute to. Similar to the “buzzing mass” of productivity that the scientific paper enabled neearly 400 years ago, Jupyter has triumphed over _Mathematica_ not along the “technical dimension”, but the “social dimension” as Nobel Laureate Paul Romer [noted](https://paulromer.net/jupyter-mathematica-and-the-future-of-the-research-paper). The _form_ of Jupyter Notebooks is dictated by an active community of developers and users who have a stake in its success. The field of computer vision stands firmly within the domain of scientific research. And for many years both academic and industry researchers have found incredible value in Jupyter Notebooks. Individual pieces of data are often the images and video themselves, and they need to be looked at and watched. Notebooks offer that. Reproducible visualizations need to be shared. Notebooks offer that. And the open ecosystem of Jupyter allows for any missing integration to be easily added. ![](https://cdn.sanity.io/images/h6toihm1/production/14b15185fca01a03973d94022492e316e1efcef1-1400x700.jpg?auto=format&dpr=2&fit=max&q=75&w=1400) Packages like [`matplotlib`](https://matplotlib.org/) and [`opencv`](https://opencv.org/) can be used to display images and video that need to be inspected. When training machine learning models, packages like [`tensorboard`](https://www.tensorflow.org/tensorboard) offer sample visualizations that extend inspection of images into the context of experiment tracking. [`matplotlib`](https://matplotlib.org/), [`opencv`](https://opencv.org/), [`tensorboard`](https://www.tensorflow.org/tensorboard), and countless other Python packages with visualization capabilities can all be seamlessly used within Jupyter Notebooks. Understanding data quality requires an informed sense of the trends in one’s data. Though, in CV/ML, a stark issue remains when using Jupyter notebooks. Data quality is critical to building great models. And understanding data quality requires an informed sense of the trends in one’s data. Looking at one or even a dozen images is almost always not enough to understand the performance and failure modes of a model. Furthermore, identifying individual mistakes in ground truth or gold standard labels that may only occur in one out of 1,000 or even 100,000 images requires rapid slicing and dicing of datasets to narrow in on problems. Fundamentally, there is a lack of tooling available for naturally working with visual datasets in notebooks that allows for this kind of problem solving. ## FiftyOne and Jupyter ![](https://cdn.sanity.io/images/h6toihm1/production/9ccb1c6c33d43598d8038211cac0424ec90f1429-1400x787.png?auto=format&dpr=2&fit=max&q=75&w=1400) So, “What is [FiftyOne](http://fiftyone.ai/)?”. It is an open-source CV/ML project that hopes to solve many of the practical problems and tooling problems faced by CV/ML researchers in industry and academia. It is the reason I am writing this article, and the reason why I think the history of computational notebooks contains valuable lessons. As a developer of [FiftyOne](http://fiftyone.ai/), I admire the quality of research that it encourages, and the collaboration it allows for. And I agree that progress is often best made in an open, sometimes noisy, forum. The opportunity to have machines make intelligent and automated decisions for us has prompted a massive manual undertaking. The brief history of notebooks and scientific publishing outlined earlier serve as a parallel to the many problems facing machine learning today. Specifically, in the domain of computer vision. Computer visions models attempt to make observations and decisions about unstructured, _human-made,_ artifacts of our world, i.e. images and video. Large datasets numbering in the millions of images have been imperfectly annotated by thousands of workers in the past decade in pursuit of training and understanding the performance of theses models. The opportunity to have machines make intelligent and automated decisions for us has prompted a massive manual undertaking. Ironic, to say the least. Recent efforts in more unsupervised approaches to computer vision tasks that do not require the arduous and costly work of human annotation such as OpenAI’s [CLIP](https://openai.com/blog/clip/) have proven to be fruitful. But performance still remains far from perfect. And, in the very least, understanding model performance will always require manual verification against trusted, ground truth or gold standard data. [Or at least it should](https://www.youtube.com/watch?v=hx7BXih7zx8&t=514). After all, these models are not being sent off into the woods to operate and infer among themselves. Models are being embedded into our everyday life, and the quality of their performance can have life and death consequences. The quality of methods and tools we use to analyze computer vision models should match, if not exceed, the quality of methods and tools used to build them. It stands, therefore, that the quality of methods and tools we use to analyze computer vision models should match, if not exceed, the quality of methods and tools used to build them. And that cumulative, collaborative, and incremental problem solving is paramount. Establishing open standards that enable the CV/ML community to work together, incrementally, as a “buzzing mass” not just on models, but on datasets is likely the only tractable way to establish trust and progress in the science of modern computer vision. [FiftyOne](https://voxel51.com/docs/fiftyone/) hopes to enable this establishment of trust and progress. What follows is a small example of the current capabilities of [FiftyOne](https://voxel51.com/docs/fiftyone/), focused on demonstrating the foundational APIs and UX that allow one to answer questions they have about their datasets and models efficiently, all within a Jupyter notebook. ## Digging Into COCO, with FiftyOne _Follow along in this [Colab notebook](https://colab.research.google.com/github/voxel51/fiftyone-examples/blob/master/examples/digging_into_coco.ipynb)!_ Notebooks offer a convenient way to analyze visual datasets. Code and visualizations can live in the same place, which is exactly what CV/ML often requires. With that in mind, being able to find problems in visual datasets is the first step towards improving them. This section walks us though each “step” (i.e., a notebook cell) of digging for problems in an image dataset. I encourage you to click through to the [Colab notebook](https://colab.research.google.com/github/voxel51/fiftyone-examples/blob/master/examples/digging_into_coco.ipynb) and experience the journey yourself. First, we’ll need to install the `fiftyone` package with `pip`. ```python 1pip install fiftyone ``` Next, we can download and load our dataset. We will be using the [`COCO-2017`](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#coco-2017) validation split. Let's also take a moment to visualize the ground truth detection labels using the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html). The following code will do all of this for us. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("coco-2017", split="validation") 5session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/4f481b29bf8d43ff26f60bb07ace8c10a1bede68-799x802.png?auto=format&dpr=2&fit=max&q=75&w=799) We have our [`COCO-2017`](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html#coco-2017) validation dataset loaded, now let's download and load our model and apply it to our validation dataset. We will be using the [`faster-rcnn-resnet50-fpn-coco-torch`](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/models.html#faster-rcnn-resnet50-fpn-coco-torch) pre-trained model from the [FiftyOne Model Zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html). Let's apply the predictions to a new label field `predictions`, and limit the application to detections with a confidence greater than or equal to `0.6`. ```python 1model = foz.load_zoo_model("faster-rcnn-resnet50-fpn-coco-torch") 2dataset.apply_model(model, label_field="predictions", confidence_thresh=0.6) ``` Let’s focus on issues related vehicle detections and consider all buses, cars, and trucks vehicles and ignore any other detections, in both the ground truth labels and our predictions. The following filters our dataset to a view containing only our vehicle detections, and renders the view in the App. ```python 1from fiftyone import ViewField as F 2 3vehicle_labels = ["bus","car", "truck"] 4only_vehicles = F("label").is_in(vehicle_labels) 5 6vehicles = ( 7 dataset 8 .filter_labels("predictions", only_vehicles, only_matches=True) 9 .filter_labels("ground_truth", only_vehicles, only_matches=True) 10) 11 12session.view = vehicles ``` ![](https://cdn.sanity.io/images/h6toihm1/production/8d2cf9b5d780566b8eec38ef7fb5ac4b2e06362f-799x803.png?auto=format&dpr=2&fit=max&q=75&w=799) Now that we have our predictions, we can evaluate the model. We’ll use an [`evaluate_detections()`](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.eval.coco.html?highlight=evaluate_detections#fiftyone.utils.eval.coco.evaluate_detections) utility method provided by FiftyOne that uses the COCO evaluation methodology. ```python 1from fiftyone.utils.eval import evaluate_detections 2 3evaluate_detections(vehicles, "predictions", gt_field="ground_truth", iou=0.75) ``` [`evaluate_detections()`](https://voxel51.com/docs/fiftyone/api/fiftyone.utils.eval.coco.html?highlight=evaluate_detections#fiftyone.utils.eval.coco.evaluate_detections) has populated various pieces of data about the evaluation into our dataset. Of note is information about which predictions were not matched with a ground truth box. The following view into the dataset lets us look at only those unmatched predictions. We'll sort by confidence, as well, in descending order. ```python 1filter_vehicles = F("ground_truth_eval.matches.0_75.gt_id") == -1 2 3unmatched_vehicles = ( 4 vehicles 5 .filter_labels("predictions", filter_vehicles, only_matches=True) 6 .sort_by(F("predictions.detections").map(F("confidence")).max(), reverse=True) 7) 8 9session.view = unmatched_vehicles ``` ![](https://cdn.sanity.io/images/h6toihm1/production/b2fd0634add61a921757d44a50e8c49ec2e58afc-800x797.png?auto=format&dpr=2&fit=max&q=75&w=800) If you were working in a notebook version of this walk through, you would see that the most common reason for an unmatched prediction is that there is a label mismatch. It is not surprising, as all three of these classes are in the `vehicle` supercategory. Trucks and cars are often confused in human annotation and model prediction. Looking beyond class confusion, though, let’s take a look at the first two samples in our unmatched predictions view. ![](https://cdn.sanity.io/images/h6toihm1/production/78a6db21bb13d84ecd3519193cd28b109fbd4680-1714x368.png?auto=format&dpr=2&fit=max&q=75&w=1600) The very first sample, found in the pictures above, has an annotation mistake. The truncated car in the right of the image has too small of a ground truth bounding box (pink). The unmatched prediction (yellow) is far more accurate, but did not meet the IoU threshold. The predicted box of the car in the shadow of the trees is correct, but it is not labeled in the ground truth. Raw image from the COCO 2017 detection dataset. (Images by author) The second sample found in our unmatched predictions view contains a different kind of annotation error. A more egregious one, in fact. The correctly predicted bounding box (yellow) in the image has no corresponding ground truth. The car in the shade of the trees was simply not annotated. Manually fixing these mistakes is out of the scope of this example, as it requires a large feedback loop. [FiftyOne](http://fiftyone.ai/) is dedicated to making that feedback loop possible (and efficient), but for now let’s focus on how we can answer questions about model performance, and confirm the hypothesis that our model does in fact confuse buses, cars, and trucks quite often. We’ll do this by reevaluating our predictions with buses, cars, and trucks all merged into a single `vehicle` label. The following creates such a view, clones the view into a separate dataset so we'll have separate evaluation results, and evaluates the merged labels. ```python 1vehicle_labels = { 2 label: "vehicle" for label in ["bus","car", "truck"] 3} 4 5merged_vehicles_dataset = ( 6 vehicles 7 .map_labels("ground_truth", vehicle_labels) 8 .map_labels("predictions", vehicle_labels) 9 .exclude_fields(["tp_iou_0_75", "fp_iou_0_75", "fn_iou_0_75"]) 10 .clone("merged_vehicles_dataset") 11) 12 13evaluate_detections( 14 merged_vehicles_dataset, "predictions", gt_field="ground_truth", iou=0.75) 15 16session.dataset = merged_vehicles_dataset ``` ![](https://cdn.sanity.io/images/h6toihm1/production/555b12cdb331c054f6585743d088ef1102c57f52-798x797.png?auto=format&dpr=2&fit=max&q=75&w=798) Now we have evaluation results for the originally segmented bus, car, and truck detections and the merged detections. We can now simply compare the number of true positives from the original evaluation, to the number of true positives in the merged evaluation. ```python 1original_tp_count = vehicles.sum("tp_iou_0_75") 2merged_tp_count = merged_vehicles_dataset.sum("tp_iou_0_75") 3 4print("Original Vehicles True Positives: %d" % original_tp_count) 5print("Merged Vehicles True Positives: %d" % merged_tp_count) ``` We can see that before merging the `bus`, `car`, and `truck` labels there were 1,431 true positives. Merging the three labels together resulted in 1,515 true positives. Original Vehicles True Positives: 1431 Merged Vehicles True Positives: 1515 We were able to confirm our hypothesis! Albeit, a quite obvious one. But we now have a data-backed understanding a common failure mode of this model. And now this entire experiment can be shared with others. In a notebook, the following will screenshot the last active App window, so all outputs can be statically viewed by others. ```python 1session.freeze() # Screenshot the active App window for sharing ``` ## Summary Notebooks have emerged as a popular medium for performing and sharing data science, especially in the field of computer vision. But working with visual datasets has historically been a challenge, a challenge we hope to address with an open tool like [FiftyOne](http://fiftyone.ai/). The notebook revolution is largely still in its infancy and will continue to evolve and become an even more compelling tool for performing and communicating ML projects in the community, hopefully, in part, because of [FiftyOne](http://fiftyone.ai/)! Thanks for following along! The [FiftyOne](http://fiftyone.ai/) project can be found on [GitHub](https://github.com/voxel51/fiftyone). If you agree that the CV/ML community needs an open tool to solve its data problems, give us a star! [Colab notebook](https://voxel51.com/blog/tag/colab-notebook) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Jupyter Notebook](https://voxel51.com/blog/tag/jupyter-notebook) [notebooks](https://voxel51.com/blog/tag/notebooks) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/e6ca14f73ab3cebac9a67e469fc0cb4f87f7b08b-960x640.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Making of Avatar: The Way of Water\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jan 19, 2023](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water) [![](https://cdn.sanity.io/images/h6toihm1/production/61b9a72c1362d209b3cb768c4ddc7f84cdf22452-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Automatically Set Up a New ML Project, Pain Free\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Feb 8, 2023](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ec2410f8c8844aa682fea044da0a14e3d9c5c7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ VoxelGPT: Your AI Assistant for Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jun 7, 2023](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-249-lllmstxt|> ## Spotlight on Jimmy Guerrero [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) People @ Voxel51: Spotlight on Jimmy Guerrero Jan 24, 2023 • 3 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/f0f8a8eaca85ea77f679a22d255bc344c572ce0e-1280x720.jpg?auto=format&dpr=2&fit=max&q=75&w=1280) At Voxel51, every day we wake up with the mission to bring transparency and clarity to the world’s data. It is exhilarating and meaningful work, but it doesn’t happen in a vacuum - it’s the product of the community and the team that create and support the [open source FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. We wanted to kick off a series of posts to take you under the hood and introduce you to some of those incredible folks on our team. It’s our hope you can get a sense of what it’s like to do this work - directly from the people who are pushing our mission forward. Oh! And if this seems like something you’d want to be part of, then definitely check out our remote, [open positions](https://voxel51.com/jobs/). We're building a fully-remote team of exceptional and diverse people who want to help bring data-centric machine learning to the world. We’re growing quickly and are looking for people to grow with us! And as always, if you like our open source and community work, please take a moment to [give the FiftyOne project a star](https://github.com/voxel51/fiftyone). Okay, this week we catch up with VP of Developer Relations and Marketing, [Jimmy Guerrero](https://www.linkedin.com/in/jiguerrero/). Can’t wait for you to meet him :) ![](https://cdn.sanity.io/images/h6toihm1/production/51abc555056fe8669511f4ecc4f7e9f64612469b-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) **What do you do at Voxel51?** I am the VP of Developer Relations and Marketing, which means doing whatever it takes to grow the FiftyOne community, keep members engaged, and create a platform for them to tell their stories. **Where are you based?** I reside in the San Francisco Bay Area. **How long have you been working at Voxel51?** About five months, and everyone has been super awesome in helping me bring a ton of projects to production in a very short amount of time! For example: - Established the [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/), which now has 2400+ members - Put into motion a weekly series of Tips and Tricks and dataset exploration blogs - Announced a [Community Rewards](https://www.voxel51.com/fiftyone-computer-vision-success-story-submission/) program to acknowledge the contributions of FiftyOne community members - Launched an initiative that enables the FiftyOne community to help guide Voxel51’s [charitable giving](https://www.voxel51.com/charitable-giving/) **What has been your favorite project so far?** Connecting with folks in the FiftyOne community. They have given me the best insights into what Voxel51 should be doing to grow and support the community. **What do you like most about your role?** Everyday brings unexpected challenges. Which is another way of saying, everyday presents itself with new opportunities. **What do you think is the most valuable lesson you have learned while working at Voxel51 so far?** Embrace and use to your advantage the asynchronous nature of the way folks in a remote-only organization work. Everyone in the company respects the fact that folks have a life outside of their job. With this type of an understanding (and expectation), it is easy to find pockets of time to focus on more involved projects that require concentration. ![](https://cdn.sanity.io/images/h6toihm1/production/f6ac6b82398228f7212b86626f27da8ea7fd9944-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) **What do you like to do when you aren’t working?** Vintage motorcycle restoration and land speed racing. **Three words that best describe you.** Ideas in motion. **What is the one thing you can’t live without?** I am going to turn this question around. What is the one thing that cannot live without me? My three-legged cat. She’s been obsessed with me since she was a kitten. **Where’s your favorite place in the world?** [Joshua Tree, California](https://www.nps.gov/jotr/index.htm) because it is quiet and you can leave most modern distractions behind for a few days. Once you get into the backcountry, it is unlikely you will run into another person for as long as you stay out there. So, it’s a bit like a silent retreat. **What’s something you’re planning on doing in the next year that you’ve never done?** I am building a time machine! Actually, I am restoring a [1931 Harley-Davidson VL](https://wallup.net/wp-content/uploads/2019/10/137385-1931-harley-davidson-v-l-retro-748x492.jpg) that I intend to ride solo across America. **Are there open positions on your team?** Yes! We are looking for developer evangelists to join the team. If you love connecting with other people, love learning and sharing, have a passion for open source, computer vision, and creating a sense of community…then consider applying for the [developer relations position](https://voxel51.com/jd/?4067392005?gh_jid=4067392005)! [company culture](https://voxel51.com/blog/tag/company-culture) [developer relations](https://voxel51.com/blog/tag/developer-relations) [Hiring](https://voxel51.com/blog/tag/hiring) [open positions](https://voxel51.com/blog/tag/open-positions) [open roles](https://voxel51.com/blog/tag/open-roles) [people at Voxel51](https://voxel51.com/blog/tag/people-at-voxel51) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/99dc871d875865fc200f14931a3cfb3118ded624-1020x1007.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Tunnel vision in computer vision: can ChatGPT see?\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 16, 2022](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see) [![](https://cdn.sanity.io/images/h6toihm1/production/a735267ad7effa9f799f850ab7c8ffa241088710-1024x1024.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Why 2022 was the most exciting year in computer vision history (so far)\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 14, 2022](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far) [![](https://cdn.sanity.io/images/h6toihm1/production/0d65f67f35ed7c2afbfc5e6978967e492b357c43-1250x705.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Computer Vision Meetup Update — November ‘22\\ \\ Product & News\\ \\ • \\ \\ Nov 9, 2022](https://voxel51.com/blog/computer-vision-meetup-update-november-22) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-250-lllmstxt|> ## Computer Vision in Agriculture [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Industry Solutions](https://voxel51.com/blog/category/industry-solutions), [Product & News](https://voxel51.com/blog/category/product-news) Why Computer Vision in Agriculture is the Future Jan 31, 2023 • 12 min read Article content In this article [Overview: Computer Vision in Agriculture](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#d901ccec2cf8) [Key industry challenges in agriculture](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#1870da8a4780) [Applications of computer vision in agriculture](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#e8748588eba7) [Companies at the cutting edge of computer vision in agriculture](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#d738fb9a63c9) [Agriculture datasets and competitions](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#70da0675ad68) [Join the FiftyOne community!](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#785f8e5b678f) In this article [Overview: Computer Vision in Agriculture](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#d901ccec2cf8) [Key industry challenges in agriculture](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#1870da8a4780) [Applications of computer vision in agriculture](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#e8748588eba7) [Companies at the cutting edge of computer vision in agriculture](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#d738fb9a63c9) [Agriculture datasets and competitions](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#70da0675ad68) [Join the FiftyOne community!](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023#785f8e5b678f) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/f0f8a8eaca85ea77f679a22d255bc344c572ce0e-1280x720.jpg?auto=format&dpr=2&fit=max&q=75&w=1280) Welcome to the first installment of [Voxel51](https://voxel51.com/)’s computer vision industry spotlight blog series. Each month, we will highlight how different industries – from construction to climate tech, from retail to robotics, and more – are using computer vision, machine learning, and artificial intelligence to drive innovation. We’ll dive deep into the main computer vision tasks being put to use, current and future challenges, and companies at the forefront. In this inaugural edition, we’ll focus on computer vision in _agriculture_! Read on to learn about how agriculture computer vision is transforming the sector. ## Overview: Computer Vision in Agriculture The agricultural industry is ripe for new computer vision-based AI applications. The world’s farmers are tasked with feeding the planet by growing healthier and more productive food, feed, fiber to feed, while also taking care of their land and resources. As in other industries, [in agriculture computer vision is also used](https://voxel51.com/computer-vision-use-cases/agriculture/) to drive new innovations and unlock new efficiencies to achieve goals against a backdrop of modern challenges. Before we dive into several popular applications of computer vision technologies in agriculture, here are some of the industry’s challenges where CV and AI can help. ## Key industry challenges in agriculture - A growing worldwide population: The number of humans is [expected to reach 9.8 billion by 2050](https://www.un.org/en/desa/world-population-projected-reach-98-billion-2050-and-112-billion-2100#:~:text=COVID%2D19-,World%20population%20projected%20to%20reach%209.8%20billion%20in,and%2011.2%20billion%20in%202100), leading to [dramatic increases in food demand](https://www.nature.com/articles/s43016-021-00322-9). - A reduction in arable land: The amount of arable land on Earth is shrinking, with some studies suggesting that [farmable land could be _halved_](https://www.nationalgeographic.com/science/article/partner-content-how-farm-our-unfarmable-land) in the next quarter century. - A shrinking workforce: The [number of people working in agriculture is falling](https://migration.ucdavis.edu/rmn/blog/post/?id=2510), from 40% of global workers in 2000 to just 27% of global workers in 2019. [Labor shortages](https://agamerica.com/blog/the-impact-of-the-farm-labor-shortage/) can result in skeleton crews overseeing hundreds of thousands of acres. - An increase in climate-related disruptions: The increasing frequency of extreme weather events is expected to lead to decreased crop productivity. - Pesky pests: According to the Food and Agriculture Organization (FAO), up to 40% of all crops worldwide are lost due to pests. Damages from plant diseases alone total $220 billion per year. In other words, the agriculture industry will need to feed far more people with fewer resources in a complex and changing environment. The invention and adoption of new technologies will be crucial to overcoming these challenges. AI in agriculture is already [valued at more than $1 billion annually](https://www.fnfresearch.com/news/global-ai-in-agriculture-market), and with a [compound annual growth rate (CAGR) of 20%](https://www.globenewswire.com/news-release/2021/01/12/2156893/0/en/Global-AI-in-Agriculture-Market-Size-Share-Estimated-to-Reach-USD-2-400-million-by-2026-Facts-Factors.html), it is projected to reach $2.6 billion by 2026. In agriculture, computer vision applications account for a large portion of this existing market, as well as the majority of its anticipated growth. Continue reading for some ways computer vision applications are helping organizations in agriculture. ## Applications of computer vision in agriculture ### Precision agriculture With escalating prices for pesticides, herbicides, and seeds, precision agriculture is helping farmers reduce their costs and get more from their land. As the name suggests, precision agriculture is all about finer-grained control over existing processes, from the placement of crops, to the constitution of the soil, to the application of chemical agents. Computer vision is coalescing with robotics and other emerging technologies to bring this level of precision to agriculture. In precision agriculture, computer techniques are used in the field to make real-time decisions. Object detection techniques are used to identify and localize individual insects and weeds for the application of pesticide, and herbicide, respectively. Precision control also goes beyond _where_, to _how much_: nonlinear regression models based on the coloring, size, and other visual attributes of plants can predict exactly what quantity of each chemical the plant should receive. This means optimizing returns while also conserving resources. For an overview of computer vision techniques in precision agriculture, see [_Machine Vision Systems in Precision Agriculture for Crop Farming_](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8321169/). ### Precision livestock farming Precision livestock farming, or PLF, aims to gain fine-grained insight into and achieve precise control over processes involving cattle, sheep, and other livestock. PLF can be applied for the purposes of maximizing yield, monitoring or ensuring animal health, or decreasing operational carbon footprint. In precision livestock agriculture farming, computer vision techniques are often used in conjunction with GPS tracking and audio signals to generate insights. Together, these techniques can be used to not only identify and track individual animals, but also to analyze their volume, gait, and activity levels. Here are just a few of the ways computer vision has been utilized in PLF: - [Kinetic depth sensors for classification and detection of aggressive behavior in pigs](https://www.mdpi.com/1424-8220/16/5/631) - [Optical flow for prediction of feather damage in laying hens](https://royalsocietypublishing.org/doi/full/10.1098/rsif.2010.0268) - [Remote thermal imaging and night vision technology for improving endangered wildlife resource management](https://iopscience.iop.org/article/10.1088/1742-6596/15/1/035/meta) For more information about precision livestock farming, see [Image Analysis and Computer Vision Applications in Animal Sciences: An Overview](https://www.frontiersin.org/articles/10.3389/fvets.2020.551269/full) and [Exploring the Potential of Precision Livestock Farming Technologies to Help Address Farm Animal Welfare](https://www.frontiersin.org/articles/10.3389/fanim.2021.639678/full). ### Autonomous farm equipment Autonomous farm equipment is important in enabling fewer farmers to farm more land. Along with advances in robot guidance and control in agriculture, computer vision is helping farmers automate their operations. Machine learning models for object detection and segmentation are [making their way](https://www.cnet.com/tech/mobile/john-deere-breaks-new-ground-with-self-driving-tractors-you-can-control-from-a-phone/) onto [tractors](https://www.mdpi.com/2075-1702/10/2/129) and harvesters. Stay tuned for our upcoming industry spotlight blog post about computer vision in autonomous vehicles! ### Crop monitoring To combat crop loss, farmers use data from suites of soil sensors, localized weather forecasts, and multi-level imagery to remotely monitor large tracts of land. This data can be synthesized into “crop intelligence” allowing farmers to take informed action before it is too late. On the computer vision side, images from satellites, drones, and high-resolution cameras are used for [early disease detection and monitoring](https://www.hindawi.com/journals/cin/2017/2917536/), soil condition monitoring, and [yield estimation](https://www.frontiersin.org/articles/10.3389/fpls.2019.00621/full). Some examples include: - [Automatic plant disease diagnosis using mobile capture devices, applied on a wheat use case](https://www.sciencedirect.com/science/article/abs/pii/S016816991631050X?via%3Dihub) - [Deep learning-based detection of seedling development](https://plantmethods.biomedcentral.com/articles/10.1186/s13007-020-00647-9) - [Soil color analysis based on a RGB camera and an artificial neural network towards smart irrigation: A pilot study](https://www.sciencedirect.com/science/article/pii/S2405844021001833) - [Deep Gaussian Process for Crop Yield Prediction Based on Remote Sensing Data](https://cs.stanford.edu/~ermon/papers/cropyield_AAAI17.pdf) For a more thorough discussion of crop monitoring research and applications, see [_Computer vision technology in agricultural automation —A review_](https://www.sciencedirect.com/science/article/pii/S2214317319301751#b0185). ### Plant phenotyping Climate change [will subject many plants](https://www.nps.gov/articles/000/plants-climateimpact.htm) to increased temperatures, higher levels of carbon dioxide, and more variable precipitation. It will also [make extreme weather events far more common](https://royalsociety.org/topics-policy/projects/climate-change-evidence-causes/question-13/). Some plants will be better equipped than others to survive and thrive. [_Plant phenotyping_](https://spj.science.org/doi/10.34133/2019/7507131) is the process of identifying and understanding how genetic and environmental factors manifest physically, in a plant’s phenome. One of the most important uses of computer vision in agriculture is becoming an important tool in what is widely recognized as a [key to global food security](https://www.nature.com/articles/nrg2897), with non-invasive detection, segmentation, and 3D reconstruction techniques giving researchers detailed information about everything from leaf area to a plant’s nutrient levels and biomass. As an example, segmentation of nuclear magnetic resonance images [can be used to](https://pubs.geoscienceworld.org/vzj/article/12/1/vzj2012.0019/111603/In-Situ-Root-System-Architecture-Extraction-from) map the structure of a plant’s three dimensional root system, which plays an important role in the flow of water and nutrients. Workshops at [ICCV](https://cvppa2021.github.io/) and [ECCV](https://www.plant-phenotyping.org/CVPPP2020) in recent years focused on problems in plant phenotyping underscore the computer vision community’s commitment to the topic. Some papers to get you started include: - [Plant phenotyping: from bean weighing to image analysis](https://plantmethods.biomedcentral.com/articles/10.1186/s13007-015-0056-8) - [Future scenarios for plant phenotyping](https://pubmed.ncbi.nlm.nih.gov/23451789/) - [Lettuce Height Regression by Single Perspective Sparse Point Cloud](https://www.cs.usask.ca/faculty/stavness/cvppa2021/abstracts/Li_52.pdf) - [Metric Learning on Field Scale Sorghum Experiments](https://www.cs.usask.ca/faculty/stavness/cvppa2021/abstracts/Zhang_40.pdf) ### Grading and sorting After produce is plucked from the field, and before it ends up in the fresh food aisle of your local grocery store, it may be subjected to quality control processes. [For fruits and vegetables](https://www.tomra.com/en/solutions/food/fruit), this takes the form of grading and sorting based on size, shape, color, and other physical characteristics. [For grains](https://millingandgrain.com/optical-sorting-a-brighter-better-and-clearer-view-19559) and beans on the other hand, similar sorting processes are used to detect defects and filter out foreign material. While grading and sorting were traditionally performed by hand, computer vision is now helping humans with much of this work. [Optical sorting](https://en.wikipedia.org/wiki/Optical_sorting) uses image processing techniques like object detection, classification, and anomaly detection to incorporate quality control into food production and preparation. By 2027, the optical sorting market is [expected to surpass $3.8 billion](https://www.globenewswire.com/news-release/2022/10/28/2543769/28124/en/The-Worldwide-Optical-Sorter-Industry-is-Expected-to-Reach-3-8-Billion-by-2027.html). That’s expected to be one of the biggest use cases for computer vision in agriculture! One final batch of papers: - [An automated machine vision based system for fruit sorting and grading](https://ieeexplore.ieee.org/document/6461669) - [Computer Vision Based Fruit Grading System for Quality Evaluation of Tomato in Agriculture industry](https://www.sciencedirect.com/science/article/pii/S1877050916001861) - [Fruits and vegetables quality evaluation using computer vision: A review](https://www.sciencedirect.com/science/article/pii/S131915781830209X) - [Prospects of Computer Vision Automated Grading and Sorting Systems in Agricultural and Food Products for Quality Evaluation](https://www.researchgate.net/publication/43656001_Prospects_of_Computer_Vision_Automated_Grading_and_Sorting_Systems_in_Agricultural_and_Food_Products_for_Quality_Evaluation) ## Companies at the cutting edge of computer vision in agriculture ### Carbon Robotics Founded in 2018 and headquartered in Seattle, agricultural robotics startup [Carbon Robotics](https://carbonrobotics.com/) has raised $35.9 million to help farmers wage war against weeds with lasers, rather than whackers. The company’s [LaserWeeder](https://carbonrobotics.com/laserweeding-technology) technology uses the thermal energy in 30 onboard lasers to target weeds without harming crops. The LaserWeeder uses object detection to identify and precisely locate weeds. High resolution images taken by mounted cameras are fed through an onboard Nvidia GPU and predictions, generated in milliseconds, are communicated to the lasers for firing. When attached to a tractor, this implement is able to eliminate 200,000 weeds per hour. Carbon Robotics’ [Autonomous LaserWeeder](https://carbonrobotics.com/autonomous-weeder) also uses computer vision techniques, in conjunction with GPS location data, for autonomous navigation. In addition to the weed detection model, this autonomous agent is also equipped with a furrow detection model, which allows it to distinguish the trail it is supposed to follow from plant beds. The company was a sponsor for ICCV in 2021. ### OneSoil Founded in 2017, Zurich-based [OneSoil](https://onesoil.ai/en) employs computer vision techniques on satellite imagery to help farmers maximize their yields and reduce costs. In 2018, OneSoil released the [OneSoil Map](https://blog.onesoil.ai/en/onesoil-map) \- a comprehensive map of farmland across 59 countries. To generate this map, they used 250 Tb of satellite imagery shot by the European Union’s [Sentinel-2 satellite](https://www.esa.int/Applications/Observing_the_Earth/Copernicus/Sentinel-2). Using proprietary computer vision models, they detected clouds, shadows, and snow, and removed these to generate clean images. To combat the low-resolution of satellite images, they combined images taken over a multi-year span. The free OneSoil App, which was crowned the 2018 Product Hunt _AI & Machine Learning Product of the Year_, allows farmers to zoom in and select their plot of land without the need to delineate the boundaries themselves. The key to this feature is OneSoil’s field boundary detection model. [To train the model](https://blog.onesoil.ai/en/how-onesoil-uses-data-science), they worked with a number of farmers to get small samples of field boundary data, and used data augmentation operations to increase the size of their training dataset by multiple orders of magnitude. Given the ease of use, it’s no wonder that more than 300,000 farmers use the app, representing around 5% of the world’s arable land. As of 2023, OneSoil’s computer vision models can identify 12 major crop types, and the company uses this information to generate [productivity zone](https://blog.onesoil.ai/en/what-are-productivity-zones) and soil brightness maps which help farmers make better use of their land. ### Taranis Located in Westfield, Indiana, [Taranis](https://taranis.ag/) has been around for almost a decade and raised more than $100 million to bring farmers leaf-level insights making it a key player in the future of computer vision in agriculture. The Taranis team has collected and digitized a dataset consisting of over 50 million submillimeter, high-resolution images, and over 200 million data points. These images are used to train custom computer vision detection algorithms for weeds, diseased plants, insects, and nutrient deficiencies. Taranis’s [AcreForward Intelligence](https://www.taranis.com/advisors/) uses real-time imagery from multiple sources, such as drones, planes, and satellites, to identify insect damage on a per-leaf basis, detect weeds before they become a problem, find nutrient deficiencies, and count the number of plants in a field so farmers can make informed decisions about planting and usage of inputs. In 2022, one of the world’s largest venture capital firms, Andreessen Horowitz, named Taranis [one of the top 50 companies](https://www.prnewswire.com/news-releases/taranis-named-one-of-the-top-50-companies-kickstarting-american-renewal-by-andreessen-horowitz-301686342.html) kickstarting the American renewal. ### Blue River Technology A subsidiary of John Deere, [Blue River Tech](https://bluerivertechnology.com/) was acquired by the agriculture powerhouse in 2017 for a cool $305 million. When the company started, they narrowed in on lettuce farming, using computer vision and machine learning models to help space plants for maximal yield. Since these relatively humble beginnings, Blue River Tech’s computer vision capabilities have expanded to include sensor fusion, object detection, and segmentation. All of these techniques come together in their [See & Spray technology](https://www.deere.com/en/sprayers/see-spray-select/). In traditional broadcast spraying, chemicals are sprayed uniformly over an entire field. This practice leads to wasted herbicide, which is costly to the farmer, can pollute the environment, and can foster resistance to the applied chemicals. See & Spray uses object detection to identify weeds in real time so that herbicide can be applied only where it is needed, resulting in 77% reduction in herbicide. Blue River Tech takes weed detection further by housing its See & Spray technology in an autonomous tractor equipped with 6 stereo cameras. Together, these cameras allow for an on-board computer to estimate depth information for objects surrounding the tractor using sensor fusion. Color (RGB) data and depth information are fed into a semantic segmentation model, which divides the world into five categories: _drivable terrain_, _sky_, _trees_, _large objects_ such as people, animals, and buildings, and _the implement_ being used by the tractor. The autonomous tractor stops whenever a large object is detected in its path, and is trained to err on the side of caution by weighing false negative large object detections more strongly than false positives. When the tractor stops, the images are sent to the cloud to be reviewed by humans. Blue River Tech’s specialized agriculture computer vision models were trained on a growing [dataset which already contains more than one million images](https://funginstitute.berkeley.edu/news/blue-river-technology-how-robotics-and-machine-learning-are-transforming-the-future-of-farming/). ## Agriculture datasets and competitions If you are interested in exploring applications of computer vision in agriculture, check out these datasets and competitions: - [Sorghum Biomass Prediction](https://www.kaggle.com/competitions/sorghum-biomass-prediction/overview) - [Root Segmentation Challenge](https://sites.google.com/sinc.unl.edu.ar/root-segmentation-challenge/home?pli=1#h.1onuu8t69n98) - [Leaf Segmentation and Counting Challenges](https://www.plant-phenotyping.org/CVPPP2017-challenge) - [Global Wheat Dataset](http://www.global-wheat.com/) and [Challenge](https://www.aicrowd.com/challenges/global-wheat-challenge-2021/leaderboards) - [Aerial Sheep Dataset](https://huggingface.co/datasets/keremberke/aerial-sheep-object-detection) - [DeepWeeds: A Multiclass Weed Species Image Dataset for Deep Learning](https://github.com/AlexOlsen/DeepWeeds) - [PlantDoc: A Dataset for Visual Plant Disease Detection](https://github.com/pratikkayal/PlantDoc-Dataset) For a detailed discussion of openly available datasets for weed control and fruit detection, check out [A survey of public datasets for computer vision tasks in precision agriculture](https://www.sciencedirect.com/science/article/pii/S0168169920312709). If you would like to see any of these, or other computer vision agriculture datasets added to the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html), get in touch and we can work together to make this happen! ## Join the FiftyOne community! Developers of agricultural applications can benefit from FiftyOne’s ability to easily filter through the huge amounts of visual data collected daily from farms and other sources. Using [open source FiftyOne](https://github.com/voxel51/fiftyone), this data can be curated into datasets for model training, or to share with experts for annotation or analysis of CV models. Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,300+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,450+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 2,700+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 231+ repositories - 55+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [agriculture](https://voxel51.com/blog/tag/agriculture) [agriculture use case](https://voxel51.com/blog/tag/agriculture-use-case) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [industry spotlight](https://voxel51.com/blog/tag/industry-spotlight) [use case](https://voxel51.com/blog/tag/use-case) ![](https://cdn.sanity.io/images/h6toihm1/production/d58692baec7c64699806d60d25a0d14f534105fa-300x300.png?auto=format&dpr=2&fit=max&q=75&w=42) Jacob Marks Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/01272b85091efcfd9e9dda5b9044b0f080831781-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Computer Vision Is Changing Manufacturing\\ \\ Industry Solutions, Product & News\\ \\ • \\ \\ Mar 9, 2023](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing) [![](https://cdn.sanity.io/images/h6toihm1/production/7ac3933208d45ce7ee8d08eb935a43eadfa2475e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Teaching Androids to Dream of Sheep\\ \\ Computer Vision, Datasets, Product & News, Vector Search\\ \\ • \\ \\ Aug 7, 2023](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep) [![](https://cdn.sanity.io/images/h6toihm1/production/6445eab4dfdba3548381f187231e576c97c3c5a5-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Computer Vision Is Changing Healthcare\\ \\ Industry Solutions, Product & News\\ \\ • \\ \\ Aug 31, 2023](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-251-lllmstxt|> ## Automate ML Project Setup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) Automatically Set Up a New ML Project, Pain Free Feb 8, 2023 • 8 min read Article content In this article [How to use a cookiecutter template for reproducibility, consistency, data standardization, and more](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#72a9f507019a) [What is cookiecutter?](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#2574a05cc741) [Core features of a good ML project](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#a7e08e526274) [Nice to haves for a good ML project](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#eec8dfb2104b) [How to use the python ML project cookiecutter](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#3c58d8625fd8) [Final thoughts](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#9dc3ca76faa1) In this article [How to use a cookiecutter template for reproducibility, consistency, data standardization, and more](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#72a9f507019a) [What is cookiecutter?](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#2574a05cc741) [Core features of a good ML project](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#a7e08e526274) [Nice to haves for a good ML project](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#eec8dfb2104b) [How to use the python ML project cookiecutter](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#3c58d8625fd8) [Final thoughts](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free#9dc3ca76faa1) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) _Editor's note: This is a guest post by [Tarmily Wen](https://www.linkedin.com/in/tarmily-wen-562a37155/), Principal Computer Vision Research Engineer at [ADT](https://www.adt.com/)_ ## How to use a cookiecutter template for reproducibility, consistency, data standardization, and more \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop How much time have you spent trying to set up an environment for your next machine learning project? There are so many things to keep track of including installing python, python packages (and the order in which they are installed), ffmpeg, opencv, nvidia packages, cuda versions, and folder structures. That doesn’t even take into account potential conflicts and installation errors with each one of these. And this is only for the bare minimum for project setup. There are other tools that should be set up to improve quality of life. This manual process presents points of failure: time consuming, environment conflicts, lack of reproducibility, and lack of consistency. That is why I have created a template that utilizes cookiecutter to help automate the setup process for each new ML project. This presents a few benefits: consistency, reproducibility, data standardization, and most importantly speed. If you’re ready to get started, check out the [cookiecutter template for quicker and consistent python ML project setup](https://github.com/ChickenTarm/cookiecutter-python-ml-project) on GitHub. Or continue reading to learn more about it and how to use it. ## What is cookiecutter? Cookiecutter is a tool that is used to fill in templates for software projects. The values are extracted from user prompts that normally ask for the name of the project and who is contributing to it. It can also ask for what software versions are desired but that gets much trickier. Really the point of [cookiecutter](https://www.cortex.io/post/an-overview-of-cookiecutter) is to enforce a standard for project setup while also being automatic. It uses Jinja for this automation by finding and replacing variables in the project template corresponding to your custom project name.. It really is as easy as using a cookie cutter to make cookies. The template is the cookie shape and the values filled in are the icing and decorations of choice. ![](https://cdn.sanity.io/images/h6toihm1/production/f0a286d8b71f07c592aa55fea1eeb29ee823fe2a-600x439.jpg?auto=format&dpr=2&fit=max&q=75&w=600) Why manually cut out a flying pig when there is a template to help make as many as you want? [The python ML project cookiecutter template](https://github.com/ChickenTarm/cookiecutter-python-ml-project) I created adds essential features including environment reproducibility, code consistency, and data standardization, as well as some “nice to haves” like documentation generation and test automation to every ML project. ## Core features of a good ML project ### Reproducibility There are so many times that the installation process for an ML repo doesn’t encompass the entire environment needed to properly run it. That is because of the aforementioned initial manual set up process for a project. It is hard to always remember to properly document the changes made. This is why the cookiecutter template uses [docker](https://www.docker.com/). It provides a consistent environment inside of the docker container. This makes it so that anyone using the project only needs a few tools in the host operating system to properly configure, build, and run the docker image. The cookiecutter template provides the setup for the host operating system to ensure it has the correct docker setup (it is trickier when docker needs to use nvidia gpus). It also provides a Dockerfile that will automatically install all defined python packages. But there is another layer of abstraction needed because using docker can be annoying with respect to remembering all the options for building and running arguments. There are plenty of times people use a docker command that is 3-4 lines long, which makes it easy to forget a few options when running it. Docker compose offers an easy way to have short and consistent build and run options. It also helps deal with the other big challenge when dealing with docker: docker networking. There are many times when one container depends on another and a multi-container workflow is required. It is easy to forget to put containers on the same network. The cookiecutter template also uses [poetry](https://python-poetry.org/) to ensure consistent python package dependencies and help prevent [dependency hell](https://www.boldare.com/blog/software-dependency-hell-what-is-it-and-how-to-avoid-it/). A requirements.txt doesn’t properly capture all of the underlying dependencies a package like pytorch needs. The packaging and dependency management made possible by poetry prevents packages from installing conflicting versions of the same dependent package. Remember reproducibility for others is also reproducibility for yourself. There have been times when I changed my own environment, which impacted my ability to rerun old projects. And untangling the broken environment state is annoying to say the least… ![](https://cdn.sanity.io/images/h6toihm1/production/d4e79d0585954aa5e543adf889b4037530402743-725x363.jpg?auto=format&dpr=2&fit=max&q=75&w=725) ### Consistency I will admit it. Sometimes I get lazy with coding and some parts of my code are sloppier than others, or some parts are written very quickly and should be revisited at some point in the future (far future). There is an “easy” way to ensure some standard coding practices, and that is done using code formatters (black, isort), linters (flake), and type checkers (pyright). These can also be incorporated automatically when you set up continuous integration (CI) to maintain code quality. But needing to setup any CI for each project is a pain which is why the CI, using [GitHub Actions](https://github.com/features/actions), is provided in this cookiecutter to check all pushed code. It will fail the merge if all checks do not pass. This will ensure a consistent standard is present at all times in the code base. This consistency allows for easier cooperation by standardizing readability, preventing improper typing, and eliminating bad practices (such as defining a variable that is never used). ![](https://cdn.sanity.io/images/h6toihm1/production/ae6d777992394ce2c360a3a5e8aca309ff387aca-820x461.jpg?auto=format&dpr=2&fit=max&q=75&w=820) ### Data standardization Data is everything in ML. But because it is everything, it is also a hassle to get right. There are many facets of data: exploration, visualization, formating, etc. Each one can be done in many different ways. This leads to headaches and assumptions about data that can be incorrect. This is why having a central store and interface for data is important. [FiftyOne](https://github.com/voxel51/fiftyone) is chosen exactly for this feature. #### Data consistency There are many ways to visualize data: matplotlib, plotly, opencv, bokeh, etc. These all require more custom and boilerplate code written just to see and explore the data. There is also the issue where for the same task type there are differing formats and structure: annotation formats with differing folder structures, bounding box formats represented as xyxy or xywh with normalized or unnormalized values, segmentations stored as masks or polylines, etc. FiftyOne helps avoid these headaches and inconsistencies by providing an easy-to-use Python SDK that lets you parse any formats into FiftyOne which you can then visualize and integrate into your downstream tasks like annotation, model training, and evaluation. #### Data duplication If you are working on a team, then you also need to solve the problem of having different users store duplicates of the same dataset. I have worked on teams where I've seen multiple copies of the same COCO dataset between different people! This is where [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) comes into play. FiftyOne Teams extends open-source FiftyOne by enabling you to work with your teammates directly on the same datasets. For example, your model results can be appended onto samples in a FiftyOne Teams dataset and be viewed and evaluated directly by others. Since a FiftyOne Teams deployment has one central database, all users access and pull the same datasets. When everyone is working with the same data, all updates to a dataset will be reflected to everyone. For open-source users of FiftyOne, you still get the advantage of having a single source of truth for all of your datasets in your local environment. So you won't find yourself with multiple copies of the COCO dataset on your machine (I may have been guilty of this before...). ![](https://cdn.sanity.io/images/h6toihm1/production/b5e6c00ac5681e57b301b42431283789e56cd73f-1000x440.jpg?auto=format&dpr=2&fit=max&q=75&w=1000) ## Nice to haves for a good ML project These aren’t really necessary compared to the other tools, but I do believe these additional components of the cookiecutter template improve code quality and experiment tracking. [Pytorch-lightning](https://www.pytorchlightning.ai/) is a wonderful library built around pytorch that reduces boilerplate code. It lets you get to writing research code much faster in a standard manner. It also has built many integrations to allow for easy, efficient, and fast training. [Weights & Biases](https://wandb.ai/site) is easily the best experimentation tracker I have seen to date. The powerful experiment tracking and visualization provides better insight in model training. It also allows for better collaboration since everyone in the group can view the results or reports can be generated and sent to others. [Sphinx](https://www.sphinx-doc.org/en/master/usage/extensions/autodoc.html) is an auto documentation framework that scans through your code and checks the doc strings under classes and functions along with the type signature to create either .rst or .html files that allows for better viewing and sharing of the API documentation. [Nox](https://nox.thea.codes/en/stable/) is an automated test runner for python that allows for tests to be run in isolated session environments. This tool can be used to check CI, unit tests, safety tests, doc generation tests, etc. ## How to use the python ML project cookiecutter I will now walk through how to quickly create a new ML project using the [python ML project cookiecutter template](https://github.com/ChickenTarm/cookiecutter-python-ml-project). ![](https://cdn.sanity.io/images/h6toihm1/production/c12532a1f3e34ec05560d5abb5551c53f1d7af1e-600x404.gif?auto=format&dpr=2&fit=max&q=75&w=600) You just need to point cookiecutter to the GitHub repository for this template and fill in the required values: project name, email, and github username. The rest of the values can just use the default values. ![](https://cdn.sanity.io/images/h6toihm1/production/ba66ee500f81e5f0955c3c658e5f489ee230c52b-600x396.gif?auto=format&dpr=2&fit=max&q=75&w=600) This script will install docker, compose, and nvidia-container-runtime. It will also configure docker settings, namely the container runtime. ![](https://cdn.sanity.io/images/h6toihm1/production/b08e0f82b12dd119adaf77fa3c6c672a07eb7d15-600x396.gif?auto=format&dpr=2&fit=max&q=75&w=600) Your docker container will automatically set up python along with python packages for ML, CI, testing, and docs. The compose file also attaches the container to the correct docker network so the container has access to FiftyOne’s [mongdb](https://docs.voxel51.com/user_guide/config.html#configuring-a-mongodb-connection) and qdrant if you use that [integration](https://docs.voxel51.com/tutorials/qdrant.html). ![](https://cdn.sanity.io/images/h6toihm1/production/296ab074a618c46228327181b6356ea22e723b8e-1999x1591.png?auto=format&dpr=2&fit=max&q=75&w=1600)![](https://cdn.sanity.io/images/h6toihm1/production/b0099c1d79387f1122b3d427eff5e57da8bab9bf-1999x1498.png?auto=format&dpr=2&fit=max&q=75&w=1600) It is very easy to connect to a running container and check that the environment is properly set up. ![](https://cdn.sanity.io/images/h6toihm1/production/65a2bff3707555b1ed7e2228b7708ed6e33fe8d9-1999x1102.png?auto=format&dpr=2&fit=max&q=75&w=1600)![](https://cdn.sanity.io/images/h6toihm1/production/48b1dbe58a3db612f60861caf292b690c084775a-1308x780.png?auto=format&dpr=2&fit=max&q=75&w=1308) The actions page will let you know what CI tests have passed or failed. It is as easy as that to get a python ML environment up and running. In addition, pairing this cookiecutter template with a modern IDE with the ability to hook into running containers makes it feel like you are working in a local development environment even though you are actually working within a container. My IDE of choice is VSCode since it has the best docker plugin I have seen, but it is not the only one with this capability. ## Final thoughts My cookiecutter is based on [hypermodern python cookiecutter](https://github.com/cjolowicz/cookiecutter-hypermodern-python) which means there are probably a few features that may not be needed for your project: nox, unit testing, and auto doc generation. Those are nice to have but really not all projects need that level of rigor, but that is offered in the template if that is desired. It can be a lot of effort setting up a project with all of these attributes, which is why I created [this cookiecutter template](https://github.com/ChickenTarm/cookiecutter-python-ml-project) to enable projects to have these attributes without the overhead of learning to set them up manually. At the end of the day we all want the same thing – to produce quality code, but not have that get in the way of research. I hope my cookiecutter proves useful for you. Try it out and if there are any issues or improvements that can be made let me know! [cookiecutter](https://voxel51.com/blog/tag/cookiecutter) [docker](https://voxel51.com/blog/tag/docker) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne Teams](https://voxel51.com/blog/tag/fiftyone-teams) [GitHub Actions](https://voxel51.com/blog/tag/github-actions) [poetry](https://voxel51.com/blog/tag/poetry) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/e6ca14f73ab3cebac9a67e469fc0cb4f87f7b08b-960x640.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Making of Avatar: The Way of Water\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jan 19, 2023](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water) [![](https://cdn.sanity.io/images/h6toihm1/production/7fabed74e8e4e741cb20a5be6a3c073bc6d71627-1400x787.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ On Notebooks and the Future of Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Feb 1, 2021](https://voxel51.com/blog/on-notebooks-and-the-future-of-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ec2410f8c8844aa682fea044da0a14e3d9c5c7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ VoxelGPT: Your AI Assistant for Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jun 7, 2023](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-252-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – Feb 10, 2023 Feb 10, 2023 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#5b0143859944) [Isolating spurious or missing objects](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#126649d859c6) [Filtering by ID](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#99408770f9c2) [Merging datasets with shared media files](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#fd5865c6d1c2) [Exporting GeoJSON annotations](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#491fab880c2e) [Picking random frames from videos](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#0e6a1006fef0) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#0ca6838b139e) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#5b0143859944) [Isolating spurious or missing objects](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#126649d859c6) [Filtering by ID](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#99408770f9c2) [Merging datasets with shared media files](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#fd5865c6d1c2) [Exporting GeoJSON annotations](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#491fab880c2e) [Picking random frames from videos](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#0e6a1006fef0) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023#0ca6838b139e) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **Isolating spurious or missing objects** Community Slack member George Pearse asked, _“Is there a way to just get bounding boxes around the possibly missing and possibly spurious objects in my dataset?”_ Here, George is asking about how to isolate potential mistakes in ground truth labels on a dataset. When working with a new dataset, it is always important to validate the quality of the ground truth annotations. Even highly regarded and well-cited datasets [can contain a plethora of errors](https://deepomatic.com/how-we-improved-computer-vision-metrics-by-more-than-5-percent-only-by-cleaning-labelling-errors). Two such common types of errors in object detection labels are: 1. A ground truth label was _spuriously_ added to the data, and does not correspond to an object in the allowed object classes 2. An object is not annotated, so the ground truth detection is _missing_ Fortunately, the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html) provides a built-in method that identifies possible spurious and missing detections. These are stored at both the sample level and the detection level. With FiftyOne’s filtering capabilities, it is easy to create a view containing only the detections that are possibly spurious, or possibly missing, or both. In these cases, you might also find it helpful to convert the filtered view to a [PatchView](https://docs.voxel51.com/api/fiftyone.core.patches.html#fiftyone.core.patches.PatchView) so you can view each potential error on its own. Here is some code to get you started: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4from fiftyone import ViewField as F 5 6## load example dataset 7dataset = foz.load_zoo_dataset("quickstart") 8 9## find possible mistakes 10fob.compute_mistakenness(dataset, "predictions") 11 12## create a view containing only objects whose 13## ground truth detections are possibly missing 14pred_field = "predictions" 15missing_view = dataset.filter_labels( 16 pred_field, 17 F("possible_missing") > 0, 18 only_matches=True 19).to_patches(pred_field) 20 21## create a view containing only objects whose 22## ground truth detections are possibly spurious 23gt_field = "ground_truth" 24spurious_view = dataset.filter_labels( 25 gt_field, 26 F("possible_spurious") > 0, 27 only_matches=True 28).to_patches(gt_field) 29 ``` We can then view these in the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html). Inspect the possibly spurious detection patches, for instance: ```python 1session = fo.launch_app(spurious_view) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/ce6156ee7a7d1d36e78f9d4c7fc33db237971048-2546x1570.png?auto=format&dpr=2&fit=max&q=75&w=1600) Learn more about [identifying detection mistakes](https://docs.voxel51.com/tutorials/detection_mistakes.html) in the FiftyOne Docs. ## **Filtering by ID** Community Slack member Sylvia Schmitt asked, _“I am storing related sample IDs as `StringField` objects in a separate field on my data and I want to use them to match sample IDs that are stored as `ObjectIdField` objects. How do I do this?”_ If you were comparing the values in two \`StringFields\`, you could use the [ViewField](https://docs.voxel51.com/api/fiftyone.core.expressions.html#fiftyone.core.expressions.ViewField) as follows: ```python 1import fiftyone as fo 2from fiftyone import ViewField as F 3 4dataset = fo.Dataset(..) 5dataset.match(F('field_a') == F('field_b')) ``` However, sample IDs are represented as `ObjectIdField` objects. They are stored under an `_id` key in the underlying database, and need to be referenced with this same syntax, prepending an underscore. Additionally, the object needs to be converted to a string for the comparison. Here is what such a matching operation might look like: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5dataset = foz.load_zoo_dataset("quickstart") 6 7# Add a `str_id` field that matches `id` on 10 samples 8view = dataset.take(10) 9view.set_values("str_id", view.values("id")) 10 11matching_view = dataset.match( 12 F("str_id") == F("_id").to_string() 13) 14 ``` Learn more about [fields](https://docs.voxel51.com/user_guide/basics.html#fields) and [filtering](https://docs.voxel51.com/user_guide/using_views.html#filtering) in the FiftyOne Docs. ## **Merging datasets with shared media files** Community Slack member Joy Timmermans asked, _“I have three datasets, and some of my samples are in multiple datasets. I’d like to combine all of these datasets into one dataset for export, retaining each copy of each of the samples. How do I do this?”_ If your datasets were created independently, even if there are samples that have the same media files (located at the same file paths), these samples will have different sample IDs. In this case, you can create a combined dataset with the `add_collection()` method without passing in any optional arguments. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("quickstart") 5ds1 = dataset[:100].clone() 6ds2 = dataset[100:].clone() 7 8## create temporary dataset for combining 9tmp = ds1.clone() 10## add ds2 samples 11tmp.add_collection(ds2) 12 13## export 14tmp.export(..) 15## delete temporary dataset 16tmp.delete() ``` If, on the other hand, your datasets have samples with the same sample IDs, then applying the `add_collection()` method without options will only lead to the “combined” dataset having a single copy of each media file. Fortunately, you can bypass this by passing in `new_ids = True` to `add_collection()`. In your case, combining three datasets would look like: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4## start with dataset1, dataset2, dataset3 5 6tmp = dataset1.clone() 7tmp.add_collection(dataset2, new_ids = True) 8tmp.add_collection(dataset3, new_ids = True) 9tmp.export(..) 10tmp.delete() ``` Learn more about [merging datasets](https://docs.voxel51.com/recipes/merge_datasets.html) in the FiftyOne Docs. ## **Exporting GeoJSON annotations** Community Slack member Kais Bedioui asked, _“I am logging some of my production data in GeoJSON format, and I want to save it in the database in that same format. Is there a way to include the `ground_truth` label in the `labels.json` file so that when I reload the GeoJSON dataset, it comes with its annotations?”_ To do this, you can use the optional `property_makers` argument of the GeoJSON exporter to include additional properties directly in GeoJSON format. For example: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("quickstart-geo") 5 6dataset.export( 7 labels_path="/tmp/labels.json", 8 dataset_type=fo.types.GeoJSONDataset, 9 property_makers={"ground_truth": lambda d: len(d.detections)}, 10) ``` Alternatively, if you want to save the annotations but do _not_ need to save everything in GeoJSON format, you can export it as a FiftyOne `Dataset`: ```python 1import fiftyone as fo 2 3dataset.export( 4 export_dir="/tmp/all", 5 dataset_type=fo.types.FiftyOneDataset, 6 export_media=False, 7) ``` When you take this approach, all you have to do to load the dataset back in is use FiftyOne’s `from_dir()` method: ```python 1import fiftyone as fo 2 3dataset = fo.Dataset.from_dir( 4 dataset_dir="/tmp/all", 5 dataset_type=fo.types.FiftyOneDataset, 6) ``` Learn more about [from\_dir()](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#basic-recipe), and [importing](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#loading-datasets-from-disk) and [exporting data](https://docs.voxel51.com/user_guide/export_datasets.html) in the FiftyOne Docs. ## **Picking random frames from videos** Community Slack member Joy Timmermans asked, _“Is there an equivalent of `take()` for frames in a video dataset so that I can randomly select a subset of frames for each sample?”_ One way to accomplish this would be to use a Python library for random sampling in conjunction with the `select_frames()` method. First, you can use random sampling without replacement to pick a set of frame numbers for each video. Then, you can get the frame `id` for each of these frames. Finally, you can pass this list of `id` s into `select_frames()`. Here’s one implementation using numpy’s random choice method: ```python 1from numpy.random import choice as nrc 2 3import fiftyone as fo 4import fiftyone.zoo as foz 5from fiftyone import ViewField as F 6 7dataset = foz.load_zoo_dataset("quickstart-video") 8 9## get nested list of frame ids for each sample 10frame_ids = dataset.values("frames.id") 11## number of frames for each sample 12nframes = dataset.values(F("frames").length()) 13## number of samples in dataset 14nsample = len(dataset) 15 16sample_frames = [nrc(nframe, 10, replace=False) for nframe in nframes] 17 18keep_frame_ids = [] 19for i in range(nsample): 20 curr_frame_ids = frame_ids[i] 21 for s in sample_frames[i]: 22 keep_frame_ids.append(curr_frame_ids[s]) 23 24kept_view = dataset.select_frames(keep_frame_ids) ``` If you’d like, at this point you can also convert the videos to frames: ```python 1kept_frames_view = kept_view.to_frames() ``` Learn more about [video views](https://docs.voxel51.com/user_guide/using_views.html#video-views) and [frame views](https://docs.voxel51.com/user_guide/using_views.html#frame-views) in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,300+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,500+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 241+ repositories - 55+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [filtering](https://voxel51.com/blog/tag/filtering) [GeoJSON](https://voxel51.com/blog/tag/geojson) [Labeling mistakes](https://voxel51.com/blog/tag/labeling-mistakes) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/3f54d0a45faa06a04b5d0244dd7c092603150cf0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Mar 10, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 11, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/a3a918e30b0553723b9392ea90763379f98480a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – April 7, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Apr 7, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/957bcedfac4ca50489f9c0dd31e48807c81e7b21-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – May 12, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ May 12, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-253-lllmstxt|> ## YOLOv8 Insights Part 1 [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tutorials](https://voxel51.com/blog/category/tutorials) Giving YOLOv8 a Second Look (Part 1) Feb 22, 2023 • 8 min read Article content In this article [Generate, load, and visualize YOLOv8 model predictions](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#2cd9c33f40e1) [Background on YOLO models](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#142c97a23f2a) [Getting started](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#62a5dde0229a) [Generating and loading YOLOv8 predictions](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#5d68d65eac96) [Visualizing YOLOv8 predictions](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#299f77ea4d6e) [Conclusion](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#8a9564af9d4a) [Join the FiftyOne community!](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#359b4937bcab) In this article [Generate, load, and visualize YOLOv8 model predictions](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#2cd9c33f40e1) [Background on YOLO models](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#142c97a23f2a) [Getting started](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#62a5dde0229a) [Generating and loading YOLOv8 predictions](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#5d68d65eac96) [Visualizing YOLOv8 predictions](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#299f77ea4d6e) [Conclusion](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#8a9564af9d4a) [Join the FiftyOne community!](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#359b4937bcab) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) _Editor's note – This is the first article in the three-part series:_ - _Part 1 - Generate, load, and visualize YOLOv8 model predictions (this article)_ - [_Part 2 - Evaluate YOLOv8 model predictions_](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2/) - [_Part 3 - Fine-tune YOLOv8 models for custom computer vision applications_](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3/) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## Generate, load, and visualize YOLOv8 model predictions Welcome to the first part in our three part series on [YOLOv8](https://docs.ultralytics.com/#ultralytics-yolov8)! In this series, we’ll show you how to work with YOLOv8, from downloading the off-the-shelf models, to fine-tuning these models for specific use cases, and everything in between. Throughout the series, we will be using two libraries: [FiftyOne](https://github.com/voxel51/fiftyone), the open source computer vision toolkit, and [Ultralytics](https://github.com/ultralytics/ultralytics), the library that will give us access to YOLOv8. In Part 1, you’ll learn how to generate, load, and visualize YOLOv8 predictions. In [Part 2](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2), we’ll show you how to evaluate the quality of YOLOv8 model predictions. In [Part 3](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3), we’ll conclude by walking you through the process of fine-tuning YOLOv8 for your computer vision applications. This post is organized as follows: - [YOLO family history](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#background) - [Getting set up with YOLOv8](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#getting-started) - [Generating YOLOv8 predictions](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#load-predictions) - [Visualizing YOLOv8 predictions with the FiftyOne App](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1#visualize-predictions) Continue reading to learn how you can leverage FiftyOne to take a deeper look into YOLOv8’s predictions! ## Background on YOLO models Since its [initial release back in 2015](https://arxiv.org/abs/1506.02640), the You Only Look Once (YOLO) family of computer vision models has been one of the most popular in the field. The core innovation of the YOLO architecture was to treat object detection tasks as regression problems, so that the model generates predictions for all object bounding boxes and class probabilities at the same time. This approach represented a major shift from prior state of the art object detection models, wherein each detected object presented an increased inference load - hence the You Only Look Once, or YOLO moniker. And by streamlining the prediction pipeline, YOLO drastically reduced the inference time. The original YOLO model was able to process 45 frames per second! Over the past eight years, this architectural innovation has spawned [a family of YOLO models](https://pyimagesearch.com/2022/04/04/introduction-to-the-yolo-family/), each generation bringing improvements in speed and accuracy. The models were [trained on new datasets](https://openaccess.thecvf.com/content_cvpr_2017/papers/Redmon_YOLO9000_Better_Faster_CVPR_2017_paper.pdf). The architecture was [extended to new domains](https://arxiv.org/abs/1803.06199#:~:text=We%20introduce%20Complex%2DYOLO%2C%20a,3D%20boxes%20in%20Cartesian%20space.) like bird’s eye views. And the end-to-end detector was used in more and more applications. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop In late 2022, [Ultralytics](https://github.com/ultralytics/ultralytics) announced the latest member of the YOLO family, [YOLOv8](https://docs.ultralytics.com/#ultralytics-yolov8), which comes with a new [backbone](https://arxiv.org/abs/2206.08016#:~:text=Many%20networks%20have%20been%20proposed,before%20and%20demonstrates%20its%20effectiveness.). The “model” is actually [a suite of models](https://github.com/ultralytics/ultralytics#models) for object detection and instance segmentation. The suite includes models of various sizes, from 3.2 million parameters up to 68.2 million parameters, which achieve state of the art performance and retain the speed of their progenitors. With YOLOv8, however, the basic detection and segmentation models are general purpose, which means for custom use cases they may not be suitable out of the box. _In this series, we’ll show you how to take a deeper look into YOLOv8’s predictions, and then use these insights to fine-tune the model for your own applications._ ## Getting started If you haven’t already done so, install the [Ultralytics](https://github.com/ultralytics/ultralytics) and [FiftyOne](https://github.com/voxel51/fiftyone) Python packages: ```bash 1pip install fiftyone ultralytics ``` We will also import all of the other relevant Python packages: ```python 1import numpy as np 2import os 3from tqdm import tqdm ``` Next, we’ll import the relevant modules from FiftyOne. The base FiftyOne library will allow us to efficiently work with our computer vision data. We will use the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html) to load subsets of the [MS COCO](https://cocodataset.org/#home) dataset. And the [ViewField](https://docs.voxel51.com/api/fiftyone.core.expressions.html#fiftyone.core.expressions.ViewField) will allow us to symbolically filter the data in our dataset. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F ``` Finally, we can import the `YOLO` object from Ultralytics and use this to instantiate pretrained detection and segmentation models in Python. Along with the YOLOv8 architecture, Ultralytics released a set of pretrained models, with different sizes, for classification, detection, and segmentation tasks. For the purposes of illustration, we will use the smallest version, YOLOv8 Nano (YOLOv8n), but the same syntax will work for any of the pretrained models on the [Ultralytics YOLOv8 GitHub repo](https://github.com/ultralytics/ultralytics). ```python 1from ultralytics import YOLO 2detection_model = YOLO("yolov8n.pt") 3seg_model = YOLO("yolov8n-seg.pt") ``` In this blog post series, we will call YOLOv8 models from the command line a majority of the time. However, as an illustration, we show how to use these models within a Python environment. In Python, you can apply a YOLOv8 model to an individual image by passing the file path into the model call. For an image with file path `path/to/image.jpg`, running ```python 1detection_model("path/to/image.jpg") ``` will generate a list containing a single `ultralytics.yolo.engine.results.Results` object. A similar result can be obtained if we apply the segmentation model to an image. These results contain bounding boxes, class confidence scores, and integers representing class labels. For a complete discussion of these results objects, see the Ultralytics YOLOv8 [Results API Reference](https://docs.ultralytics.com/reference/results/). If we want to run tasks on all images in a directory, then we can do so from the command line with the [YOLO Command Line Interface](https://docs.ultralytics.com/cli/) by specifying the task `[detect, segment, classify]` and mode `[train, val, predict, export]`, along with other arguments. To run inference on a set of images, we must first put the data in the appropriate format. The best way to do so is to load your images into a FiftyOne `Dataset`, and then export the dataset in [YOLOv5Dataset](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#yolov5dataset) format, as YOLOv5 and YOLOv8 use the same data formats. As an example, if all of your images are in a `my_image_dir` directory, you can load the images in using the `from_dir()` method: ```python 1dataset = fo.Dataset.from_dir( 2 dataset_dir="my_image_dir", 3 dataset_type=fo.types.ImageDirectory 4) ``` And then export the dataset into a new directory, `my_yolo_dir` in the right format, which will create the directory and populate it with an `images` subdirectory, as well as a `dataset.yaml` [YAML](https://yaml.org/) file: ```python 1dataset.export( 2 export_dir=`my_yolo_dir`, 3 dataset_type=fo.types.YOLOv5Dataset 4) ``` Shortly, we will use a slightly more general export in order to account for ground truth labels, label classes, and dataset splits. Now we can run inference with the YOLOv8n detection model: ```bash 1yolo task=detect mode=predict model=yolov8n.pt source=/my_yolo_dir/images/val save_txt=True save_conf=True ``` But how do we know if these predictions are _good_? Before we deploy a model in production, we need to understand what it is doing and what its limitations may be. So, let’s take a deeper look at what YOLOv8 is doing! ## Generating and loading YOLOv8 predictions The first step to understanding YOLOv8 is visualizing its predictions. To do this, we will use the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html). For the sake of simplicity, we will look at YOLOv8’s predictions on a subset of the MS COCO dataset. This is the dataset on which these models were trained, which means that they are likely to show close to peak performance on this data. Additionally, working with COCO data makes it easy for us to map model outputs to class labels. Let’s load the images and ground truth object detections in COCO’s validation set from the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html). ```python 1dataset = foz.load_zoo_dataset( 2 'coco-2017', 3 split='validation', 4) ``` We can also generate a mapping from YOLO class predictions to COCO class labels. [COCO has 91 classes](https://cocodataset.org/#home), and YOLOv8, just like YOLOv3 and YOLOv5, ignores all of the numeric classes and [focuses on the remaining 80](https://imageai.readthedocs.io/en/latest/detection/). ```python 1coco_classes = [c for c in dataset.default_classes if not c.isnumeric()] ``` Before adding in YOLOv8’s predictions, we can visualize this data in the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html) by launching a session: ```python 1session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/a45a0d3814229b0a524057e139041d2ad200c7fd-3426x1816.png?auto=format&dpr=2&fit=max&q=75&w=1600) Now let’s generate detection predictions for these images from the command line as in the example above. For later convenience, we use this more general `export_yolo_data()` method. ```python 1def export_yolo_data( 2 samples, 3 export_dir, 4 classes, 5 label_field = "ground_truth", 6 split = None 7 ): 8 9 if type(split) == list: 10 splits = split 11 for split in splits: 12 export_yolo_data( 13 samples, 14 export_dir, 15 classes, 16 label_field, 17 split 18 ) 19 else: 20 if split is None: 21 split_view = samples 22 split = "val" 23 else: 24 split_view = samples.match_tags(split) 25 26 split_view.export( 27 export_dir=export_dir, 28 dataset_type=fo.types.YOLOv5Dataset, 29 label_field=label_field, 30 classes=classes, 31 split=split 32 ) ``` For this data, the export and inference commands look like: ```python 1coco_val_dir = "coco_val" 2export_yolo_data(dataset, coco_val_dir, coco_classes) ``` And: ```bash 1yolo task=detect mode=predict model=yolov8n.pt source=coco_val/images/val save_txt=True save_conf=True ``` Now it’s time to add YOLOv8’s predictions for these images into our dataset. When we run a YOLOv8 inference task from the command line, the predictions are stored in a `.txt` file. For detections, these text files contain one line per object detection in the image: an integer for the class label, a class confidence score, and four values representing the bounding box. We can read a YOLOv8 detection prediction file with `N` detections into an `(N, 6)` numpy array: ```python 1def read_yolo_detections_file(filepath): 2 detections = [] 3 if not os.path.exists(filepath): 4 return np.array([]) 5 6 with open(filepath) as f: 7 lines = [line.rstrip('n').split(' ') for line in f] 8 9 for line in lines: 10 detection = [float(l) for l in line] 11 detections.append(detection) 12 return np.array(detections) ``` From here, we need to convert these detections into FiftyOne’s [`Detections`](https://docs.voxel51.com/user_guide/using_datasets.html#object-detection) format. [YOLOv8 represents bounding boxes](https://docs.ultralytics.com/reference/results/#boxes-api-reference) in a centered format with coordinates `[center_x, center_y, width, height]`, whereas [FiftyOne stores bounding boxes](https://docs.voxel51.com/user_guide/using_datasets.html#object-detection) in `[top-left-x, top-left-y, width, height]` format. We can make this conversion by “un-centering” the predicted bounding boxes: ```python 1def _uncenter_boxes(boxes): 2 '''convert from center coords to corner coords''' 3 boxes[:, 0] -= boxes[:, 2]/2. 4 boxes[:, 1] -= boxes[:, 3]/2. ``` Additionally, we can convert a list of class predictions (indices) to a list of class labels (strings) by passing in the class list: ```python 1def _get_class_labels(predicted_classes, class_list): 2 labels = (predicted_classes).astype(int) 3 labels = [class_list[l] for l in labels] 4 return labels ``` Given the output of a `read_yolo_detections_file()` call, `yolo_detections`, we can generate the FiftyOne Detections object that captures this data: ```python 1def convert_yolo_detections_to_fiftyone( 2 yolo_detections, 3 class_list 4 ): 5 6 detections = [] 7 if yolo_detections.size == 0: 8 return fo.Detections(detections=detections) 9 10 boxes = yolo_detections[:, 1:-1] 11 _uncenter_boxes(boxes) 12 13 confs = yolo_detections[:, -1] 14 labels = _get_class_labels(yolo_detections[:, 0], class_list) 15 16 for label, conf, box in zip(labels, confs, boxes): 17 detections.append( 18 fo.Detection( 19 label=label, 20 bounding_box=box.tolist(), 21 confidence=conf 22 ) 23 ) 24 return fo.Detections(detections=detections) ``` The final ingredient is a function that takes in the file path of an image, and returns the file path of the corresponding YOLOv8 detection prediction text file. ```python 1def get_prediction_filepath(filepath, run_number = 1): 2 run_num_string = "" 3 if run_number != 1: 4 run_num_string = str(run_number) 5 filename = filepath.split("/")[-1].split(".")[0] 6 return "runs/detect/predict{}/labels/".format(run_num_string) + filename + ".txt" ``` Note that if you run multiple inference calls for the same task, the predictions results are stored in a directory with the next available integer appended to `predict` in the file path. You can account for this in the above function by passing in the `run_number` argument. Putting the pieces together, we can write a function that adds these YOLOv8 detections to all of the samples in our dataset efficiently by batching the read and write operations to the underlying [MongoDB database](https://docs.voxel51.com/environments/index.html#connecting-to-a-localhost-database). ```python 1def add_yolo_detections( 2 samples, 3 prediction_field, 4 prediction_filepath, 5 class_list 6 ): 7 8 prediction_filepaths = samples.values(prediction_filepath) 9 yolo_detections = [read_yolo_detections_file(pf) for pf in prediction_filepaths] 10 detections = [convert_yolo_detections_to_fiftyone(yd, class_list) for yd in yolo_detections] 11 samples.set_values(prediction_field, detections) ``` Now we can rapidly add the detections in a few lines of code: ```python 1filepaths = dataset.values("filepath") 2prediction_filepaths = [get_prediction_filepath(fp) for fp in filepaths] 3dataset.set_values( 4 "yolov8n_det_filepath", 5 prediction_filepaths 6) 7 8add_yolo_detections( 9 dataset, 10 "yolov8n", 11 "yolov8n_det_filepath", 12 coco_classes 13) 14 ``` ## Visualizing YOLOv8 predictions The first step to understanding what a computer vision model is doing should also be to visualize the model’s predictions on your data. We can do so in the FiftyOne App: ```python 1session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/aab2c34b834a984190ba6e8e0365dc6b0873a093-3426x1816.png?auto=format&dpr=2&fit=max&q=75&w=1600) If we want to show only the predictions, we can uncheck the `ground_truth` label in the sidebar on the left. We can also view more details about the labels on an individual image by clicking on the image in the sample grid, which opens in an expanded modal. ![](https://cdn.sanity.io/images/h6toihm1/production/c18e83dd3a5518bd17a8036b7eb597b3a3faad0e-3426x1816.png?auto=format&dpr=2&fit=max&q=75&w=1600) Here we can see the ground truth and predicted class labels, as well as the class confidence probabilities. We can filter for high confidence predictions using the `confidence` slider in the sidebar on the left under the `yolov8n` label header: ![](https://cdn.sanity.io/images/h6toihm1/production/71a2e103ae9867cf3a64e94277219ec29fce63d1-3426x1816.png?auto=format&dpr=2&fit=max&q=75&w=1600) We can see that if we filter for predictions with `confidence >= 0.9`, we get only 2,008 out of the 26k+ predictions generated by running the model on the dataset. It is also worth noting that it is possible to convert YOLOv8 predictions directly from the output of a YOLO model call in Python, without first generating external prediction files and reading them in. Let’s see how this can be done for instance segmentations. Like detections, YOLOv8 stores instance segmentations with centered bounding boxes. In addition, [YOLOv8 stores a mask](https://docs.ultralytics.com/reference/results/#masks-api-reference) that covers the entire image, with only a rectangular region of that mask containing nonzero values. FiftyOne, on the other hand, [stores instance segmentations](https://docs.voxel51.com/user_guide/using_datasets.html#instance-segmentations) at `Detection` labels with a mask that only covers the given bounding box. We can convert from YOLOv8 instance segmentations to FiftyOne instance segmentations with this `convert_yolo_segmentations_to_fiftyone()` function: ```python 1def convert_yolo_segmentations_to_fiftyone( 2 yolo_segmentations, 3 class_list 4 ): 5 6 detections = [] 7 boxes = yolo_segmentations.boxes.xywhn 8 if not boxes.shape or yolo_segmentations.masks is None: 9 return fo.Detections(detections=detections) 10 11 _uncenter_boxes(boxes) 12 masks = yolo_segmentations.masks.masks 13 labels = _get_class_labels(yolo_segmentations.boxes.cls, class_list) 14 15 for label, box, mask in zip(labels, boxes, masks): 16 ## convert to absolute indices to index mask 17 w, h = mask.shape 18 tmp = np.copy(box) 19 tmp[2] += tmp[0] 20 tmp[3] += tmp[1] 21 tmp[0] *= h 22 tmp[2] *= h 23 tmp[1] *= w 24 tmp[3] *= w 25 tmp = [int(b) for b in tmp] 26 y0, x0, y1, x1 = tmp 27 sub_mask = mask[x0:x1, y0:y1] 28 29 detections.append( 30 fo.Detection( 31 label=label, 32 bounding_box = list(box), 33 mask = sub_mask.astype(bool) 34 ) 35 ) 36 return fo.Detections(detections=detections) ``` Looping through all samples in the dataset, we can add the predictions from our `seg_model`, and then view these predicted masks in the FiftyOne App. ![](https://cdn.sanity.io/images/h6toihm1/production/6f658930c64ff32c506f11676b6bb65cb396495f-2560x1357.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Conclusion In this article, we demonstrated how to start using YOLOv8 models on your data, and how to visualize YOLOv8 model predictions on images. In [Part 2](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2), using the same COCO validation images and labels, we’ll show you how to evaluate the quality of a YOLOv8 model’s predictions, identify edge cases, and assess potential modes of failure. Continue to [Part 2](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2)! ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,350+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,500+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 250+ repositories - 55+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Instance segmentation](https://voxel51.com/blog/tag/instance-segmentation) [MS COCO](https://voxel51.com/blog/tag/ms-coco) [object detection](https://voxel51.com/blog/tag/object-detection) [Ultralytics](https://voxel51.com/blog/tag/ultralytics) [YOLO](https://voxel51.com/blog/tag/yolo) [YOLOv8](https://voxel51.com/blog/tag/yolov8) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/de34ad70ef11fe3d7daf5b3c8ff6214e76c85522-2560x1440.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 2)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2) [![](https://cdn.sanity.io/images/h6toihm1/production/803cb935ddffbb5b29b6d3c73104b3a1221ddbfc-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 3)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3) [![](https://cdn.sanity.io/images/h6toihm1/production/047b21a97f6c858334f9f35ed89fa7655ebf5767-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ State-of-the-Art Object Detection with YOLO-NAS & FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ May 4, 2023](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-254-lllmstxt|> ## Evaluating YOLOv8 Predictions [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tutorials](https://voxel51.com/blog/category/tutorials) Giving YOLOv8 a Second Look (Part 2) Feb 22, 2023 • 5 min read Article content In this article [Evaluate YOLOv8 model predictions](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#483e5da7718f) [Part 1 recap](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#ad242f8a5b3a) [Printing YOLOv8 model performance metrics](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#d19df562ee52) [Viewing concerning classes](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#a887b655136c) [Finding poorly performing samples](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#5e42c8ec1888) [Conclusion](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#2553e1ab06cb) [Join the FiftyOne community!](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#e3796f75b148) In this article [Evaluate YOLOv8 model predictions](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#483e5da7718f) [Part 1 recap](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#ad242f8a5b3a) [Printing YOLOv8 model performance metrics](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#d19df562ee52) [Viewing concerning classes](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#a887b655136c) [Finding poorly performing samples](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#5e42c8ec1888) [Conclusion](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#2553e1ab06cb) [Join the FiftyOne community!](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#e3796f75b148) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) _Editor 's note – This is the second post in the three-part series:_ - [_Part 1 - Generate, load, and visualize YOLOv8 model predictions_](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1/) - _Part 2 - Evaluate YOLOv8 model predictions (this article)_ - [_Part 3 - Fine-tune YOLOv8 models for custom computer vision applications_](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3/) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## Evaluate YOLOv8 model predictions Welcome to the second part in our three part series on [YOLOv8](https://docs.ultralytics.com/#ultralytics-yolov8)! In this series, we’ll show you how to work with YOLOv8, from downloading the off-the-shelf models, to fine-tuning these models for specific use cases, and everything in between. Throughout the series, we will be using two libraries: [FiftyOne](https://github.com/voxel51/fiftyone), the open source computer vision toolkit, and [Ultralytics](https://github.com/ultralytics/ultralytics), the library that will give us access to YOLOv8. In [Part 1](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1), we generated, loaded, and visualized YOLOv8 model predictions. Here in Part 2, we’ll delve deeper into evaluating the quality of the YOLOv8n detection model’s predictions, from one-number metrics to class-wise performance and identifying edge cases. This post is organized as follows: - [Part 1 recap](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#part1) - [Printing performance metrics](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#printing-metrics) - [Viewing concerning classes](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#view-concerning-classes) - [Finding poorly performing samples](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2#poor-samples) Continue reading to learn how you can leverage FiftyOne to take a deeper look into YOLOv8’s predictions! ## Part 1 recap In [Part 1](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1), after importing the necessary modules, ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F ``` we loaded the validation split of the [COCO 2017 dataset](https://cocodataset.org/#home) into FiftyOne with ground truth object detections. We then generated predictions with the YOLOv8n detection model from the [Ultralytics YOLOv8 GitHub](https://github.com/ultralytics/ultralytics) and added them to our dataset in the `yolov8n` label field on our samples. When we left off, we had launched the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html) to visualize our images and predictions. ![](https://cdn.sanity.io/images/h6toihm1/production/aab2c34b834a984190ba6e8e0365dc6b0873a093-3426x1816.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Printing YOLOv8 model performance metrics Now that we have YOLOv8 predictions loaded onto the images in our dataset from [Part 1](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1), we can evaluate the quality of these predictions using FiftyOne’s [Evaluation API](https://docs.voxel51.com/user_guide/evaluation.html). To evaluate the object detections in the `yolov8_det` field relative to the `ground_truth` detections field, we can run: ```python 1detection_results = dataset.evaluate_detections( 2 "yolov8n", 3 eval_key="eval", 4 compute_mAP=True, 5 gt_field="ground_truth", 6) ``` We can then get the [mean average precision](https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173) (mAP) of the model’s predictions: ```python 1mAP = detection_results.mAP() 2print("mAP = {}".format(mAP)) ``` mAP = 0.3121319189417518 We can also look at the model’s performance on the 20 most common object classes in the dataset, where it has seen the most examples so the statistics are most meaningful: ```python 1counts = dataset.count_values("ground_truth.detections.label") 2 3top20_classes = sorted( 4 counts, 5 key=counts.get, 6 reverse=True 7)[:20] 8 9detection_results.print_report(classes=top20_classes) 10 ``` precision recall f1-score support person 0.85 0.68 0.76 11573 car 0.71 0.52 0.60 1971 chair 0.62 0.34 0.44 1806 book 0.61 0.12 0.20 1182 bottle 0.68 0.39 0.50 1051 cup 0.61 0.44 0.51 907 dining table 0.54 0.42 0.47 697 traffic light 0.66 0.36 0.46 638 bowl 0.63 0.49 0.55 636 handbag 0.48 0.12 0.19 540 bird 0.79 0.39 0.52 451 boat 0.58 0.29 0.39 430 truck 0.57 0.35 0.44 415 bench 0.58 0.27 0.37 413 umbrella 0.65 0.52 0.58 423 cow 0.81 0.61 0.70 397 banana 0.68 0.34 0.45 397 carrot 0.56 0.29 0.38 384 motorcycle 0.77 0.58 0.66 379 backpack 0.51 0.16 0.24 371 micro avg 0.76 0.52 0.61 25061 macro avg 0.64 0.38 0.47 25061 weighted avg 0.74 0.52 0.60 25061 ## Viewing concerning classes From the output of the `print_report()` call above, we can see that this model performs decently well, but certainly has its limitations. While its precision is relatively good on average, it is lacking when it comes to recall. This is especially pronounced for certain classes like the `book` class. Fortunately, we can dig deeper into these results with FiftyOne. Using the FiftyOne App, we can for instance filter by class for both ground truth and predicted detections so that only `book` detections appear in the samples. ![](https://cdn.sanity.io/images/h6toihm1/production/20bc2c89f7d0f3f9498abc542dcb98c62f8320f1-3426x1816.png?auto=format&dpr=2&fit=max&q=75&w=1600) Scrolling through the samples in the sample grid, we can see that a lot of the time, COCO’s purported _ground truth_ labels for the `book` class appear to be imperfect. Sometimes, individual books are bounded, other times rows or whole bookshelves are encompassed in a single box, and yet other times books are entirely unlabeled. Unless our desired computer vision application specifically requires good `book` detection, this should probably not be a point of concern when we are assessing the quality of the model. After all, the quality of a model is limited by the quality of the data it is trained on - this is why data-centric approaches to computer vision are so important! For other classes like the `bird` class, however, there appear to be challenges. One way to see this is to filter for `bird` ground truth detections and then convert to an [`EvaluationPatchesView`](https://docs.voxel51.com/api/fiftyone.core.patches.html#fiftyone.core.patches.EvaluationPatchesView). Some of these recall errors appear to be related to small features, where the resolution is poor. In other cases though, quick inspection confirms that the object is clearly a bird. This means that there is likely room for improvement. ![](https://cdn.sanity.io/images/h6toihm1/production/98de7c46717d4f999d87ea8fd2371c13e533a84f-3426x1816.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Finding poorly performing samples Beyond summary statistics, visualizing model predictions in FiftyOne allows for deeper exploration, such as looking at samples that contain the highest number of false negative detections: ```python 1fn_view = dataset.sort_by("det_fn", reverse=True) 2session = fo.launch_app(fn_view) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/e1aab68da71df2e6308085969e2f5e8529990d75-3426x1602.png?auto=format&dpr=2&fit=max&q=75&w=1600) Looking at these samples, it is immediately evident that these samples are crowded, so there are a lot of detections to possibly be missed in prediction. Instead, it might be more useful to sort by something like precision, which does not have the same dependence on the total number of objects in the image. We can compute the precision by using the number of true positives and false positives in a given image: ```python 1non_empty_view = dataset.match(F("eval_tp") > 0) 2 3## get true and false positive counts by image 4tp_vals = np.array(non_empty_view.values("eval_tp")) 5fp_vals = np.array(non_empty_view.values("eval_fp")) 6## compute precision by image 7precision_vals = tp_vals/(tp_vals+fp_vals) 8 9## set precision values 10non_empty_view.set_values("precision", precision_vals) 11 12## get lowest precision images 13low_precision_view = non_empty_view.sort_by("precision") ``` ![](https://cdn.sanity.io/images/h6toihm1/production/0fb5dc4d872931e6abe77098b9f7a7619f354a5b-3422x1808.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Conclusion In this article, we demonstrated how to evaluate a YOLOv8 model’s performance on your data. One-number metrics like mAP provide a good starting point for evaluating model quality, but fail to give a complete picture of the model’s effectiveness on a class by class or image by image basis. By looking at individual samples, we can build an intuition for _why_ a model fails or succeeds. In [Part 3](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3), we’ll push the story forward and fine-tune our YOLOv8 detection model to detect birds. Finish with [Part 3](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3)! ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,350+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,550+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,200+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 245+ repositories - 56+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [mean average precision](https://voxel51.com/blog/tag/mean-average-precision) [model evaluation](https://voxel51.com/blog/tag/model-evaluation) [MS COCO](https://voxel51.com/blog/tag/ms-coco) [Ultralytics](https://voxel51.com/blog/tag/ultralytics) [YOLO](https://voxel51.com/blog/tag/yolo) [YOLOv8](https://voxel51.com/blog/tag/yolov8) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/803cb935ddffbb5b29b6d3c73104b3a1221ddbfc-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 3)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3) [![](https://cdn.sanity.io/images/h6toihm1/production/98e839c6e81bb9c4fa3ad96bf0d5d1b77ee11f6c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 1)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1) [![](https://cdn.sanity.io/images/h6toihm1/production/047b21a97f6c858334f9f35ed89fa7655ebf5767-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ State-of-the-Art Object Detection with YOLO-NAS & FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ May 4, 2023](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-255-lllmstxt|> ## Fine-tuning YOLOv8 Models [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tutorials](https://voxel51.com/blog/category/tutorials) Giving YOLOv8 a Second Look (Part 3) Feb 22, 2023 • 10 min read Article content In this article [Fine-tune YOLOv8 models for custom computer vision applications](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#930c9811a9da) [Parts 1 and 2 recap](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#280707b6d877) [Detecting birds with YOLOv8](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#ce41bcfb0ee0) [Choosing the training data for your YOLOv8 model](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#54f045563958) [Fine-tuning YOLOv8 for a custom use case](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#bb87c579f436) [Assessing YOLOv8 model performance improvement](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#5733764e11ef) [Conclusion](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#c4151096ae43) [Join the FiftyOne community!](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#b8f6d7732850) In this article [Fine-tune YOLOv8 models for custom computer vision applications](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#930c9811a9da) [Parts 1 and 2 recap](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#280707b6d877) [Detecting birds with YOLOv8](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#ce41bcfb0ee0) [Choosing the training data for your YOLOv8 model](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#54f045563958) [Fine-tuning YOLOv8 for a custom use case](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#bb87c579f436) [Assessing YOLOv8 model performance improvement](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#5733764e11ef) [Conclusion](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#c4151096ae43) [Join the FiftyOne community!](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#b8f6d7732850) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) _Editor’s note – This is the third article in the three-part series:_ - [_Part 1 – Generate, load, and visualize YOLOv8 model predictions_](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1/) - [_Part 2 – Evaluate YOLOv8 model predictions_](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2/) - _Part 3 – Fine-tune YOLOv8 models for custom computer vision applications (this article)_ \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## Fine-tune YOLOv8 models for custom computer vision applications Welcome to the third and final installment in our three part series on [YOLOv8](https://docs.ultralytics.com/#ultralytics-yolov8)! In this series, we’ll show you how to work with YOLOv8, from downloading the off-the-shelf models, to fine-tuning these models for specific use cases, and everything in between. Throughout the series, we will be using two libraries: [FiftyOne](https://github.com/voxel51/fiftyone), the open source computer vision toolkit, and [Ultralytics](https://github.com/ultralytics/ultralytics), the library that will give us access to YOLOv8. Here in Part 3, we’ll demonstrate how to fine-tune a YOLOv8 model for your specific use case. This post is organized as follows: - [Parts 1 and 2 recap](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#recap) - [Defining our use case](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#use-case) - [Choosing training data](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#training-data) - [Fine-tuning the YOLOv8 model](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#fine-tuning) - [Assessing the improvement](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3#assessing-improvement) Continue reading to learn how you can incorporate YOLOv8 models into your computer vision workflows! ## Parts 1 and 2 recap In [Part 1](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1), we loaded the validation split of the [COCO 2017 dataset](https://cocodataset.org/#home), into FiftyOne with ground truth object detections. We then generated predictions with the YOLOv8n detection model from the [Ultralytics YOLOv8 Github](https://github.com/ultralytics/ultralytics) and added them to our dataset in the `yolov8n` label field on our samples. In [Part 2](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2), we used FiftyOne’s [Evaluation API](https://docs.voxel51.com/user_guide/evaluation.html) to evaluate the quality of these detections. We found that this base model had relatively good precision, but struggled with recall for certain classes. When we dug into the details, we saw that for certain classes like the `book` class, imperfections in the COCO ground truth labels were at least part of the problem, but for other classes like `bird`, the model could benefit from fine-tuning on a broader set of examples. ![](https://cdn.sanity.io/images/h6toihm1/production/98de7c46717d4f999d87ea8fd2371c13e533a84f-3426x1816.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Detecting birds with YOLOv8 precision recall f1-score support person 0.85 0.68 0.76 11573 car 0.71 0.52 0.60 1971 chair 0.62 0.34 0.44 1806 book 0.61 0.12 0.20 1182 bottle 0.68 0.39 0.50 1051 cup 0.61 0.44 0.51 907 dining table 0.54 0.42 0.47 697 traffic light 0.66 0.36 0.46 638 bowl 0.63 0.49 0.55 636 handbag 0.48 0.12 0.19 540 bird 0.79 0.39 0.52 451 boat 0.58 0.29 0.39 430 truck 0.57 0.35 0.44 415 bench 0.58 0.27 0.37 413 umbrella 0.65 0.52 0.58 423 cow 0.81 0.61 0.70 397 banana 0.68 0.34 0.45 397 carrot 0.56 0.29 0.38 384 motorcycle 0.77 0.58 0.66 379 backpack 0.51 0.16 0.24 371 micro avg 0.76 0.52 0.61 25061 macro avg 0.64 0.38 0.47 25061 weighted avg 0.74 0.52 0.60 25061 As we saw in the previous section, while YOLOv8 has decent performance out of the box, it may not be suitable for specific use cases without some modification. For the `bird` class, for example, the base YOLOv8n model only achieved 39% recall. Suppose you’re working for a bird conservancy group, putting computer vision models in the field to track and protect endangered species. Your goal is to detect, in real time, as many birds as possible. Given its inference speed, the YOLOv8 architecture seems like the obvious choice. However, you are not satisfied with the 39% recall in the evaluation report on the COCO validation data. Because this is a high-stakes application, you want to squeeze every last bit of performance you can out of this real-time detection architecture. If you wanted to, you could train a new YOLOv8 detection model from scratch, as illustrated in the [YOLOv8 Quickstart guide](https://docs.ultralytics.com/quickstart/#use-with-python), but ideally you would like to leverage the pretrained model’s existing knowledge. Fortunately, it is pretty straightforward to fine-tune an existing YOLOv8 model. Before continuing, let’s pare down our task. The flexible query language built into FiftyOne makes it easy to slice and dice your datasets to find interesting views in just a line of code or the click of a button in the App. At this point, we have a FiftyOne [`Dataset`](https://docs.voxel51.com/user_guide/using_datasets.html) with our COCO validation images, ground truth detections, and YOLOv8n predictions in a `yolov8n` label field on each sample. Given that in our use case we are only concerned with detecting birds, let’s create a test set by filtering out all non- `bird` ground truth detections using [`filter_labels()`](https://docs.voxel51.com/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.filter_labels). We will also filter out the non- `bird` predictions, but will pass the `only_matches = False` argument into `filter_labels()` to make sure we keep images that have ground truth `bird` detections without YOLOv8n `bird` predictions. ```python 1test_dataset = dataset.filter_labels( 2 "ground_truth", 3 F("label") == "bird" 4).filter_labels( 5 "yolov8n", 6 F("label") == "bird", 7 only_matches=False 8).clone() 9 10test_dataset.name = "birds-test-dataset" 11test_dataset.persistent = True 12 13## set classes to just include birds 14classes = ["bird"] ``` We then give the dataset a name, make it persistent, and save it to the underlying database. This test set has only 125 images, which we can visualize in the [FiftyOne App](https://voxel51.com/docs/fiftyone/user_guide/app.html). ![](https://cdn.sanity.io/images/h6toihm1/production/4c28e3c5615a01c4166aff9f3409079e91d7e77c-3422x1808.png?auto=format&dpr=2&fit=max&q=75&w=1600) We can also run [`evaluate_detections()`](https://docs.voxel51.com/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.evaluate_detections) on this data to evaluate the YOLOv8n model’s performance on images with ground truth bird detections. We will store the results under the `base` evaluation key: ```python 1base_bird_results = test_dataset.evaluate_detections( 2 "yolov8n", 3 eval_key="base", 4 compute_mAP=True, 5) 6 7print(base_bird_results.mAP()) 8## 0.24897924786479841 9 10base_bird_results.print_report(classes=classes) ``` precision recall f1-score support bird 0.87 0.39 0.54 451 We note that while the recall is the same as in the initial evaluation report over the entire COCO validation split, the precision is higher. This means there are images that have YOLOv8n `bird` predictions but _not_ ground truth `bird` detections. The final step in preparing this test set is exporting the data into YOLOv8 format so we can run inference on just these samples with our fine-tuned model when we are done training. We will do so using the `export_yolo_data()` function we defined in [Part 1](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1). ```python 1export_yolo_data( 2 test_dataset, 3 "birds_test", 4 classes 5) ``` ## Choosing the training data for your YOLOv8 model The most important component in fine-tuning a model is the data on which the model is trained. If we want our model to exhibit high performance on a specific subset of data, then our goal should be to generate a high-quality training dataset whose examples cover all expected scenarios in that subset. This is both an art and a science. It can involve pulling in data from other datasets, annotating more data that you’ve already collected with ground truth labels, augmenting your data with tools like [Albumentations](https://albumentations.ai/), or generating synthetic data with [diffusion models](https://blog.roboflow.com/synthetic-data-with-stable-diffusion-a-guide/) or [GANs](https://towardsai.net/p/l/gans-for-synthetic-data-generation). In this article, we’ll take the first approach and incorporate existing high-quality data from Google’s [Open Images dataset](https://storage.googleapis.com/openimages/web/index.html). For a thorough tutorial on how to work with Open Images data, see [Loading Open Images V6 and custom datasets with FiftyOne](https://medium.com/voxel51/loading-open-images-v6-and-custom-datasets-with-fiftyone-18b5334851c3). The COCO training data on which YOLOv8 was trained contains 3237 images with `bird` detections. Open Images is more expansive, with the train, test, and validation splits together housing 20k+ images with `Bird` detections. Let’s create our training dataset. First, we’ll create a dataset, `train_dataset`, by loading the `bird` detection labels from the COCO train split using the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html), and cloning this into a new `Dataset` object: ```python 1train_dataset = foz.load_zoo_dataset( 2 'coco-2017', 3 split='train', 4 classes=classes 5).clone() 6train_dataset.name = "birds-train-data" 7train_dataset.persistent = True 8train_dataset.save() ``` Then, we’ll load Open Images samples with `Bird` detection labels, passing in `only_matching=True` to only load the `Bird` labels. We then map these labels into COCO label format by changing `Bird` into `bird`. ```python 1oi_samples = foz.load_zoo_dataset( 2 "open-images-v6", 3 classes = ["Bird"], 4 only_matching=True, 5 label_types="detections" 6).map_labels( 7 "ground_truth", 8 {"Bird":"bird"} 9) ``` We can add these new samples into our training dataset with `merge_samples()`: ```python 1train_dataset.merge_samples(oi_samples) ``` This dataset contains 24,226 samples with `bird` labels, or more than seven times as many birds as the base YOLOv8n model was trained on. In the next section, we’ll demonstrate how to fine-tune the model on this data using the [YOLO Trainer class](https://docs.ultralytics.com/reference/base_trainer/). ## Fine-tuning YOLOv8 for a custom use case The final step in preparing our data is splitting it into training and validation sets and exporting it into YOLO format. We will use an 80-20 train-val split, which we will select randomly using [FiftyOne’s random utils](https://docs.voxel51.com/api/fiftyone.utils.random.html). ```python 1import fiftyone.utils.random as four 2 3## delete existing tags to start fresh 4train_dataset.untag_samples(train_dataset.distinct("tags")) 5 6## split into train and val 7four.random_split( 8 train_dataset, 9 {"train": 0.8, "val": 0.2} 10) 11 12## export in YOLO format 13export_yolo_data( 14 train_dataset, 15 "birds_train", 16 classes, 17 split = ["train", "val"] 18) ``` Now all that is left is to do the fine-tuning! We will use the same [YOLO command line syntax](https://docs.ultralytics.com/cli/), but instead of setting `mode=predict`, we will set `mode=train`. We will specify the initial weights as the starting point for training, the number of epochs, image size, and batch size. ```bash 1yolo task=detect mode=train model=yolov8n.pt data=birds_train/dataset.yaml epochs=100 imgsz=640 batch=16 ``` For my fine-tuning, I used an [NVIDIA TITAN GPU](https://www.nvidia.com/en-us/titan/titan-v/). I set the training to run for 100 epochs, but it had basically converged after 60 epochs, so I stopped there. You may find that your specific use case requires fewer or potentially more iterations to reach your desired performance. Image sizes 640 train, 640 val Using 8 dataloader workers Logging results to runs/detect/train Starting training for 100 epochs... Epoch GPU\_mem box\_loss cls\_loss dfl\_loss Instances Size 1/100 6.65G 1.392 1.627 1.345 22 640: 1 Class Images Instances Box(P R mAP50 m all 4845 12487 0.677 0.524 0.581 0.339 Epoch GPU\_mem box\_loss cls\_loss dfl\_loss Instances Size 2/100 9.58G 1.446 1.407 1.395 30 640: 1 Class Images Instances Box(P R mAP50 m all 4845 12487 0.669 0.47 0.54 0.316 Epoch GPU\_mem box\_loss cls\_loss dfl\_loss Instances Size 3/100 9.58G 1.54 1.493 1.462 29 640: 1 Class Images Instances Box(P R mAP50 m all 4845 12487 0.529 0.329 0.349 0.188 ...... Epoch GPU\_mem box\_loss cls\_loss dfl\_loss Instances Size 58/100 9.59G 1.263 0.9489 1.277 47 640: 1 Class Images Instances Box(P R mAP50 m all 4845 12487 0.751 0.631 0.708 0.446 Epoch GPU\_mem box\_loss cls\_loss dfl\_loss Instances Size 59/100 9.59G 1.264 0.9476 1.277 29 640: 1 Class Images Instances Box(P R mAP50 m all 4845 12487 0.752 0.631 0.708 0.446 Epoch GPU\_mem box\_loss cls\_loss dfl\_loss Instances Size 60/100 9.59G 1.257 0.9456 1.274 41 640: 1 Class Images Instances Box(P R mAP50 m all 4845 12487 0.752 0.631 0.709 0.446 With fine-tuning complete, we can generate predictions on our test data with the “best” weights found during the training process, which are stored at `runs/detect/train/weights/best.pt`: ```bash 1yolo task=detect mode=predict model=runs/detect/train/weights/best.pt source=birds_test/images/val save_txt=True save_conf=True ``` And load these predictions onto our data and visualize the predictions in the FiftyOne App: ```python 1filepaths = test_dataset.values("filepath") 2prediction_filepaths = [get_prediction_filepath(fp, run_number=2) for fp in filepaths] 3test_dataset.set_values( 4 "yolov8n_bird_det_filepath", 5 prediction_filepaths 6) 7 8add_yolo_detections( 9 birds_test_dataset, 10 "yolov8n_bird", 11 "yolov8n_bird_det_filepath", 12 classes 13) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/c40274aa74722b289ff92b5db0623552bdfa7085-3422x1808.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Assessing YOLOv8 model performance improvement On a holistic level, we can compare the performance of the fine-tuned model to the original, pretrained model by stacking their standard metrics against each other. The easiest way to get these metrics is with FiftyOne’s Evaluation API: ```python 1finetune_bird_results = test_dataset.evaluate_detections( 2 "yolov8n_bird", 3 eval_key="finetune", 4 compute_mAP=True, 5) ``` From this, we can immediately see improvement in the mean average precision (mAP): ```python 1print("yolov8n mAP: {}.format(base_bird_results.mAP())) 2print("fine-tuned mAP: {}.format(finetune_bird_results.mAP())) ``` yolov8n mAP: 0.24897924786479841 fine-tuned mAP: 0.31339033693212076 Printing out a report, we can see that the recall has improved from 0.39 to 0.56. This major improvement offsets a minor dip in precision, giving an overall higher F1 score (0.67 compared to 0.54). precision recall f1-score support bird 0.81 0.56 0.67 506 ![](https://cdn.sanity.io/images/h6toihm1/production/7de7d44f384ebce6d49ad24be575b7b6c8161556-3396x1764.png?auto=format&dpr=2&fit=max&q=75&w=1600) We can also look more closely at individual images to see where the fine-tuned model is having trouble. When we do so, we can see that the model struggles to correctly handle small features. This is true for both false positives and false negatives. ![](https://cdn.sanity.io/images/h6toihm1/production/7c3d61454f88b1735276ab9813bb38b8937b9d1e-3422x1808.png?auto=format&dpr=2&fit=max&q=75&w=1600) This poor performance could be in part due to quality of the data, as many of these features are grainy. It could also be due to the training parameters, as both the pretraining and fine-tuning for this model used an image size of 640 pixels, which might not allow for fine-grained details to be captured. To further improve the model’s performance, we could try a variety of approaches, including: - Using image augmentation to increase the proportion of images with small birds - Gathering and annotating more images with small birds - Increasing the image size during fine-tuning ## Conclusion The fine-tuning presented in the previous section is only for the purpose of illustration. While YOLOv8 represents a step forward for real-time object detection and segmentation models, out-of-the-box it’s aimed at general purpose uses. Before deploying the model, it is essential to understand how it performs on your data. Only then can you effectively fine-tune the YOLOv8 architecture to suit your specific needs. In this series, we have shown you how you can use FiftyOne to visualize, evaluate, and better understand YOLOv8 model predictions. While YOLO may only look once, a conscientious computer vision engineer or researcher certainly looks twice (or more)! ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,350+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,550+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,200+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 245+ repositories - 56+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Albumentations](https://voxel51.com/blog/tag/albumentations) [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [Diffusion models](https://voxel51.com/blog/tag/diffusion-models) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [fine-tune models](https://voxel51.com/blog/tag/fine-tune-models) [GANs](https://voxel51.com/blog/tag/gans) [model evaluation](https://voxel51.com/blog/tag/model-evaluation) [MS COCO](https://voxel51.com/blog/tag/ms-coco) [Open Images](https://voxel51.com/blog/tag/open-images) [Ultralytics](https://voxel51.com/blog/tag/ultralytics) [YOLO](https://voxel51.com/blog/tag/yolo) [YOLOv8](https://voxel51.com/blog/tag/yolov8) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/de34ad70ef11fe3d7daf5b3c8ff6214e76c85522-2560x1440.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 2)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2) [![](https://cdn.sanity.io/images/h6toihm1/production/98e839c6e81bb9c4fa3ad96bf0d5d1b77ee11f6c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 1)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1) [![](https://cdn.sanity.io/images/h6toihm1/production/9fafae668ef4cd6175d5009f998d07c7bddf9719-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Better Way to Visualize 3D Point Clouds and Work with OpenAI’s Point-E\\ \\ Tutorials\\ \\ • \\ \\ Mar 22, 2023](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-256-lllmstxt|> ## February 2023 Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup – February 2023 Feb 14, 2023 • 13 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#4d46ff59f9c8) [Computer Vision Meetup Recap at a Glance](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#8e5ca2c200da) [Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#41edbdd9b155) [Q&A Recap](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#d27c636556a4) [Additional Resources](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#d06e037309b8) [Understanding Speech Recognition Using OpenAI's Whisper Model](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#fa8fbca9861c) [Q&A Recap](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#d2516c187dad) [Additional Resources](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#959025962ed8) [Computer Vision Meetup Locations](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#51025a27efe5) [Upcoming Computer Vision Meetup Speakers & Schedule](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#05bb0d61c020) [Get Involved!](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#5b2520382e9b) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#4d46ff59f9c8) [Computer Vision Meetup Recap at a Glance](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#8e5ca2c200da) [Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#41edbdd9b155) [Q&A Recap](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#d27c636556a4) [Additional Resources](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#d06e037309b8) [Understanding Speech Recognition Using OpenAI's Whisper Model](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#fa8fbca9861c) [Q&A Recap](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#d2516c187dad) [Additional Resources](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#959025962ed8) [Computer Vision Meetup Locations](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#51025a27efe5) [Upcoming Computer Vision Meetup Speakers & Schedule](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#05bb0d61c020) [Get Involved!](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#5b2520382e9b) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Last week Voxel51 hosted the February 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/). In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes by an overwhelming majority this month was Direct Relief. We are sending this month’s charitable donation of $200 to Direct Relief’s [Turkey-Syria Earthquake Relief](https://www.directrelief.org/emergency/turkey-syria-earthquake/) program on behalf of the computer vision community. ![](https://cdn.sanity.io/images/h6toihm1/production/b840f86383550954bb4a0be48a61f4bad9bbf801-1024x193.png?auto=format&dpr=2&fit=max&q=75&w=1024) ## Computer Vision Meetup Recap at a Glance #### Paula Ramos // Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO - [Video replay](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#edgeai-video) - [Presentation recap](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#edgeai-summary) - [Q&A recap](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#edgeai-qa) - [Additional resources](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#edgeai-resources) #### Vishal Rajput // Understanding Speech Recognition Using OpenAI's Whisper Model - [Video replay](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#speechrecognition-video) - [Presentation recap](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#speechrecognition-summary) - [Q&A recap](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#speechrecognition-qa) - [Additional resources](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#speechrecognition-resources) #### Next steps - [Computer Vision Meetup Locations](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#locations) - [Computer Vision Meetup Speakers — March, April, and May](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#schedule) - [Get Involved!](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap#get-involved) ## Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO ### Video Replay https://www.youtube.com/watch?v=9TzdGz1NXac ### Executive Summary One of the biggest problems in computer vision is data. As [Paula Ramos](https://www.linkedin.com/in/paula-ramos-41097319/), Computer Vision, AI and IoT Evangelist at Intel, notes, “good datasets make good models; bad datasets will affect the model’s performance and accuracy, and could result in frustration for you.” Getting quality data is a common challenge in AI use cases across industries. For example, [Eigen Innovations](https://www.intel.com/content/www/us/en/partner/showcase/offering/a5b3b0000004fKSAAY/eigen-machine-vision.html), an Intel partner, helps manufacturers prevent quality issues by accurately detecting defects. But in real world scenarios there can be an imbalance of data where there are not enough samples of defects in order to train an accurate model. So what can we do? It’s real world datasets like this that motivated the creation of Anomalib. #### An Introduction to Anomalib [Anomalib](https://github.com/openvinotoolkit/anomalib), part of the OpenVINO toolkit, is a library for unsupervised anomaly detection from data collection to deployment. As part of a tutorial Paula prepared for CVPR last year, she checked for defects in a production system. To train the model, she made use of Anomalib and was able to train her model with just 10 total images, none of which were samples of defects. How does this magic happen? Paula explores the main components of Anomalib so we can learn more. #### Exploring Anomalib: Algorithms Anomalib includes state-of-the-art anomaly detection algorithms across four main categories: knowledge based models, clustering models, reconstruction based models, and probabilistic models. Choose a model depending on your use case. ![](https://cdn.sanity.io/images/h6toihm1/production/982e5387f858b6a44c4c9ad78913921c163a3d23-1824x754.png?auto=format&dpr=2&fit=max&q=75&w=1600) #### Exploring Anomalib: Modules, Tools, and Tests Anomalib includes modules for data, pre-processing, models, and post-processing, and deployment. It also includes tools and tests. Paula describes each of these seven areas. ![](https://cdn.sanity.io/images/h6toihm1/production/b119756e60e53bb1c51a14bc98b355d76f6583d1-1880x780.png?auto=format&dpr=2&fit=max&q=75&w=1600) Anomalib’s **data** component provides dataset adapters for a growing number of public benchmark datasets in both the image and video domains. Custom datasets are also supported. **Pre-processing** applies transformations to input images before training and optionally divides the image into overlapping or non-overlapping tiles. Paula shares a common use case for image tiling in real world datasets: high resolution images that include anomalies in a relatively small pixel area. Scenarios like these can be challenging for deep learning model to handle, so to address this issue, Anomalib can tie images to patches to support high resolution image training. Anomalib’s **model** component contains a selection of state-of-the art anomaly detection and localization algorithms, as well as a set of modular components that serve as building blocks to compose custom algorithms. Anomalib also offers **post-processing** features for normalization, thresholding, and visualization outputs. Anomalib’s **deployment** options include using Torch, ONNX, Gradio, or OpenVINO. Within the Anomalib library there are **tools** that include entry points for training, testing, inferring, benchmarking, and hyper parameter optimization. Additionally, the Anomalib library constantly undergoes unit integration and regression **tests** to capture any potential defects. #### Getting Started with Anomalib Paula shows how easy it is to get started with Anomalib, including a demo: - Create an environment to run Anomalib (Python version 3.8) - Clone the Anomalib repo and install it locally (with OpenVINO requirements) - Install Jupyter Notebooks and ipywidgets - Download the MVTec-AD dataset needed to run the demonstration before you follow along with the demo in the notebook - Head to the getting\_started Anomalib notebook ( [available on GitHub](https://github.com/openvinotoolkit/anomalib/tree/main/notebooks)) to get started and run the demo to see it in action for yourself The demo shows how easy it is to install Anomalib and the other packages you need, choose a model, update the config file, start model training, visualize results, and perform inference. Now it’s your turn to try! Plus, giving Anomalib a whirl could win you one of five limited edition hoodies if you follow these steps before February 16, 2023. ![](https://cdn.sanity.io/images/h6toihm1/production/20c6d8ab849546f2769d3eacaca35d00cc8fa5ad-1774x1044.png?auto=format&dpr=2&fit=max&q=75&w=1600) #### Other Exciting Tools: OpenVINO and Intel Geti Anomalib is part of the OpenVINO ecosystem. If you're interested in using the OpenVINO toolkit in your deployments, Paula invites you to check out the examples in [the OpenVINO tutorial notebooks](https://docs.openvino.ai/latest/tutorials.html). More than 60 demos are there, including object detection, pose estimation, human action recognition, style transfer, text spotting, OCR, stable diffusion, YOLOv8, and many more. In her presentation, Paula also introduces another exciting tool for computer vision: [the Intel® Geti™](https://geti.intel.com/) platform. Intel Geti enables users to build and optimize computer vision models by abstracting away technical complexity with an intuitive interface. Learn more about Intel Geti in less than 3 minutes [in this video](https://www.linkedin.com/feed/update/urn:li:activity:6981300515385085952/) showcasing how the computer vision AI platform is used to help optimize coffee production. ## Q&A Recap Here’s a recap of the live Q&A from this presentation during the virtual Computer Vision Meetup: **After anomaly detection, can we use DC-GANs to remove the anomaly in the image or solve the data imbalance problem?** Using a DC-GAN it may be possible to remove the anomaly, but it would not be able to solve the imbalance problem 100%. There is another way to try to balance the data; in the Anomalib library, we have the option to create scientific abnormalities, but some imbalancing would still exist. **Can OpenVINO speed up deep learning models that are deployed to CPU up to the same performance as GPU?** Yes, with OpenVINO we have the flexibility to run models on different hardware, so we can load the model in the CPU at the beginning, then we can load the model in GPU, and we can have the advantage to accelerate the performance of the model. Join the [Computer Vision Meetup on May 11](https://us02web.zoom.us/webinar/register/2516761651886/WN_D0MKbn3eTcSRtyHW1A6z1Q) at 10 AM PT for part 2 of today’s presentation, focused on performant ML models for Edge Apps using OpenVINO. **Can we have access to this Jupyter Notebook?** Yes, you can access the Jupyter Notebook in the [Anomalib repo on GitHub](https://github.com/openvinotoolkit/anomalib/tree/main/notebooks), including the [getting\_started](https://github.com/openvinotoolkit/anomalib/blob/main/notebooks/000_getting_started/001_getting_started.ipynb) notebook we shared today. **What real world examples do you see this type of anomaly detection being useful in? Other than the industry example you already showed?** In our example, we showed how useful Anomalib is in a manufacturing or factory setting. Other real-world use cases include security, such as screening for anomalies in luggage at airports. And also healthcare; for example, there are scenarios where we need to detect the presence of cancer in medical images. So Anomalib is not just for factory data, it's also popular for healthcare and security use cases. **What about edge inference? Are many anomaly detection networks capable of "compression" and therefore can be effectively deployed via tinyML?** OpenVINO gives us the flexibility to write models once and deploy them everywhere, including the ability to run inference at the edge. With OpenVINO you can use less memory, while also being able to make use of different types of hardware; so with just one model, you can deploy it everywhere, regardless of the hardware required to run it. Although there are competitors in this space, OpenVINO has differentiating features that make it attractive for a variety of use cases, including edge AI. **In coffee AI work, how was the labeling handled? Does it need thousands of images to be labeled?** Back when Paula was pursuing her PhD in computer vision and machine learning, she built a system for coffee production detection, before the launch of Intel Geti made the same use case significantly easier. Paula explains the requirements for image labeling across both scenarios: “During my PhD, I needed to annotate tens of thousands of images in video. (However) using Intel Geti, I just needed to annotate 20 images to start with in the first round of training.” If you are interested to learn more about Intel Geti, you visit geti.intel.com. **How easy is it for anyone who doesn’t have wide knowledge in deep learning to start with Intel Geti?** If somebody doesn't have knowledge in deep learning, they can start with Intel Geti in a simple way. Intel Geti has this flexibility because it abstracts away the technical complexity with an intuitive user interface. For example, in the use case regarding coffee production, we can involve accounting, the farmer, and the data scientist, each with differing levels of knowledge in deep learning. ## Additional Resources Check out the additional resources on the presentation: - [Talk transcript](https://www.rev.com/transcript-editor/shared/0QtDmPOtWEwHzoaPNobzvv8n0c6Wbn2E2EsK8BFZG3ifC3bnNxHxQ81bcKfznf0PGfnRFQonsb9JKfuK7-9kh1LVjYI?loadFrom=SharedLink) - [Connect with Paula on LinkedIn](https://www.linkedin.com/in/paula-ramos-41097319/) Thank you Paula on behalf of the entire Computer Vision Meetup community for sharing your expertise on working with Anomalib, OpenVINO, and Intel Geti to optimize computer vision models, especially at the edge! ## Understanding Speech Recognition Using OpenAI's Whisper Model ### Video Replay https://www.youtube.com/watch?v=BGio73nVvAQ ### Executive Summary [Vishal Rajput](https://www.linkedin.com/in/vishal-rajput-999164122/?originalSubdomain=be), AI Vision Engineer, Author, presents an overview of speech recognition technology and its importance in today's digital world. To start, Vishal provides an overview of what speech recognition is, as described by a Google search he performed that surfaced the top result from Oxford Languages: “the ability of a computer to identify and respond to the sounds produced in human speech”. Why do we need speech recognition systems and how can they help us? Vishal hits the highlights. First, we can, so why not? Next, they can make our lives easier and more comfortable. Additionally, they can be used in scenarios where we cannot use hands. Lastly, they speed things up – speech can be almost three times faster than typing. In order to get us up to speed in modern speech recognition, Vishal first takes us back to 2013-2015 to give us a glimpse into early speech recognition systems. Early speech recognition technologies, such as Cortana, had some challenges, including low signal-to-noise ratio, variability in speaker accents, and difficulty in understanding natural conversational speech. This is in part because of how they operated. Each model was trained differently, with different objectives. Changing one model created blind spots in terms of how it affected the others. ![](https://cdn.sanity.io/images/h6toihm1/production/4ee3e92f150e98bce8e049f9313d18f2a33ad55e-1776x738.png?auto=format&dpr=2&fit=max&q=75&w=1600) To achieve better results, deep learning was introduced, first into the acoustic models. Then the next step was to make an end-to-end deep learning speech engine. And you can see as the data and model size increased, speech recognition accuracy increased significantly. ![](https://cdn.sanity.io/images/h6toihm1/production/01780d827240a1a517c10c4778d892edfacf2479-1088x804.png?auto=format&dpr=2&fit=max&q=75&w=1088) Vishal then discusses various deep learning models used in speech recognition systems and specifically calls out two models, which are very important in the development of speech recognition: Connectionist Temporal Classification (CTC), which solves the alignment problem in speech recognition, and sequence to sequence (listen attend and spell, or LAS). Vishal notes that these two important models predated attention, which came out in the [_Attention Is All You Need_](https://arxiv.org/abs/1706.03762) paper in 2017. Next, Vishal takes time to appreciate other important work in speech recognition, including: wav2vec, vq-wav2vec, wav2vec2, and XLSR-wav2vec. Before diving into OpenAI Whisper, Vishal notes that there are two ways to make speech recognition systems: _supervised_ and _unsupervised_. Unsupervised offers more than a million hours of audio data; while supervised has only about 5000 hours of data available. What was OpenAI Whisper able to do to outperform previous models? First, OpenAI Whisper introduced _weak supervision_ to scale its data from 5,000 hours to 680,000 hours, bringing it close to the scale of unsupervised systems. Additionally, OpenAI used automated filtering methods on the automated transcripts available on the internet, as well as manual inspection, to improve the quality of the transcripts. Looking at Whisper’s transformer block, the architecture is similar to an off-the-shelf transformer with the main differences being the use of a Mel Spectrogram and the use of specialized tokens. ![](https://cdn.sanity.io/images/h6toihm1/production/5d03f365530c4976781959bed7f29ad04d9f82fc-1900x932.png?auto=format&dpr=2&fit=max&q=75&w=1600) Vishal explores the tokens (noting that you don't actually see this happening; this is happening in the back end): language tag, no speech, transcribe, translate, and timestamp tokens. The model is trained on 99 different languages and can detect which language is being spoken, transcribe it into text, translate it into a different language if needed, and provide timestamps for each sound or word in the audio sample. ![](https://cdn.sanity.io/images/h6toihm1/production/9af4fea7e520cf3981ff59b1a6332ef7fd0e0ff1-1900x950.png?auto=format&dpr=2&fit=max&q=75&w=1600) Additionally, OpenAI fine-tuned the Whisper model to better differentiate between different speakers talking and to standardize text (example, to standardize color and colour) before calculating the word error rate (WER) metric. Speech recognition models still have room for improvement in the areas of improving decoding strategies on long-form transcriptions, increasing training data for lower-resource languages, studying fine-tuning, and the impact of language models on robustness. Nonetheless, Whisper still produces impressive results. To give us a taste, Vishal runs 25 seconds of audio through Whisper to show how accurately it transcribes audio to text. In the example of a narrative spoken in a Scottish accent, Whisper misconstrued only one word (mistaking Eldons as Yildens) in the entire audio file! It’s easy to get started with the Whisper model (start [here on GitHub](https://github.com/openai/whisper), for example). ## Q&A Recap **For new ASR pipelines, is there vector quantization? Gaussian mixture models remind me of diffusion models.** Yes, for ASR pipelines, Gaussian mixture models are similar to diffusion models. **Can you please explain how using FFT preserves the sequence in the speech?** FFT doesn't need to preserve the sequences because the window size is very, very small, like 20 milliseconds. FFT doesn't preserve, but RNN does. **Besides contrastive loss, there is reconstruction loss. L1 reconstruction performs better than L2 reconstruction and why?** Generally what happens is you combine the losses to have a better performance rather than just using L1 loss or L2 loss. It's a combination of both, which performs better than one of them individually. And performance depends on the task. Sometimes L2 will perform better and sometimes L1 will perform better. **What is the difference between GPT and BERT?** As far as I understand GPT, GPT is not bidirectional because it is predicting the future, but BERT is bidirectional. (This notion is mentioned in the name–bidirectional encoder representations for transformers.) That is the primary difference between the two. **In what real world scenarios does the Whisper model have advantages over traditional ASR?** The Whisper model can definitely handle noise much better. It is actually on par with human understanding or in some cases even better than humans. ## Additional Resources Check out these additional resources: - [Talk transcript](https://www.rev.com/transcript-editor/shared/RfBmeLUpx14C0lCJMxOIuTViqDKOEPOzhpEv_R5VVCpVk8mSULBWM2iGDDVWXP4ErF_h-fkG-KY0RYT-HQ3ExM78wNU?loadFrom=SharedLink) - [Presentation slides](https://voxel51.com/wp-content/uploads/2023/02/Speech-Recognition-Computer-Vision-Meetup-Vishal-Rajput.pdf) - [Vishal on LinkedIn](https://www.linkedin.com/in/vishal-rajput-999164122/) - [AI Medium publication](https://medium.com/aiguys) A big thank you to Vishal on behalf of the entire Computer Vision Meetup community for getting us up-to-speed on speech recognition and the latest Whisper model by OpenAI. ## Computer Vision Meetup Locations Computer Vision Meetup membership has grown to nearly [3,000 members](https://www.meetup.com/pro/computer-vision-meetups/) in just a few months! The goal of the meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. New Meetup Alert – We just added a Computer Vision Meetup location in Singapore! Join one of the (now) 13 Meetup locations closest to your timezone. - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) ## Upcoming Computer Vision Meetup Speakers & Schedule We have exciting speakers already signed up over the next few months! Become a member of the [Computer Vision Meetup closest to you](https://www.meetup.com/pro/computer-vision-meetups/), then register for the Zoom for the Meetups of your choice. ### March 9 @ 10AM PT - Lighting up Images in the Deep Learning Era — [Soumik Rakshit](https://www.linkedin.com/in/soumikrakshit/), ML Engineer (Weights & Biases) - Training and Fine Tuning Vision Transformers Efficiently with Colossal AI — [Sumanth P](https://www.linkedin.com/in/sumanth-p-09b339173/) (ML Engineer) - [Zoom Link](https://us02web.zoom.us/webinar/register/8816708728020/WN_mTdNXxTSR-e7bDG5EkH1XQ) ### April 13 @ 10AM PT - Emergence of Maps in the Memories of Blind Navigation Agents — [Dhruv Batra](https://www.linkedin.com/in/dhruv-batra-dbatra/) (Meta & Georgia Tech) - Talk 2 — coming soon! - [Zoom Link](https://us02web.zoom.us/webinar/register/3416762523346/WN_oXTCgqQiQT6ouNxygQkpsg) ### May 11 @ 10AM PT - Machine Learning for Fast, Motion-Robust MRI — [Nalini Singh](https://www.linkedin.com/in/nalinimsingh/) (MIT) - Quick and Performant Machine Learning Models for Edge Applications using OpenVINO — [Paula Ramos](https://www.linkedin.com/in/paula-ramos-41097319/), PhD (Intel) - [Zoom Link](https://us02web.zoom.us/webinar/register/2516761651886/WN_D0MKbn3eTcSRtyHW1A6z1Q) ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [Anomalib](https://voxel51.com/blog/tag/anomalib) [ASR](https://voxel51.com/blog/tag/asr) [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [Edge AI](https://voxel51.com/blog/tag/edge-ai) [OpenAI Whisper model](https://voxel51.com/blog/tag/openai-whisper-model) [OpenVINO](https://voxel51.com/blog/tag/openvino) [speech recognition](https://voxel51.com/blog/tag/speech-recognition) [Whisper](https://voxel51.com/blog/tag/whisper) [Whisper model](https://voxel51.com/blog/tag/whisper-model) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/991b89d515d9f7fe4eea26c14396eb51116b4a8b-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup – April 27, 2023\\ \\ Event Recaps\\ \\ • \\ \\ Apr 28, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ed751b5c0fbc3d2cb74f0b30e6418d3319564c-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — December 2022\\ \\ Event Recaps\\ \\ • \\ \\ Dec 13, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/bbb1d9add0b0b9aa12682acac795df7c2ba760a9-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — November 2022\\ \\ Event Recaps\\ \\ • \\ \\ Nov 16, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-257-lllmstxt|> ## FiftyOne 0.19 Release [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Announcing FiftyOne 0.19 with Spaces, In-App Embeddings Visualization, Saved Views, and More! Feb 16, 2023 • 8 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-fiftyone-0-19#aca80875d890) [tl;dr: What’s new in FiftyOne 0.19?](https://voxel51.com/blog/announcing-fiftyone-0-19#f07fdfd1c628) [Live demo & AMA on Feb. 28 @ 10 AM PT](https://voxel51.com/blog/announcing-fiftyone-0-19#8af7c10f7f55) [Spaces](https://voxel51.com/blog/announcing-fiftyone-0-19#c347cac4b10e) [In-App embeddings visualization](https://voxel51.com/blog/announcing-fiftyone-0-19#e277876a1ba3) [Saved views](https://voxel51.com/blog/announcing-fiftyone-0-19#8ba511d1318e) [On-disk segmentations](https://voxel51.com/blog/announcing-fiftyone-0-19#154b571f691c) [New UI filtering options](https://voxel51.com/blog/announcing-fiftyone-0-19#85043492b251) [FiftyOne Teams documentation](https://voxel51.com/blog/announcing-fiftyone-0-19#2fc511d8042d) [Community contributions](https://voxel51.com/blog/announcing-fiftyone-0-19#1ecddef143c1) [FiftyOne community updates](https://voxel51.com/blog/announcing-fiftyone-0-19#421316920897) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-fiftyone-0-19#aca80875d890) [tl;dr: What’s new in FiftyOne 0.19?](https://voxel51.com/blog/announcing-fiftyone-0-19#f07fdfd1c628) [Live demo & AMA on Feb. 28 @ 10 AM PT](https://voxel51.com/blog/announcing-fiftyone-0-19#8af7c10f7f55) [Spaces](https://voxel51.com/blog/announcing-fiftyone-0-19#c347cac4b10e) [In-App embeddings visualization](https://voxel51.com/blog/announcing-fiftyone-0-19#e277876a1ba3) [Saved views](https://voxel51.com/blog/announcing-fiftyone-0-19#8ba511d1318e) [On-disk segmentations](https://voxel51.com/blog/announcing-fiftyone-0-19#154b571f691c) [New UI filtering options](https://voxel51.com/blog/announcing-fiftyone-0-19#85043492b251) [FiftyOne Teams documentation](https://voxel51.com/blog/announcing-fiftyone-0-19#2fc511d8042d) [Community contributions](https://voxel51.com/blog/announcing-fiftyone-0-19#1ecddef143c1) [FiftyOne community updates](https://voxel51.com/blog/announcing-fiftyone-0-19#421316920897) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Voxel51 in conjunction with the FiftyOne community is excited to announce the general availability of [FiftyOne 0.19](https://docs.voxel51.com/release-notes.html#fiftyone-0-19-0). This release is packed with new features that make it even easier and faster to visualize your computer vision datasets and boost the performance of your machine learning models. How? Read on! ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is the open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into the release. ## **tl;dr: What’s new in FiftyOne 0.19?** This release includes: - [**Spaces**](https://voxel51.com/blog/announcing-fiftyone-0-19#spaces): an all-new customizable framework for organizing interactive information panels within the FiftyOne App, allowing you to visualize and query your datasets in powerful new ways through a convenient interface - [**In-App embeddings visualization**](https://voxel51.com/blog/announcing-fiftyone-0-19#in-app-embeddings): you can now interactively explore embeddings visualizations natively in the App by opening an embeddings panel with one click - [**Saved views**](https://voxel51.com/blog/announcing-fiftyone-0-19#saved-views): you can now save views into your datasets and switch between them natively in the App - [**On-disk segmentations**](https://voxel51.com/blog/announcing-fiftyone-0-19#on-disk-segmentations): you can now store your semantic segmentation masks and heatmaps on disk, rather than in the database - [**New UI filtering options**](https://voxel51.com/blog/announcing-fiftyone-0-19#new-ui-filtering-options): the App’s sidebar now contains upgraded options for filtering datasets - [**FiftyOne Teams documentation**](https://voxel51.com/blog/announcing-fiftyone-0-19#teams-documentation): documentation for [FiftyOne Teams](https://docs.voxel51.com/teams/index.html#fiftyone-teams) is now publicly available! Check out the [release notes](https://voxel51.com/docs/fiftyone/release-notes.html#fiftyone-0-19-0) for a full rundown of additional enhancements and bugfixes. ## **Live demo & AMA on Feb. 28 @ 10 AM PT** \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop See all of the new features in action in the webinar and AMA we held on February 28, 2023. I demoed all of the new features in FiftyOne 0.19, which you can check out in [the recap and the video playback](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision/). Now, here’s a quick overview of some of the new features we packed into this release. ## **Spaces** FiftyOne 0.19 debuts Spaces, a customizable framework for organizing interactive information Panels in the App. As of FiftyOne 0.19, the following Panel types are included natively: - [Samples Panel](https://docs.voxel51.com/user_guide/app.html#app-samples-panel): the media grid that loads by default when you launch the App - [Histograms Panel](https://docs.voxel51.com/user_guide/app.html#app-histograms-panel): a dashboard of histograms for the fields of your dataset - **(New!)** [Embeddings Panel](https://docs.voxel51.com/user_guide/app.html#app-embeddings-panel): a canvas for working with embeddings visualizations - [Map Panel](https://docs.voxel51.com/user_guide/app.html#app-map-panel): visualizes the geolocation data of datasets that have a GeoLocation field - You can also configure [custom Panels via plugins](https://docs.voxel51.com/plugins/index.html#fiftyone-plugins)! In the screenshot below, for example, we’ve added the Embeddings and Map Panels to the default Samples Panel so we can visualize all three together seamlessly in the App. ![](https://cdn.sanity.io/images/h6toihm1/production/8651d28f4a2978eb72cca25ef09e1f01f81847ea-2968x2042.png?auto=format&dpr=2&fit=max&q=75&w=1600) You can configure Spaces visually in the App in a variety of ways described below. 1\. Click the + icon in any Space to add a new Panel: ![](https://cdn.sanity.io/images/h6toihm1/production/ae4cb879e96ddb7e9861fe7b180b9b8f4634bebc-1168x831.gif?auto=format&dpr=2&fit=max&q=75&w=1168) 2\. When you have multiple Panels open in a Space, you can use the divider buttons to split the Space either horizontally or vertically: ![](https://cdn.sanity.io/images/h6toihm1/production/d1a67549bdfe309f54136029c73929c4a6a44e6c-1171x831.gif?auto=format&dpr=2&fit=max&q=75&w=1171) 3\. You can rearrange Panels at any time by dragging their tabs between Spaces, or close a Panel by clicking on its x icon: ![](https://cdn.sanity.io/images/h6toihm1/production/16fe3c77059f1f9d3abab0729b0b874dfb344b23-1169x831.gif?auto=format&dpr=2&fit=max&q=75&w=1169) You can also programmatically configure your Spaces layout from Python! The code sample below shows an end-to-end example of loading a dataset, generating an embeddings visualization [via the FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html), and launching the App with a customized Spaces layout that includes the Samples Panel, Histograms Panel, and Embeddings Panel with the Brain result already loaded: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4 5dataset = foz.load_zoo_dataset("quickstart") 6fob.compute_visualization(dataset, brain_key="img_viz") 7 8samples_panel = fo.Panel( 9 type="Samples", 10 pinned=True, # don’t allow closing 11) 12 13histograms_panel = fo.Panel( 14 type="Histograms", 15 state=dict(plot="Labels"), # open label fields by default 16) 17 18# Open the visualization we generated above by default 19embeddings_panel = fo.Panel( 20 type="Embeddings", 21 state=dict(brainResult="img_viz", colorByField="metadata.size_bytes"), 22) 23 24spaces = fo.Space( 25 children=[\ 26 fo.Space(\ 27 children=[\ 28 fo.Space(children=[samples_panel]),\ 29 fo.Space(children=[histograms_panel]),\ 30 ],\ 31 orientation="horizontal",\ 32 ),\ 33 fo.Space(children=[embeddings_panel]),\ 34 ], 35 orientation="vertical", 36) 37 38session = fo.launch_app(dataset, spaces=spaces) ``` Check out [the docs](https://docs.voxel51.com/user_guide/app.html#spaces) for more information about using and configuring Spaces layouts. ## **In-App embeddings visualization** New in FiftyOne 0.19 (and enabled by the Spaces feature above), when you load a dataset in the App that contains an [embeddings visualization](https://docs.voxel51.com/user_guide/brain.html#brain-embeddings-visualization), you can open the [Embeddings Panel](https://docs.voxel51.com/user_guide/app.html#embeddings-panel) to visualize and interactively explore a scatterplot of the embeddings in the App. For example, try running the code below to download a dataset, generate two embeddings visualizations on it, and launch the App: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4 5dataset = foz.load_zoo_dataset("quickstart") 6 7# Image embeddings 8fob.compute_visualization(dataset, brain_key="img_viz") 9 10# Object patch embeddings 11fob.compute_visualization( 12 dataset, patches_field="ground_truth", brain_key="gt_viz" 13) 14 15session = fo.launch_app(dataset) ``` Then click on the + icon next to the Samples tab to open the Embeddings Panel and use the two menus in the upper-left corner of the Panel to configure your plot: - **Brain key**: the Brain key associated with the [compute\_visualization()](https://docs.voxel51.com/user_guide/brain.html#visualizing-embeddings) run to display - **Color by**: an optional sample field (or label attribute, for patches embeddings) to color the points by From there you can lasso points in the plot to show only the corresponding samples/patches in the Samples Panel: ![](https://cdn.sanity.io/images/h6toihm1/production/94c168b1408e66a8d61a767d02e7ba541cd34cd6-1235x770.gif?auto=format&dpr=2&fit=max&q=75&w=1235) The Embeddings Panel also provides a number of additional controls: - Press the **pan** icon in the menu (or type g) to switch to pan mode, in which you can click and drag to change your current field of view - Press the **lasso** icon (or type s) to switch back to lasso mode - Press the **locate** icon to reset the plot’s viewport to a tight crop of the current view’s embeddings - Press the **x** icon (or double click anywhere in the plot) to clear the current selection When coloring points by categorical fields (strings and integers) with fewer than 100 unique classes, you can also use the legend to toggle the visibility of each class of points: - Single click on a legend trace to show/hide that class in the plot - Double click on a legend trace to show/hide all other classes in the plot ![](https://cdn.sanity.io/images/h6toihm1/production/5fda8c54ad1d40a588a9f1a3f633e5616a924ded-1234x769.gif?auto=format&dpr=2&fit=max&q=75&w=1234) As demonstrated in the previous section, the Embeddings Panel can also be programmatically configured via Python. Check out [the docs](https://docs.voxel51.com/user_guide/app.html#embeddings-panel) for more information about working with embeddings visualizations in the App. ## **Saved views** In FiftyOne 0.19 you can use a new menu in the upper-left of the App to record the current state of the App’s view bar and filters sidebar as a **saved view** into your dataset: ![](https://cdn.sanity.io/images/h6toihm1/production/7a96144a1be2b0b2d7dc10994507899a69888b96-917x605.gif?auto=format&dpr=2&fit=max&q=75&w=917) Saved views are persisted on your dataset under a name of your choice so that you can quickly load them in a future session via the UI or Python. Saved views are a convenient way to record semantically relevant subsets of a dataset, such as: - Samples in a particular state, e.g. with certain tag(s) - A subset of a dataset that was used for a task, e.g. training a model - Samples that contain content of interest, e.g. object types or image characteristics Remember that saved views only store the rules used to extract content from the underlying dataset, not the actual content itself. You can save hundreds of views into a dataset if desired without worrying about storage space. You can load a saved view at any time by selecting it from the saved view menu: ![](https://cdn.sanity.io/images/h6toihm1/production/63d831cd5dbac6e383f99c8e96e96a3814bb056d-916x639.gif?auto=format&dpr=2&fit=max&q=75&w=916) You can also edit or delete saved views by clicking on their pencil icon: ![](https://cdn.sanity.io/images/h6toihm1/production/2bf6abc7ae44564977ede3379354b582058b3df7-915x638.gif?auto=format&dpr=2&fit=max&q=75&w=915) You can also programmatically create saved views [via Python](https://docs.voxel51.com/user_guide/using_views.html#saving-views): ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5dataset = foz.load_zoo_dataset("quickstart") 6dataset.persistent = True 7 8# Create a view 9cats_view = ( 10 dataset 11 .select_fields("ground_truth") 12 .filter_labels("ground_truth", F("label") == "cat") 13 .sort_by(F("ground_truth.detections").length(), reverse=True) 14) 15 16# Save the view 17dataset.save_view("cats-view", cats_view) ``` And load them in future sessions (including saved views created via the App): ```python 1import fiftyone as fo 2 3dataset = fo.load_dataset("quickstart") 4 5# Retrieve a saved view 6cats_view = dataset.load_saved_view("cats-view") 7print(cats_view) ``` Check out [the docs](https://docs.voxel51.com/user_guide/app.html#saving-views) for more information about using saved views in the App and Python. ## **On-disk segmentations** In prior FiftyOne versions, [semantic segmentations](https://docs.voxel51.com/user_guide/using_datasets.html#semantic-segmentation) and [heatmaps](https://docs.voxel51.com/user_guide/using_datasets.html#heatmaps) could only be stored as compressed bytes directly in the database. Now in FiftyOne 0.19, you can store segmentations and heatmaps as images on disk and store (only) their paths on your FiftyOne datasets, just like you do for the primary media of each sample: ```python 1import cv2 2import numpy as np 3 4import fiftyone as fo 5 6# Example segmentation mask 7mask_path = "/tmp/segmentation.png" 8mask = np.random.randint(10, size=(128, 128), dtype=np.uint8) 9cv2.imwrite(mask_path, mask) 10 11sample = fo.Sample(filepath="/path/to/image.png") 12sample["segmentation"] = fo.Segmentation(mask_path=mask_path) 13 14print(sample) ``` Segmentation masks can be stored in either of these formats on disk: - 2D 8-bit or 16-bit images - 3D 8-bit RGB images When you load datasets with segmentation fields containing 2D masks in the App, each pixel value is rendered as a different color from the App’s color pool so that you can visually distinguish the classes. When you view RGB segmentation masks in the App, the mask colors are always used. You can also [store semantic labels](https://docs.voxel51.com/user_guide/using_datasets.html#storing-mask-targets) for your segmentation fields on your dataset. Then, when you view the dataset in the App, label strings will appear in the App’s tooltip when you hover over pixels. If you are working with 2D segmentation masks, specify target keys as integers: ```python 1import fiftyone as fo 2 3dataset = fo.Dataset() 4dataset.default_mask_targets = {1: "cat", 2: "dog"} ``` And if you are working with RGB segmentation masks, specify target keys as RGB hex strings: ```python 1import fiftyone as fo 2 3dataset = fo.Dataset() 4dataset.default_mask_targets = {"#499CEF": "cat", "#6D04FF": "dog"} ``` The entire FiftyOne API was upgraded to support on-disk and/or RGB segmentations: - Evaluation via [evaluate\_segmentations()](https://docs.voxel51.com/user_guide/evaluation.html#semantic-segmentations) natively supports on-disk and/or RGB segmentations - The [apply\_model()](https://docs.voxel51.com/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.apply_model) method now has an optional output\_dir argument specifying where to store semantic segmentation inferences as images on disk - There’s a new [export\_segmentations()](https://docs.voxel51.com/api/fiftyone.utils.labels.html#fiftyone.utils.labels.export_segmentations) utility for conveniently exporting in-database segmentations to on-disk images - Other new utilities like [transform\_segmentations()](https://docs.voxel51.com/api/fiftyone.utils.labels.html#fiftyone.utils.labels.transform_segmentations) are now available for manipulating segmentations Check out [the docs](https://docs.voxel51.com/user_guide/using_datasets.html#semantic-segmentation) for more information about adding on-disk segmentations to your FiftyOne datasets. ## **New UI filtering options** We’re constantly improving and extending the filtering options available natively in the App to provide more powerful and intuitive ways to query datasets. In FiftyOne 0.19, we added a new selector that allows you to fine-tune your filters in the sidebar. For example, when filtering by the label attribute of a Detections field, you can choose between the following options: - (default): Filter to only show objects with the specified labels (omitting samples with no matching objects) - Exclude objects with the specified labels - Show samples that contain the specified labels (without filtering) - Omit samples that contain the specific labels All applicable filtering options are available from both the grid view and the [sample modal](https://docs.voxel51.com/user_guide/app.html#viewing-a-sample), and for all field types, including top-level fields and [dynamic label attributes](https://docs.voxel51.com/user_guide/using_datasets.html#dynamic-attributes)! ![](https://cdn.sanity.io/images/h6toihm1/production/52efef8c31aa755a449c0343c2882822e18f164a-916x640.gif?auto=format&dpr=2&fit=max&q=75&w=916) ## **FiftyOne Teams documentation** Exciting news! Documentation for FiftyOne Teams is now publicly available at [https://docs.voxel51.com/teams](https://docs.voxel51.com/teams). FiftyOne Teams enables multiple users to securely collaborate on the same datasets and models, either on-premises or in the cloud, all built on top of the open source FiftyOne workflows that you’re already relying on. Look interesting? [Schedule a demo](https://voxel51.com/get-fiftyone-teams) to get started with FiftyOne Teams yourself. ## **Community contributions** Shoutout to the following community members who contributed to this release! - [kalpit-S](https://github.com/kalpit-S) contributed [#2354 - added help link for Mapbox configuration in App](https://github.com/voxel51/fiftyone/pull/2354) - [flakeice](https://github.com/flakeice) contributed [#2359 - fix bug when loading datasets in VOC format](https://github.com/voxel51/fiftyone/pull/2359) - [Rustem Galiullin](https://github.com/Rusteam) contributed [#2353 - add support for custom CVAT task names](https://github.com/voxel51/fiftyone/pull/2353) - [Rustem Galiullin](https://github.com/Rusteam) contributed [#2373 - exact frame count support](https://github.com/voxel51/fiftyone/pull/2373) - [Oguz-hanoglu](https://github.com/oguz-hanoglu) contributed [#2297 - improved explanation of sidebar modes in the App](https://github.com/voxel51/fiftyone/pull/2297) - [Jamie Werther](https://github.com/jwertherUM) contributed [#2427 - show only supported eval keys](https://github.com/voxel51/fiftyone/pull/2427) - [Nikita Manovich](https://github.com/nmanovic) contributed [#2478 - Fix several CVAT links](https://github.com/voxel51/fiftyone/pull/2478) - [Chris Hall](https://github.com/shortcipher3) contributed [#2561 - updated CVAT links](https://github.com/voxel51/fiftyone/pull/2561) ## **FiftyOne community updates** The FiftyOne community continues to grow! - 1,300+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,500+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 245+ repositories - 56+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne 0.19](https://voxel51.com/blog/tag/fiftyone-0-19) [on-disk segmentations](https://voxel51.com/blog/tag/on-disk-segmentations) [open source](https://voxel51.com/blog/tag/open-source) [product release](https://voxel51.com/blog/tag/product-release) [saved views](https://voxel51.com/blog/tag/saved-views) [Spaces](https://voxel51.com/blog/tag/spaces) [UI filtering](https://voxel51.com/blog/tag/ui-filtering) ![](https://cdn.sanity.io/images/h6toihm1/production/8d61ff90b31d151405f9e21a33c2802509f34651-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Brian Moore Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a4c2bee9ed053c5be2a1c161e5abf758c9a12ff8-1400x923.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.18 with App Performance Improvements, Sidebar Modes, and Custom Attributes\\ \\ Product & News\\ \\ • \\ \\ Nov 15, 2022](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes) [![](https://cdn.sanity.io/images/h6toihm1/production/e94f20fa81716294c7a6caccf7256e9106cb6e89-967x800.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.17 with Grouped Datasets, 3D, Geolocation, and Custom Plugins\\ \\ Product & News\\ \\ • \\ \\ Sep 21, 2022](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins) [![](https://cdn.sanity.io/images/h6toihm1/production/ed0b9dc4072cfa3d1d1bd2837c0b834fccec607f-1400x775.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Introducing FiftyOne: A Tool for Rapid Data & Model Experimentation\\ \\ Product & News\\ \\ • \\ \\ Sep 12, 2020](https://voxel51.com/blog/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-258-lllmstxt|> ## FiftyOne Community Update [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) FiftyOne Computer Vision Community Update – Feb ‘23 Feb 21, 2023 • 7 min read Article content In this article [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#1c44789224f7) [Community Spotlights](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#9b604216c75a) [FiftyOne Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#2e77d52e2d06) [FiftyOne 0.19 Is Here!](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#53c021d3b523) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#bc32ac99d1e3) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#601c364a38c9) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#4bc155b6703d) [Event Calendar](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#fec8101a9d26) [Community-Powered Charitable Contributions](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#47c1dc2376f8) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#a69efeb0cb2d) [We Are Hiring!](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#cc155c1db6ac) In this article [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#1c44789224f7) [Community Spotlights](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#9b604216c75a) [FiftyOne Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#2e77d52e2d06) [FiftyOne 0.19 Is Here!](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#53c021d3b523) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#bc32ac99d1e3) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#601c364a38c9) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#4bc155b6703d) [Event Calendar](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#fec8101a9d26) [Community-Powered Charitable Contributions](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#47c1dc2376f8) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#a69efeb0cb2d) [We Are Hiring!](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023#cc155c1db6ac) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the first post in our monthly blog series where we bring you up to speed on recent happenings in the FiftyOne community and celebrate noteworthy milestones. 🙌 🚀 ## Voxel51’s Commitment to Open Source and Community If you’re new to Voxel51, open source, transparency, and giving back to the computer vision community are what we are all about! Whether it’s developing the open source [FiftyOne computer vision toolset](https://github.com/voxel51/fiftyone) to help engineers and data scientists build high-quality datasets and models, sponsoring [Meetups](https://www.meetup.com/pro/computer-vision-meetups/) to help members boost their computer vision knowledge, or [giving to charitable causes](https://voxel51.com/charitable-giving/) on behalf of the community, Voxel51 is committed to bringing transparency and clarity to the world's data. ## Community Spotlights We love hearing how FiftyOne helps you solve challenges and reach new heights! Curious what sorts of use cases are possible with FiftyOne? Here are just a few highlights from what people in the community have to say. ### **Research and Development - Raytheon** ![](https://cdn.sanity.io/images/h6toihm1/production/be35eb6d8955faa2e1ba93540de65e3a71af27ef-285x177.png?auto=format&dpr=2&fit=max&q=75&w=285) The [Raytheon Technologies Research Center](https://www.rtx.com/who-we-are/what-we-do/transformative-technologies/rtrc) serves as the company’s innovation hub. RTRC’s engineers, scientists, and researchers anticipate the discoveries destined to change everything, and they transform that research into the solutions and products that help the company’s businesses shape the future. > _“We use FiftyOne to organize large research datasets. My favorite feature is the ability to view distributions over image attributes in the dataset, and filter the dataset by those attributes.”_ — Brett Israelsen, Principal Research Scientist ### **Agriculture - Taranis** ![](https://cdn.sanity.io/images/h6toihm1/production/a3826946773c311c461f043c64f5ff73559a2bb0-600x314.jpg?auto=format&dpr=2&fit=max&q=75&w=600) [Taranis](https://www.taranis.com/) is a leading precision agriculture intelligence platform that uses sophisticated computer vision, data science, and deep learning algorithms to effectively monitor fields. Taranis offers a full-stack solution for high precision aerial surveillance imagery to prevent crop yield loss due to insects, diseases, and weeds. > _"We’ve been using FiftyOne for over a year and it has drastically changed the way we work. The ability to easily display and analyze our images and their metadata, including experiment results, has been a refreshing change compared to the way we've worked before - mainly writing our own metrics and viewers. I've personally used FiftyOne for a segmentation model I've trained - trying to analyze the results and visually see what my model outputs has been really easy and fluid thanks to FiftyOne.”_ — Ido Greenfeld, AI Team Lead ### **Waste Management/Materials Recovery - Binit** \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop [Binit](https://binit.ai/) is the first fully integrated analytics and data platform for waste management, helping Materials Recovery Facilities analyze materials and boost revenues via data. > _“FiftyOne is the backbone of our Data Engine. It helps us to clean and relabel our datasets efficiently, and convert our model predictions into large scale training datasets with a little help from CVAT.”_ — George Pearse, Machine Learning Engineer ## **FiftyOne Community Rewards** ![](https://cdn.sanity.io/images/h6toihm1/production/e3ed352e497d0e500ff4a1484b8422b3c9bef5cb-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) Is your organization using FiftyOne to solve interesting computer vision problems? [Share your success story](https://voxel51.com/fiftyone-computer-vision-success-story-submission/) and claim a box of community rewards as a thank you! ## FiftyOne 0.19 Is Here! FiftyOne adoption continues to accelerate with **more than 590k downloads** to date! Here’s the latest news on the product side: [FiftyOne 0.19](https://voxel51.com/blog/announcing-fiftyone-0-19/) is here and it’s packed with a bunch of new features including an all-new customizable Spaces framework, in-App embeddings visualization, saved views, on-disk segmentations, new UI filtering options, and more! Check out the new features in the [announcement blog post](https://voxel51.com/blog/announcing-fiftyone-0-19/), the latest [release notes](https://docs.voxel51.com/release-notes.html), and in the [live demo & AMA](https://voxel51.com/computer-vision-events/whats-new-in-fiftyone-0-19/?utm_source=blog) with Voxel51 CTO Brian Moore on February 28 @ 10 AM PT. Oh yea, you can also see it for yourself! It’s easy to [get up and running](https://docs.voxel51.com/index.html) in just a few minutes. ## FiftyOne on GitHub GitHub is home to the open source FiftyOne project. Here’s the latest snapshot of what’s happening in the [FiftyOne GitHub repo](https://github.com/voxel51/fiftyone): - **Total stars:** 2,562 - **Total contributors:** 56 - Shout out to recent contributors [shortcipher3](https://github.com/voxel51/fiftyone/commits?author=shortcipher3&since=2022-12-01&until=2023-01-31), [nmanovic](https://github.com/voxel51/fiftyone/commits?author=nmanovic&since=2022-12-01&until=2023-01-31), [jwertherUM](https://github.com/voxel51/fiftyone/commits?author=jwertherUM&since=2022-12-01&until=2023-01-31) and [kalpit-S](https://github.com/voxel51/fiftyone/commits?author=kalpit-S&since=2022-12-01&until=2023-01-31)! - **Total used by:** 245 repositories - **Total forks:** 306 - **Total issues closed so far this year**: 185 - **Total commits so far this year:** 392 ## FiftyOne Community Slack The FiftyOne Community [Slack channel](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) is where you can join more than 1,360 machine learning engineers and data scientists using FiftyOne to improve the quality of their computer vision data and build better models. Ask questions, answer questions, or simply follow along with the discussion! To make it easy to catch the highlights, every Friday we recap interesting questions and answers from Slack in [Tips & Tricks blog series](https://voxel51.com/blog/category/tips-tricks/). Recent posts include: - [Adding & Merging Data Tips & Tricks – Feb 17](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023/) - [Tips & Tricks – Feb 10](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023/) - [Model Evaluation Tips & Tricks – Feb 3](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023/) - [Tips & Tricks – Jan 27](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023/) - [View Stages Tips & Tricks – Jan 20](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023/) ## Computer Vision Meetups \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Voxel51 sponsors 13 virtual [Computer Vision Meetups](https://www.meetup.com/pro/computer-vision-meetups/) around the world. (To join, visit the Meetup link and scroll down to find the location friendliest to your time zone.) The Computer Vision Meetups are geared towards data scientists, machine learning engineers, and open source enthusiasts who want to expand their knowledge of computer vision and complementary technologies. We put an emphasis on open source software, and speakers who are computer vision practitioners or academics doing research in the field. ### **Recapping the Feb ‘23 Meetup** We recently held the February ‘23 Computer Vision Meetup that showcased these topics and speakers: - Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO — [Paula Ramos, PhD](https://www.linkedin.com/in/paula-ramos-41097319/) (Intel) - Understanding Speech Recognition with OpenAI’s Whisper Model — [Vishal Rajput](https://www.linkedin.com/in/vishal-rajput-999164122/) (AI-Vision Engineer) [Get the Meetup recap](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap/), including video playbacks, executive summaries, slides, and Q&A. ### **Upcoming: March ‘23 Meetup** Now that the event has passed, [check out the Meetup Recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023/) to tune into these talks: - Lightning Talk: A Recycling Max Pooling Module for 3D Point Cloud Analysis — [Jiajing Chen](https://www.linkedin.com/in/jiajing-chen-560189193/) (PhD Candidate, Syracuse University) - Lighting up Images in the Deep Learning Era — [Soumik Rakshit](https://www.linkedin.com/in/soumikrakshit/), ML Engineer (Weights & Biases) - Taking Computer Vision Models in Notebooks to Production — [Sumanth P](https://www.linkedin.com/in/sumanth-p-09b339173/) (ML Engineer) ## Event Calendar Check out these [upcoming](https://voxel51.com/computer-vision-events/) FiftyOne and computer vision events. - Feb 28: [FiftyOne 0.19 Live Demo & AMA](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision/) - Mar 9: [March Computer Vision Meetup](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023/) - Mar 22, NVIDIA GTC: Creating Digital Twins & Simulations of Industrial Robotic Workcells for Smart Factories, RIOS - Apr 13: [April Computer Vision Meetup](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023/) - May 11: [May Computer Vision Meetup](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023/) ## Community-Powered Charitable Contributions In the spirit of giving and in response to global circumstances over the last few years, we made some enhancements to our event strategy. Two notable changes: we added more virtual events so that more people can participate without the requirement to travel; and we applied our swag budget to donate to charitable causes on behalf of the community. During each of our virtual Meetups, we ask attendees to vote for their favorite charity, and then we make a donation to the charity that receives the most votes. In the last four months, on behalf of the computer vision community, we’ve made donations to these admirable causes: [Children International](https://www.children.org/), [Foundation Fighting Blindness](https://www.fightingblindness.org/), [Turkey-Syria Earthquake Relief](https://www.directrelief.org/emergency/turkey-syria-earthquake/) by Direct Relief, and [World Literacy Foundation](https://worldliteracyfoundation.org/). ## New Docs, Blogs, Videos, and Tutorials We want everyone to be successful with FiftyOne, and one of the ways we try to do that is by publishing resources that you might find helpful and handy. Here’s a list of some of the new [documentation](https://docs.voxel51.com/), [blogs](https://voxel51.com/blog/), [videos](https://www.youtube.com/@voxel51/videos), [tutorials](https://docs.voxel51.com/tutorials/index.html), [integrations](https://docs.voxel51.com/integrations/index.html), and [cheat sheets](https://docs.voxel51.com/cheat_sheets/index.html) that you may want to check out. ### **Blogs** - [Automatically Set Up a New ML Project, Pain Free](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free/) - [How Computer Vision is Changing Agriculture in 2023](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023/) - [People @ Voxel51: Spotlight on Jimmy Guerrero](https://voxel51.com/blog/people-voxel51-spotlight-on-jimmy-guerrero/) - [The Making of Avatar: The Way of Water](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water/) - [Recapping the Computer Vision Meetup – January 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023/) - [Finding Images with Words](https://voxel51.com/blog/finding-images-with-words/) - [Exploring the Berkeley Deep Drive Autonomous Vehicle Dataset](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset/) ### **Videos** - [Computer Vision Meetup: AI Deployment at the Edge with OpenVINO](https://www.youtube.com/watch?v=9TzdGz1NXac) - [Computer Vision Meetup: Speech Recognition Using OpenAI's Whisper Model](https://www.youtube.com/watch?v=BGio73nVvAQ) - [Exploring the Berkeley Deep Drive Autonomous Vehicle Dataset](https://www.youtube.com/watch?v=MX7VS_-oSFM) - [Computer Vision Meetup: Intro to Hugging Face Transformers](https://www.youtube.com/watch?v=yGUQSa6Emnc) - [Computer Vision Meetup: Hyperparameter Scheduling](https://www.youtube.com/watch?v=gwAk66W129Q) - [Pandas-Style Queries for Computer Vision Data](https://www.youtube.com/watch?v=0MhuMQuhYSw) ### **Tutorials and Cheat Sheets** - [Nearest Neighbor Embeddings Classification with Qdrant](https://docs.voxel51.com/tutorials/qdrant.html) - [Exploring Open Images V7 Tutorial](https://docs.voxel51.com/tutorials/open_images.html) - [Performing pandas-style queries in FiftyOne Tutorial](https://docs.voxel51.com/tutorials/pandas_comparison.html) - [pandas vs FiftyOne Cheat Sheet](https://docs.voxel51.com/cheat_sheets/index.html) - [FiftyOne Terminology Cheat Sheet](https://docs.voxel51.com/cheat_sheets/fiftyone_terminology.html) - [Views Cheat Sheet](https://docs.voxel51.com/cheat_sheets/views_cheat_sheet.html) - [Filtering Cheat Sheet](https://docs.voxel51.com/cheat_sheets/filtering_cheat_sheet.html) ### **New & Updated Documentation** We published a bunch of new docs to help you make the most of your FiftyOne experience: - Using [custom plugins](https://docs.voxel51.com/plugins/index.html#fiftyone-plugins) - Working with [sidebar modes](https://docs.voxel51.com/user_guide/app.html#app-sidebar-mode) - Customizing the sidebar layout with [sidebar groups](https://docs.voxel51.com/user_guide/app.html#app-sidebar-groups) - Viewing field-level descriptions with a [field tooltip](https://docs.voxel51.com/user_guide/app.html#app-fields-sidebar) - Declaring [custom dynamic attributes](https://docs.voxel51.com/user_guide/using_datasets.html#dynamic-attributes) on datasets - Storing [field-level metadata](https://docs.voxel51.com/user_guide/using_datasets.html#storing-field-metadata) on datasets - Option to import annotation IDs when loading data stored in [COCO format](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#cocodetectiondataset-import) - Including the export directory in the dataset.yaml file generated by [YOLOv5](https://docs.voxel51.com/user_guide/export_datasets.html#yolov5dataset-export) exports - Using CUDA devices when running the [CLIP model](https://docs.voxel51.com/user_guide/model_zoo/models.html#model-zoo-clip-vit-base32-torch) from the Model Zoo And [FiftyOne Teams](https://docs.voxel51.com/teams/index.html#fiftyone-teams) documentation is now publicly available! ## We Are Hiring! Voxel51 is on a mission to bring transparency and clarity to the world’s data. We’re growing quickly and are looking for people to grow with us. Voxel51 team members are fueled by learning, adapt quickly to face new challenges, and aim to shatter the status quo. Here’s a sample of some of our [open, remote positions](https://voxel51.com/jobs/): - Account Executive - DevOps Engineer - VP of Engineering [Community Update](https://voxel51.com/blog/tag/community-update) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [open source](https://voxel51.com/blog/tag/open-source) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/869a02098d1898869a250f4a5a23648c479af7da-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – April ‘23\\ \\ Product & News\\ \\ • \\ \\ Apr 6, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/338b38d41e6072dd11af86f21f5309337c52f36b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – May ‘23\\ \\ Product & News\\ \\ • \\ \\ May 5, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/7685a2b9b8681c3c641a1118ad1c4685f1af21b7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – July 2023\\ \\ Product & News\\ \\ • \\ \\ Jul 6, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-259-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – Feb 24, 2023 Feb 25, 2023 • 6 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#b38767f8be15) [Omitting classes with few detection instances](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#99ce2b7a910d) [Saving changes to sample fields](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#6513e27f8aa1) [Predicting class labels in homogeneous images](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#0367f11cbb74) [Matching classification results](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#5399712b0b2b) [Shutting down a session](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#90ac91d161c2) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#027902ce7260) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#b38767f8be15) [Omitting classes with few detection instances](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#99ce2b7a910d) [Saving changes to sample fields](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#6513e27f8aa1) [Predicting class labels in homogeneous images](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#0367f11cbb74) [Matching classification results](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#5399712b0b2b) [Shutting down a session](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#90ac91d161c2) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023#027902ce7260) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **Omitting classes with few detection instances** Community Slack member Sylvia Schmitt asked, _“When grouping samples by values in a specific field, I would like to omit samples which have values that rarely occur in the dataset. How can this be done?”_ One way to accomplish this would be to use `count_values()` to get a count of the number of occurrences of each unique value in the given field across the entire `Dataset` or `DatasetView` object, take the values that occur more frequently than your desired cutoff, and use the `match()` method to get samples that contain these. For instance, if you wanted to get samples from the test split of the [Families in the Wild](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#families-in-the-wild) dataset with `name` values that occur more than ten times in the dataset, you can do so as follows: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5## load the dataset 6dataset = foz.load_zoo_dataset("fiw", split="test") 7 8counts = dataset.count_values("name") 9keep_names = [name for name, count in counts.items() if count > 10] 10 11## filter for samples with these names 12view = dataset.match(F("name").is_in(keep_names)) 13 14session = fo.launch_app(view) ``` \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop You could then pass this resulting view into `group_by()` to group by the values in the field, or any other [Aggregation](https://docs.voxel51.com/user_guide/using_aggregations.html) you’d like. Learn more about [count\_values()](https://docs.voxel51.com/user_guide/using_aggregations.html#count-values), [is\_in()](https://docs.voxel51.com/api/fiftyone.core.expressions.html#fiftyone.core.expressions.ViewExpression.is_in), and [using aggregations](https://docs.voxel51.com/user_guide/using_aggregations.html#advanced-usage) in the FiftyOne Docs. ## **Saving changes to sample fields** Community Slack member Sylvia Schmitt asked, _“When adding sample fields and later on changing these values within a view, do the changes have to be made persistent by calling `save()` on the \`Dataset\` object, or will these changes be saved if the dataset is already persistent?”_ Great question, Sylvia! In general, when changes are made to an individual sample in a `Dataset` or `DatasetView`, the changes need to be saved by calling `save()` on the _sample_, not the dataset. This is the case even if the dataset is persistent, i.e. if ```python 1dataset.persistent = True ``` As an example, you could change the class label for the first detection of the first sample in the [Quickstart dataset](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#quickstart) as follows: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4## load dataset 5dataset = foz.load_zoo_dataset("quickstart") 6 7## get sample 8sample = dataset.first() 9 10## change label 11sample.ground_truth.detections[0].label = "bear" 12 13## save changes to dataset 14sample.save() ``` Using the `save()` method on a dataset is only necessary when editing dataset-level metadata such as `dataset.info`. There are a few cases, however, in which it is not necessary to explicitly run `sample.save()` to propagate changes back to the dataset. These include the `view.set_values(field_name, field_vals)` method, which takes in a list of values, `field_vals`, and writes these to the field `field_name` for the samples in the view, as well as the `view.tag_samples(tags)` method, which adds the tags `tags` to all samples in the view. If you know that you need to iterate through a `Dataset` or `DatasetView` and make changes to each sample, rather than call `save()` on each sample, it is more efficient to pass `autosave=True` into `iter_samples()`, which batches the operations. For example, to set a `random` field with a random number for each sample in our dataset, we can run: ```python 1import random 2 3import fiftyone as fo 4import fiftyone.zoo as foz 5 6## load dataset 7dataset = foz.load_zoo_dataset("quickstart") 8 9## Automatically saves sample edits in efficient batches 10for sample in dataset.select_fields().iter_samples(autosave=True): 11 sample["random"] = random.random() ``` Learn more about [set\_values()](https://docs.voxel51.com/api/fiftyone.core.view.html#fiftyone.core.view.DatasetView.set_values) and [tagging samples](https://docs.voxel51.com/user_guide/using_datasets.html#using-tags) in the FiftyOne Docs. ## **Predicting class labels in homogeneous images** Community Slack member George Pearse asked, _“What would be the best way to handle an application where the label for an object is deeply intertwined with the labels of the other objects in the sample? For instance, I might have images that are typically either crowds of all cats, or crowds of all dogs, but not crowds containing both cats and dogs.”_ Great question, George! There are many ways to deal with data like this. One approach would be to accumulate a lot of examples like this and train a model on this data. Given enough high quality examples, the model should (theoretically) be able to learn these relationships. As an alternative approach using just your existing data, you could perform post-processing to the labels in your samples based on the outputs of your model’s predictions. For instance, if your model’s predictions are stored in a `model_raw` field on your samples, you can create a new label field `model_processed` and populate the contents of this new field based on the contents of `model_raw` for that sample. For each sample, check if there are, say, three or more objects with the same class label. For the sake of simplicity, we’ll assume that `dog` is this class. If there are, then for all objects that are not labeled as `dog` s in `model_raw`, if their class confidence score is below some threshold, set their class label to `dog` in `model_processed`. Here’s what this might look like: ```python 1import numpy as np 2 3import fiftyone as fo 4import fiftyone.zoo as foz 5from fiftyone import ViewField as F 6 7## create or load your dataset 8dataset = fo.Dataset(..) 9 10## clone predictions into new field 11dataset.clone_sample_field( 12 "model_raw", 13 "model_processed" 14) 15 16## set a class confidence threshold 17conf_thresh = 0.3 18 19## iterate through samples in dataset 20for sample in dataset.iter_samples(autosave=True): 21 dets = sample.model_processed.detections 22 labels = [det.label for det in dets] 23 unique_labels, label_counts = np.unique(labels, return_counts=True) 24 25 ## find samples with at least 3 labels of same class 26 if max(label_counts) > 2: 27 crowd_label = unique_labels[np.argmax(label_counts)] 28 for det in dets: 29 if (det.label != crowd_label) and 30 (det.confidence < conf_thresh): 31 det.label = crowd_label 32 det.confidence = None 33 34 ## tag samples to look at later 35 sample.tags.append("possible homogeneous crowd") ``` You can then compare these tagged samples whose processed model predictions differ from the raw predictions, and inspect them in the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html). Learn more about [saving, keeping, and cloning sample fields](https://docs.voxel51.com/user_guide/using_views.html#saving-and-cloning) in the FiftyOne Docs. ## **Matching classification results** Community Slack member Nadav asked, _“I have a dataset with two kinds of classification. What is the best way to create a view, in code or in the app, that only contains samples on which the two classifications agree?”_ One way to do this in code is to use FiftyOne’s built-in [filtering and matching capabilities](https://docs.voxel51.com/user_guide/using_views.html#filtering). The `dataset.match(my_condition)` method will return a view consisting of all samples on which the condition `my_condition` is true. In your case, you can use the [ViewField](https://docs.voxel51.com/api/fiftyone.core.expressions.html#fiftyone.core.expressions.ViewField) to create the agreement condition between the two classifications. Here is what it could look like: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5# create or load your dataset with 6# classifications in field1 and field2 7dataset = fo.Dataset(...) 8 9view = dataset.match( 10 F("field1.label") == F("field2.label") 11) 12 13session = fo.launch_app(view) ``` If you instead wanted a view containing all samples where the two classifications did _not_ line up, you could replace the equality operator `==` with the inequality operator `!=`. Learn more about [filtering](https://docs.voxel51.com/cheat_sheets/index.html) in the FiftyOne Docs. ## **Shutting down a session** Community Slack member Scott asked, _“How can I disconnect the launched session?”_ In FiftyOne, a [Session](https://docs.voxel51.com/user_guide/app.html#app-sessions) is an instance of the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html#) connected to a specific `Dataset` or `DatasetView`. You can launch a session for a particular dataset or view with the `launch_app()` method: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4## load dataset 5dataset = foz.load_zoo_dataset("quickstart") 6 7## launch one session 8session1 = fo.launch_app(dataset) 9 10## create a view 11view = dataset.take(20) 12 13## launch another session 14session2 = fo.launch_app(view) ``` You can also view all registered sessions with `fo.core.session.session._subscribed_sessions`: defaultdict(set, {5151: {Dataset: quickstart Media type: image Num samples: 20 Selected samples: 0 Selected labels: 0 Session URL: http://localhost:5151/ View stages: 1\. Take(size=20, seed=None), Dataset: quickstart Media type: image Num samples: 20 Selected samples: 0 Selected labels: 0 Session URL: http://localhost:5151/ View stages: 1\. Take(size=20, seed=None)}}) When you terminate the Python process on which FiftyOne is running, all sessions are shut down, so typically you do not need to shut sessions down explicitly. However, if you would like to terminate a session at any point, you can do so using the private `_unregister_session()` method: ```python 1from fiftyone.core.session.session import _unregister_session 2_unregister_session(session1) ``` Learn more about sessions, including how to [launch multiple App instances](https://docs.voxel51.com/faq/index.html#faq-multiple-apps) on a remote machine, in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,350+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,550+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,200+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 246+ repositories - 56+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [classification](https://voxel51.com/blog/tag/classification) [Families in the Wild](https://voxel51.com/blog/tag/families-in-the-wild) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne App](https://voxel51.com/blog/tag/fiftyone-app) [object detection](https://voxel51.com/blog/tag/object-detection) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/205569e4c6b9ed68023e0ed430acbb40a981026f-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Tips and Tricks for Customizing your Computer Vision Workflows – Mar 03, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 4, 2023](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/5eca7f455938fa6e0b4755f60529b4123ae91046-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – May 26, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ May 26, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023) [FiftyOne Computer Vision Tips and Tricks – Mar 24, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 25, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-24-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-260-lllmstxt|> ## Spotlight on Lanny Wang [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) People @ Voxel51: Spotlight on Lanny Wang Feb 28, 2023 • 3 min read ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Last month we kicked off a series of posts to take you under the hood and introduce you to some of those incredible people on our team. It’s our hope you can get a sense of what it’s like to work at Voxel51 - directly from the team that is pushing our mission forward. What is our mission? Every day we wake up with the mission to bring transparency and clarity to the world’s data. It is exhilarating and meaningful work, but it doesn’t happen in a vacuum - it’s the product of the community and the team here that create and support the [open source FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. Oh! And if this seems like something you’d want to be part of, then definitely check out our remote, [open positions](https://voxel51.com/jobs/). We're building a fully-remote team of exceptional and diverse people who want to help bring data-centric machine learning to the world. We’re growing quickly and are looking for people to grow with us! And as always, if you like our open source and community work, please take a moment to [give the FiftyOne project a star](https://github.com/voxel51/fiftyone). Okay, this week we catch up with Software Engineer, [Lanny Wang](https://www.linkedin.com/in/lzwang/). Can’t wait for you to meet her :) ![](https://cdn.sanity.io/images/h6toihm1/production/48072eb9862ca80b22c8bce1c071cb0ebc02b2bc-671x671.png?auto=format&dpr=2&fit=max&q=75&w=671) **What do you do at Voxel51?** I am a software engineer at Voxel51. I work mostly on improving the features of the open source application [FiftyOne](https://github.com/voxel51/fiftyone). **Where are you based?** Houston, TX **What about the opportunity at Voxel51 inspired you to join?** I am very enthusiastic about the product. As someone who is visual-driven, I am interested in extracting meaningful insights from data, visualizing and curating them to help people make better decisions. And that’s what fiftyone is about — bringing **clarity** and **transparency** to all data. **What has been your favorite project to work on so far?** I have been at Voxel51 for three months. My favorite project so far is the feature of “Search by text similarity”, which allows the user to type in anything and search their samples using NLP search. It’s coming out in version 0.20. **What do you like most about your role?** I enjoy having the agency and independence to work on my projects. And when I want suggestions, people are always very open and accessible to discuss creative solutions. Also I am happy to work on open source projects, and I find it thrilling to receive feedback from the active community. It's rewarding to be part of a collaborative effort that helps make our app more extensible and flexible. **What do you like to do when you aren’t working?** I love plants and have a tropical garden. Caring for plants is very relaxing and healing. It's a balancing act of four variables: sunlight, humidity, medium, and temperature. Then all it needs is time. ![](https://cdn.sanity.io/images/h6toihm1/production/7ad2910093a9a0eb978b238ac278bd78459f2287-2248x1304.png?auto=format&dpr=2&fit=max&q=75&w=1600) **Three words that best describe you.** Observant, analytical, adaptable. **What is the one thing you can’t live without?** Fruits. Can’t name a fruit that I don’t like. Sugar + water = happiness. **Where’s your favorite place in the world?** Art Museum storage, where one can examine the objects closely. Objects tell their own stories. Plus, there are always some surprise findings. Prior to becoming a software engineer, I was trained in art history and material culture. I admire great craftsmanship and technological innovation, and aspire to create something that lasts through time. **What’s something you’re planning on doing in the next year that you’ve never done?** Grow mushrooms. **Are there open positions on your team?** VP Engineering and DevOps Engineer. [Check them out](https://voxel51.com/jobs/) and consider joining us! MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/99dc871d875865fc200f14931a3cfb3118ded624-1020x1007.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Tunnel vision in computer vision: can ChatGPT see?\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 16, 2022](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see) [![](https://cdn.sanity.io/images/h6toihm1/production/a735267ad7effa9f799f850ab7c8ffa241088710-1024x1024.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Why 2022 was the most exciting year in computer vision history (so far)\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 14, 2022](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far) [![](https://cdn.sanity.io/images/h6toihm1/production/0d65f67f35ed7c2afbfc5e6978967e492b357c43-1250x705.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Computer Vision Meetup Update — November ‘22\\ \\ Product & News\\ \\ • \\ \\ Nov 9, 2022](https://voxel51.com/blog/computer-vision-meetup-update-november-22) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-261-lllmstxt|> ## UCF101 Action Recognition Dataset [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) Exploring the UCF101 Dataset: A Large-Scale, YouTube-Based Action Recognition Dataset Mar 1, 2023 • 5 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#ef0faf030315) [About the UCF101 action recognition dataset](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#7e3c08a7c31e) [What is human action recognition?](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#47edc8d3708b) [Dataset quick facts](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#38cce0cf7cee) [Step 1: Install FiftyOne](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#e3a471113bbe) [Step 2: Import the dataset](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#e318105258c4) [Sample details](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#a06e5baafa1c) [Filtering by ID](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#f23bb11f7a54) [Filtering by label](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#91ece162ce3a) [Start working with the dataset](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#d567c80c5f3b) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#ef0faf030315) [About the UCF101 action recognition dataset](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#7e3c08a7c31e) [What is human action recognition?](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#47edc8d3708b) [Dataset quick facts](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#38cce0cf7cee) [Step 1: Install FiftyOne](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#e3a471113bbe) [Step 2: Import the dataset](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#e318105258c4) [Sample details](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#a06e5baafa1c) [Filtering by ID](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#f23bb11f7a54) [Filtering by label](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#91ece162ce3a) [Start working with the dataset](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset#d567c80c5f3b) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Welcome to the latest installment of our ongoing blog series where we highlight computer vision datasets in the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html)! FiftyOne provides a Dataset Zoo that contains a collection of common datasets that you can download and load into FiftyOne via a few simple commands. In this post, we explore the [UCF101 Action Recognition](https://www.crcv.ucf.edu/research/data-sets/ucf101/) video dataset. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop The FiftyOne Dataset Zoo comprises more than 30 datasets, with new datasets being added all the time! They cover a variety of computer vision use cases including: - Video - Images - Location - Point-cloud - Action-recognition - Classification - Detection - Segmentation - Relationships - And more! ![](https://cdn.sanity.io/images/h6toihm1/production/6e255bf059a91e67187551d90f4e81ae788aa38c-752x476.png?auto=format&dpr=2&fit=max&q=75&w=752) ## About the UCF101 action recognition dataset UCF101 is a human action recognition (HAR) dataset of realistic action videos, collected from YouTube, with 101 action categories. At the time of its release in 2012, it was the largest video action recognition dataset available to the research community. The dataset is made up of 13,320 videos (27 total hours) and is characterized by a large diversity of actions, as well as large variations in camera motion, object appearance and pose, object scale, viewpoint, cluttered background, and illumination conditions. Because at the time of its curation, most of the available action recognition datasets were not realistic or were staged by actors, UCF101 aimed to encourage further research into action recognition by learning and exploring new realistic action categories. The videos in the 101 action categories are grouped into 25 groups, where each group can consist of 4-7 videos of an action. Videos from the same group can share some common features, such as a similar background, viewpoint, etc. The action categories are divided into five types: - Human-Object Interaction - Body-Motion Only - Human-Human Interaction - Playing Musical Instruments - Sports **Note:** The UCF101 dataset is an extension of the [UCF50 Action Recognition Data Set](https://www.crcv.ucf.edu/data/UCF50.php), which has 50 action categories. https://www.youtube.com/watch?v=xArphgd\_hVs Video tutorial: How to get started with the UCF101 action recognition video dataset ## What is human action recognition? As you can imagine, it is very easy for a human to watch a video and recognize humans and the actions they are performing. However, having a machine do the same thing is a very challenging problem in the “video understanding” subfield of computer vision. More concretely, human action recognition for the purposes of this dataset is the problem of automatically assigning a video into one of the 101 different action categories. Human action recognition in video has a variety of real-world applications like surveillance (military, industrial and civilian), healthcare (for example: the monitoring of patients as they move around a facility), human-computer interaction, content-based video retrieval, and video summarization. ## **Dataset quick facts** - **Research Paper:** [UCF101: A Dataset of 101 Human Action Classes From Videos in The Wild](https://www.crcv.ucf.edu/wp-content/uploads/2019/03/UCF101_CRCV-TR-12-01.pdf) - **Authors:** Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah - **Download Dataset:** [RAR file](https://www.crcv.ucf.edu/datasets/human-actions/ucf101/UCF101.rar) - **Revised Annotations:** [Available on thumos.info](http://www.thumos.info/download.html) - **Action Recognition:** [ZIP file](https://www.crcv.ucf.edu/wp-content/uploads/2019/03/UCF101TrainTestSplits-RecognitionTask.zip) - **Action Detection:** [ZIP file](https://www.crcv.ucf.edu/wp-content/uploads/2019/03/UCF101TrainTestSplits-DetectionTask.zip) - **Video-Level Annotations:** [ZIP file](https://www.crcv.ucf.edu/wp-content/uploads/2019/06/Datasets_UCF101-VideoLevel.zip) - **STIP Features:** [Part 1](https://www.crcv.ucf.edu/datasets/human-actions/ucf101/UCF101_STIP_Part1.rar), [Part 2](https://www.crcv.ucf.edu/datasets/human-actions/ucf101/UCF101_STIP_Part2.rar) - **Dataset Size:** 6.48 GB - **Last Release:** 2012 - **FiftyOne Dataset Name:** `ucf101` - **Tags:** `video`, `action-recognition` - **Supported Splits:** `train`, `test` - **Zoo Dataset class:** [`UCF101Dataset`](https://docs.voxel51.com/api/fiftyone.zoo.datasets.base.html?highlight=ucf101dataset#fiftyone.zoo.datasets.base.UCF101Dataset) ## Step 1: Install FiftyOne If you don’t already have FiftyOne installed on your laptop, it takes just a few minutes! For example on macOS: - [Verify your version](https://docs.voxel51.com/getting_started/virtualenv.html#creating-a-virtual-environment-using-venv) of Python - Create and activate a [virtual environment](https://docs.voxel51.com/getting_started/virtualenv.html#creating-a-virtual-environment-using-venv) - [Install IPython](https://docs.voxel51.com/getting_started/troubleshooting.html#ipython-installation) (optional) - [Upgrade](https://docs.voxel51.com/getting_started/virtualenv.html#creating-a-virtual-environment-using-venv) your `Setuptools` - [Install FiftyOne](https://docs.voxel51.com/getting_started/install.html#installing-fiftyone) - [Install FFmpeg](https://docs.voxel51.com/getting_started/troubleshooting.html#videos-do-not-load-in-the-app) - Install a utility to uncompress .rar files (optional) ![](https://cdn.sanity.io/images/h6toihm1/production/0b42c9111ab2e6b9cf510251eff787b957c7f675-1728x1080.gif?auto=format&dpr=2&fit=max&q=75&w=1600) **Note:** In order to work with this video dataset you’ll need to have [FFmpeg](https://docs.voxel51.com/getting_started/troubleshooting.html#videos-do-not-load-in-the-app) installed. Also, if you don’t have a package already installed to uncompress the UCF101 dataset .rar files, you can install a utility to accomplish this (for example on Mac) using: ```python 1brew install rar ``` You may also need to restart your IPython kernel and/or authorize the rar app in your macOS settings for the utility to be recognized during the dataset import step. Or on Linux: ```python 1sudo apt install rar ``` Learn more about how to [get up and running with FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html) in the Docs. ## Step 2: Import the dataset Now that you have the dataset downloaded and FiftyOne installed, let’s import the dataset into FiftyOne and launch the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html). This should take just a few minutes and a few more lines of code. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3import fiftyone.utils.video as fouv 4 5dataset = foz.load_zoo_dataset("ucf101", split="test") 6 7# Re-encode source videos as H.264 MP4s so they can be viewed in the App 8fouv.reencode_videos(dataset) 9 10print(dataset.name) # ucf101-test 11 12session = fo.launch_app(dataset) ``` The last line in the code snippet will launch the FiftyOne App in your default browser. You should see the following initial view of the `test` dataset in the FiftyOne App: ![](https://cdn.sanity.io/images/h6toihm1/production/ea4d856fa7ac2f0dfb447497a82443d1e6b4b716-1369x976.png?auto=format&dpr=2&fit=max&q=75&w=1369) Your directory of videos in `~/fiftyone/ucf101` should have a `test` and `train` folder, plus an `info.json` file: ![](https://cdn.sanity.io/images/h6toihm1/production/1d453518f041beaf17ef60de36781667571165c7-140x108.png?auto=format&dpr=2&fit=max&q=75&w=140) Both the `test` and `train` folders will contain video files broken up into 101 action categories. ![](https://cdn.sanity.io/images/h6toihm1/production/149c812d088c0f6ec63051275c78e0ad9aa30b98-182x317.png?auto=format&dpr=2&fit=max&q=75&w=182) **Tip:** If you want to persist the dataset so you don’t have to repeat the re-encode process when you load it in your next session, add the following to your initial load command: ```python 1dataset.persistent = True ``` Now, you can load the dataset quickly and launch the App in your next session. ```python 1import fiftyone as fo 2 3dataset = fo.load_dataset("ucf101-test") 4session = fo.launch_app(dataset) ``` Ok, let’s do a quick exploration of the UCF101 dataset! ## Sample details Click on any of the samples to get additional detail like tags, metadata, labels, frame labels and primitives. ![](https://cdn.sanity.io/images/h6toihm1/production/39c7976be5a8892844f7418aedfba8bf5b3b6f8d-257x257.jpg?auto=format&dpr=2&fit=max&q=75&w=257) ## Filtering by ID FiftyOne makes it very easy to filter the samples to find the ones that meet your specific criteria. For example we can filter by a specific `id`: ![](https://cdn.sanity.io/images/h6toihm1/production/bae53b9e6fb387bbb9bb262108e620189293e606-804x977.gif?auto=format&dpr=2&fit=max&q=75&w=804) ## Filtering by label In this example we filter the samples by the `SkateBoarding` action category. ![](https://cdn.sanity.io/images/h6toihm1/production/f11eb69116dd4002c99683f9b2a51ce20fc17dfa-804x420.gif?auto=format&dpr=2&fit=max&q=75&w=804) ## Start working with the dataset Now that you have a general idea of what the dataset contains, you can start using FiftyOne to perform a variety tasks including: - [Creating dataset views](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) - [Creating aggregations](https://voxel51.com/docs/fiftyone/user_guide/using_aggregations.html) - [Creating interactive plots](https://voxel51.com/docs/fiftyone/user_guide/plots.html) - [Annotating datasets](https://voxel51.com/docs/fiftyone/user_guide/annotation.html) - [Evaluating models](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html) You can also start making use of the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html) which provides powerful machine learning techniques you can apply to your workflows like visualizing embeddings, finding similarity, uniqueness and mistakenness. [action recognition](https://voxel51.com/blog/tag/action-recognition) [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [UCF101](https://voxel51.com/blog/tag/ucf101) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/33d08c7b16ab5bfa0e4c5a4f936be4592b8e0a90-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the Berkeley Deep Drive Autonomous Vehicle Dataset\\ \\ Datasets\\ \\ • \\ \\ Jan 11, 2023](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/129c3574861e6c307549e14106753af13ecfa2bf-1308x1044.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How to Download ActivityNet and Evaluate Video Understanding Models\\ \\ Datasets\\ \\ • \\ \\ Feb 8, 2022](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models) [![](https://cdn.sanity.io/images/h6toihm1/production/40381f5f37fa5fcd70eddca2f63b6710568f5d2c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual Kinship Recognition with the Families in the Wild Computer Vision Dataset\\ \\ Datasets\\ \\ • \\ \\ Dec 7, 2022](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-262-lllmstxt|> ## FiftyOne 0.19 Webinar Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Webinar Recap: What’s New in FiftyOne 0.19 for Computer Vision Mar 2, 2023 • 9 min read Article content In this article [What Is FiftyOne?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#8067fcc309e7) [What Are the New Features in FiftyOne 0.19?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#6f68bb2447b7) [1\. Spaces](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#d7c389f5c1e1) [2\. In-App Embeddings Visualization](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#259d878851bc) [3\. Saved Views](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#3700451a8d80) [4\. On-Disk Segmentations](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#6b97497bba98) [5\. New UI Filtering Options](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#af26ad108596) [Other Notes](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#7b17261c059c) [Q&A from the Webinar](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#f5ca1e7a8937) In this article [What Is FiftyOne?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#8067fcc309e7) [What Are the New Features in FiftyOne 0.19?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#6f68bb2447b7) [1\. Spaces](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#d7c389f5c1e1) [2\. In-App Embeddings Visualization](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#259d878851bc) [3\. Saved Views](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#3700451a8d80) [4\. On-Disk Segmentations](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#6b97497bba98) [5\. New UI Filtering Options](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#af26ad108596) [Other Notes](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#7b17261c059c) [Q&A from the Webinar](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#f5ca1e7a8937) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We recently released [FiftyOne 0.19](https://voxel51.com/blog/announcing-fiftyone-0-19/), which is packed with new features that make it even easier and faster to visualize your computer vision datasets and boost the performance of your machine learning models. How? Read on! Voxel51 Co-Founder and CTO [Brian Moore](https://www.linkedin.com/in/brimoor/) walked us through the new features in a live webinar, with plenty of live demos and code examples to show it in action. You can watch [the playback on YouTube](https://www.youtube.com/watch?v=6hlJO8wW0YU), take a look at [the slides](https://docs.google.com/presentation/d/1LIS9ASit1LG4_3q7cVii8c3ZffTsr1CQ7rzTX43i5Zg/edit?usp=sharing), read [the transcript](https://www.rev.com/transcript-editor/shared/y9m9UtyA3Y9R13xyk3w5KP6bLpE3jlRPx_XwDsT3MoIWCWLFLmvGh5B3W1OcMtRtY0eo00aTMdCaPVHZdzXYcTDFAeY?loadFrom=SharedLink), and read the recap below for the highlights. Enjoy! https://www.youtube.com/watch?v=6hlJO8wW0YU ## What Is FiftyOne? Brian starts with a quick overview of what [FiftyOne](https://voxel51.com/fiftyone/) is for those who might be new to it. It’s a data-centric open source toolset that enables you to: - Visualize, query, and analyze computer vision datasets - Streamline annotation workflows - Identify and correct labeling mistakes - Analyze model performance, both visually and programmatically … And dozens of additional workflows to help you curate high quality computer vision datasets and improve model performance. As a tool for visual data (images, videos, and 3D data), it supports all popular computer vision tasks, including: classification, detection, instance segmentation, polygons and polylines, keypoints, point clouds and annotations, geolocation, embeddings, and multiview datasets. Before diving into the fresh new features, Brian gives a quick shout out to a couple of other capabilities: - FiftyOne is more than just an application. It's also a very powerful Python API. This means you can move seamlessly between code and the UI. - FiftyOne comes with the Brain, a component that provides powerful machine learning techniques and helps you automatically surface potential issues in your datasets ## What Are the New Features in FiftyOne 0.19? Here are the new features in FiftyOne 0.19 that were demonstrated in the webinar and are explained in the sections below: 1. [**Spaces**](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#spaces): a new feature that enables you to organize panels of information within the FiftyOne App 2. [**In-App embeddings visualization**](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#in-app-embeddings): the ability to visualize embeddings directly in the App through a panel made possible by Spaces 3. [**Saved views**](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#saved-views): the ability to save views under a name of your choice and recall those later either in the App or through code 4. [**On-disk segmentations**](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#on-disk-segmentations): as of this release you can store semantic segmentation masks and heatmaps on disk, including as RGB images, rather than in the database 5. [**New UI filtering options**](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision#new-ui-filtering): we’re we're continually bringing new filtering options to the App’s sidebar, which now contains upgraded options for filtering datasets ## 1\. Spaces In FiftyOne, it was already possible to load different types of modalities like a map using the previous layout, which was similar to a persistent tab-based menu at the top. Those tabs have now been converted into Panels that you can arrange horizontally and vertically within your App window. You can drag Panels to reorganize them, and you can have multiple tabs open in a Panel to toggle between them. Brian provides an analogy for the Spaces feature: “it's like having a code editor for your data, so when you're looking at FiftyOne now, you can think more like VSCode, and less like a grid of images alone.” FiftyOne natively includes these Panel types: - [Samples panel](https://docs.voxel51.com/user_guide/app.html#app-samples-panel): the media grid that loads by default when you launch the App - [Histograms panel](https://docs.voxel51.com/user_guide/app.html#app-histograms-panel): a dashboard of histograms for the fields of your dataset - (Detailed below!) [Embeddings panel](https://docs.voxel51.com/user_guide/app.html#app-embeddings-panel): a canvas for working with [embeddings visualizations](https://docs.voxel51.com/user_guide/brain.html#brain-embeddings-visualization) - [Map panel](https://docs.voxel51.com/user_guide/app.html#app-map-panel): visualizes the geolocation data of datasets that have a [GeoLocation](https://docs.voxel51.com/api/fiftyone.core.labels.html#fiftyone.core.labels.GeoLocation) field Other nifty things to know about Spaces: You can configure custom Panels [via plugins](https://docs.voxel51.com/plugins/index.html#fiftyone-plugins). And, not only can you configure Spaces through the App, you can also configure them through code. Learn more about configuring Spaces from [~13:15 - 21:37](https://www.youtube.com/watch?v=6hlJO8wW0YU&t=795s) in the webinar replay, in the [release announcement blog post](https://voxel51.com/blog/announcing-fiftyone-0-19/#spaces), and in the [FiftyOne Spaces docs](https://docs.voxel51.com/user_guide/app.html#spaces). ![](https://cdn.sanity.io/images/h6toihm1/production/c1df31887bc375fa960d6e7a41e695f1404e32ec-2512x1384.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## 2\. In-App Embeddings Visualization Brian walks us through how to interactively explore embeddings scatterplots in the App through the Embeddings Panel mentioned above. The Embeddings Panel: - Supports any visualization generated by compute\_visualization() - Supports image and object-level embeddings - Enables you to lasso points in the plot to show only the corresponding samples/patches - Enables you to show/hide specific classes via the legend - Automatically updates when your view changes - Gracefully handles new/missing data To show us the Embeddings Panel in action, Brian demos a subset of the Berkeley Deep Drive dataset with embeddings he had already generated using the [`compute_visualization()` method](https://docs.voxel51.com/user_guide/brain.html#visualizing-embeddings) available in the FiftyOne Brain ahead of time. In the demo, he color codes the samples by time of day – red for night, pink for day. Looking at the Samples Panel and Embeddings Panel side by side enables you to pull out interesting findings. For example, Brian turns off all time-of-day labels except daytime labels, then lassos the samples that are labeled as daytime but clustered with nighttime samples in the Embeddings Panel, then explores the lassoed samples in the Samples Panel. Some samples are incorrectly labeled. Other samples represent driving through a tunnel so it is daytime, but visually darker. ![](https://cdn.sanity.io/images/h6toihm1/production/c22078cada6245eac44dfa3fe576456a31856a34-2510x1378.png?auto=format&dpr=2&fit=max&q=75&w=1600) How might you act on findings like these in FiftyOne? Brian describes some of the ways: “This could be a powerful workflow to identify label mistakes. If these were instead model predictions, it could be an equally interesting way to pick out particular samples that were classified incorrectly. And then maybe grab close samples, and use them as hard examples, especially if you're viewing a data set that contains a bunch of points, only some of which were in your training dataset.” Brian shows a few other examples from this dataset of exploring outliers in the embeddings clusters to find mislabeled samples (images with a highly visible dashboard, images in the rain, etc.). Finally, he shows a dataset with geolocation data in order to show the Samples Panel, Map Panel, and Embeddings Panel in a single view and all working in concert with each other as you filter and explore. Learn more about working with the Embeddings Panel from [~21:38 - 29:34](https://www.youtube.com/watch?v=6hlJO8wW0YU&t=1298s) in the webinar replay, in the [release announcement blog post](https://voxel51.com/blog/announcing-fiftyone-0-19/#in-app-embeddings), and in the [Embeddings Panel docs](https://docs.voxel51.com/user_guide/app.html#embeddings-panel). ## 3\. Saved Views Next, Brian explains the new saved views features. As of FiftyOne 0.19, you can save any view. Simply give your saved view a name, and then you can load that saved view later, either through the App or through code. What kind of workflows can you accomplish with saved views? Brian explains, “I've seen workflows where users tag data and then create a saved view that simply matches a certain tag, which would give you a quick way to pull up specific subsets of your data set encoded by tags. Or you might use saved views to remember what subset of your data set you trained a model on by creating the view, and then naming it the name of the model you're training. Maybe I want to use a saved view as a shorthand to pull up the samples in my dataset that were the nighttime images in the example before, or really anything else you can imagine.” In the demo, Brian shows how to save a view of samples in a dataset that are labeled as cats, sorted by the number of cats with the most cats shown first, and then load that saved view in the App. ![](https://cdn.sanity.io/images/h6toihm1/production/5e9ec0a28e9fd91966be4d2e2b025cdef897a55b-2512x1376.png?auto=format&dpr=2&fit=max&q=75&w=1600) You may notice in the screenshot above that the name of the saved view is in the URL bar in the App. This gives you a one-click way to pull up a specific subset of a data set. (Additionally, for anyone using [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) to securely collaborate on datasets, you could share this link with members of your team.) Get the details and demo on saved views in the [webinar replay from ~29:37 - 34:53](https://www.youtube.com/watch?v=6hlJO8wW0YU&t=1777s), in the [release announcement blog post](https://voxel51.com/blog/announcing-fiftyone-0-19/#saved-views), and also in the [Saving Views docs](https://docs.voxel51.com/user_guide/app.html#saving-views). ## 4\. On-Disk Segmentations Until now, it has been possible to work with semantic segmentations in FiftyOne. The previous way to do that resulted in storing segmentations in the database, which was convenient but not efficient for large datasets. However, many people prefer to work with segmentations stored on disk, so as of FiftyOne 0.19 you can! Brian walks us through some notable characteristics of the new on-disk segmentation feature: - Storing segmentations on disk is significantly more performant - RGB segmentation masks now also supported - Entire API upgraded to support on-disk and RGB segmentations - `evaluate_segmentations()` - `apply_model()` - `export_segmentations()` Did you know? Before showing on-disk segmentations, Brian gives a shout out to a powerful feature in FiftyOne you may not yet know about: FiftyOne provides a number of utility methods to convert between different representations of certain label types, such as converting between instance segmentations, semantic segmentations, and polylines – each implemented with a simple command. Learn more about [converting label types](https://docs.voxel51.com/user_guide/using_datasets.html#converting-label-types) in the docs. Now, to show on-disk segmentations, Brian first loads 25 samples from the COCO 2017 validation split. The dataset has samples with a field called instances (for instance segmentations), with detection objects and masks. Brian converts instance segmentations to semantic segmentations and passes the `output_dir` argument so that these are stored on disk as follows. ```python 1import fiftyone.utils.labels as foul 2 3# Convert instance segmentations to semantic segmentations stored on disk 4foul.objects_to_segmentations( 5 dataset, 6 "instances", 7 "segmentations", 8 output_dir="/tmp/segmentations", 9 mask_targets=mask_targets, 10) ``` Then Brian demonstrates the new segmentation instances on the dataset with a `mask_path` on disk. You can learn more about the new on-disk feature in the demo from [~38:19 - 42:05 in the webinar replay](https://www.youtube.com/watch?v=6hlJO8wW0YU&t=2299s), in the [release announcement blog post](https://voxel51.com/blog/announcing-fiftyone-0-19/#on-disk-segmentations), and in the docs: [instance segmentations](https://docs.voxel51.com/user_guide/using_datasets.html#instance-segmentations), [semantic segmentation](https://docs.voxel51.com/user_guide/using_datasets.html#semantic-segmentation), and [heatmaps](https://docs.voxel51.com/user_guide/using_datasets.html#heatmaps). ## 5\. New UI Filtering Options Next, Brian covers the new builtin UI filtering options added in the App’s sidebar as of FiftyOne 0.19: - Only show objects with the specified labels (omitting samples with no matching objects) - Exclude objects with the specified labels - Show samples that contain the specified labels (without filtering) - Omit samples that contain the specific labels All applicable filtering options are available from both the grid view and the sample modal, and for all field types, including top-level fields and dynamic label attributes. This has always been possible through code, but now you can do it through the App as well! To learn more about the new filtering options, check out the demo from [~42:05 - 45:30 in the webinar replay](https://www.youtube.com/watch?v=6hlJO8wW0YU&t=2525s) and in the [release announcement blog post](https://voxel51.com/blog/announcing-fiftyone-0-19/#new-ui-filtering-options). ## Other Notes After demonstrating the new features, Brian shares a few additional points before concluding the presentation. Open source software like FiftyOne doesn’t happen without an amazing community supporting it. Brian gives a shout out to the community members who contributed to FiftyOne 0.19. In addition to FiftyOne, Voxel51 also builds [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) that adds collaboration and security features built specifically for teams, including cloud-backed media, dataset permissions, versioning, sharing, and more – all on top of the goodness available in open source FiftyOne. ## Q&A from the Webinar **You mentioned an analogy: FiftyOne can be the pandas for computer vision. Can you elaborate?** Yes! While they apply to different types of data, the pandas DataFrame and FiftyOne Dataset classes share many similar functionalities. As a result, we prepared a side-by-side comparison of common operations in the two libraries in a [tutorial](https://docs.voxel51.com/tutorials/pandas_comparison.html), [blog post](https://voxel51.com/blog/why-fiftyone-is-the-pandas-of-computer-vision/), and [cheat sheet](https://docs.voxel51.com/cheat_sheets/pandas_vs_fiftyone.html). **Are there any plans to add support for text annotation visualizations for scene text recognition problems?** There are already FiftyOne users working on scene text recognition today! You can store StringFields on your samples or labels that you can use to store and visualize arbitrary text. One example is storing the bounding boxes as Detection labels, then storing the result of your text recognition model as a string attribute of your detection. Adding fields to sample: [https://docs.voxel51.com/user\_guide/using\_datasets.html#adding-fields-to-a-sample](https://docs.voxel51.com/user_guide/using_datasets.html#adding-fields-to-a-sample) Label attributes: [https://docs.voxel51.com/user\_guide/using\_datasets.html#labels](https://docs.voxel51.com/user_guide/using_datasets.html#labels) **Does FiftyOne have support for 3D images other than video? ( i.e. time is not one of the dimensions).** Yes we support [3D point clouds](https://docs.voxel51.com/user_guide/groups.html#point-cloud-slices), polylines, bounding boxes, etc. But to clarify: FiftyOne’s support for 3D point clouds does not support any specific visualizer for 3D volumetric images (e.g. MRI). For this, the current practice is to slice the volume according to the dimension of minimum size and render these into frames of a video. FiftyOne does support plugins based on media type. So one could envision adding a 3D volumetric imaging visualizer via a [custom plugin](https://docs.voxel51.com/plugins/index.html). **So, I could build my own visualization tool as a plugin? Cool!** Yes! You could write your own [custom plugin of any kind](https://docs.voxel51.com/plugins/index.html) and expose it directly in the FiftyOne App. If you need any help along the way, please reach out to us [in Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) individually (Jacob, Brian, anyone at Voxel51) or generally in the #help channel where we hang out and we would be happy to assist you. [embeddings](https://voxel51.com/blog/tag/embeddings) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne 0.19](https://voxel51.com/blog/tag/fiftyone-0-19) [FiftyOne App](https://voxel51.com/blog/tag/fiftyone-app) [heatmaps](https://voxel51.com/blog/tag/heatmaps) [saved views](https://voxel51.com/blog/tag/saved-views) [segmentations](https://voxel51.com/blog/tag/segmentations) [Spaces](https://voxel51.com/blog/tag/spaces) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/8651d28f4a2978eb72cca25ef09e1f01f81847ea-2968x2042.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.19 with Spaces, In-App Embeddings Visualization, Saved Views, and More!\\ \\ Product & News\\ \\ • \\ \\ Feb 16, 2023](https://voxel51.com/blog/announcing-fiftyone-0-19) [![](https://cdn.sanity.io/images/h6toihm1/production/308698a5aece1d5b1b95ee1bf52811b24448458c-1200x672.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne 0.18 for Computer Vision\\ \\ Event Recaps\\ \\ • \\ \\ Dec 6, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/17422a76c76f14096dce21e43da51945f498a811-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne & FiftyOne Teams\\ \\ Event Recaps\\ \\ • \\ \\ Oct 8, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-263-lllmstxt|> ## Exploring Open Images V7 [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) Exploring Google’s Open Images V7 Mar 8, 2023 • 8 min read Article content In this article In this article ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Google’s Open Images dataset just got a major upgrade. The dataset that gave us more than one million images with detection, segmentation, classification, and visual relationship annotations has added 22.6 million _point labels_ spanning 4171 classes. With [Open Images V7](https://ai.googleblog.com/2022/10/open-images-v7-now-featuring-point.html), Google researchers make a move towards a new paradigm for semantic segmentation: rather than densely labeling every pixel in an image, which leads to expensive and time consuming annotation, in [the accompanying paper](https://arxiv.org/pdf/2210.14142.pdf) they show that sparse annotations of a variety they dub _pointillism_ can lead to comparable model performance. This addition to Open Images reflects the ongoing shift in computer vision towards data-centric approaches. In this article, we’ll show you how to get started working with Open Images V7 and point labels, and explore some features of the dataset. For this exploration, we’ll be using the open source computer vision library [FiftyOne](https://voxel51.com/docs/fiftyone/), which is [one of the official download and visualization tools](https://storage.googleapis.com/openimages/web/download_v7.html#download-fiftyone) recommended by the Open Images team. For a deep-dive into Open Images V6, check out this [Medium article](https://medium.com/voxel51/loading-open-images-v6-and-custom-datasets-with-fiftyone-18b5334851c3) and [tutorial](https://voxel51.com/docs/fiftyone/tutorials/open_images.html). Keep reading for a look at point labels and how to navigate what’s new in Open Images V7! ### **Loading in the data** The easiest way to get started is to import FiftyOne and download [Open Images V7](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#open-images-v7) from the FiftyOne Dataset Zoo. By default, this will download (if necessary) all splits of the data — train, test, and validation — including all available label types for each, and the associated metadata. As with the Open Images V6 dataset in the FiftyOne Dataset Zoo, however, we can also specify what subsets of the data we would like to download and load! In this article, we’ll be working with the validation split, which consists of 41,620 images. We can download and load in this split in FiftyOne by passing in the `split` argument: If we only wanted to download a thousand images, for the purposes of getting a feel for the data, we could specify this with the `max_samples` argument: Let’s take a look at this data by launching the FiftyOne App, a powerful graphical user interface that enables you to visualize, browse, and interact directly with your datasets: \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop As we can see, there’s a lot going on here. Every one of these images has object bounding boxes, segmentation masks, classification labels, relationship labels, and point labels. However, not every image in the Open Images dataset is annotated with every one of these label types; some images, for instance, do not have point labels. The reason the data that we’ve loaded has all of these is that, by default, FiftyOne prioritizes downloading images that have as many label types as possible! If we only care about some of the types of annotations, we can isolate these in one of a few ways. We can use the sidebar on the left hand side to toggle label types on/off: \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Or we can create a view of the data using `select_labels()` to choose the label types we want, and then view this in the FiftyOne App: Another alternative is to explicitly specify which label types we want when loading the dataset: ### **What’s the point?** Now that we’ve loaded in the data, let’s get a better sense for what these point labels are, and what they actually mean. In the FiftyOne App, click on one of the images in the sample grid and hover over one of the points in the image. You’ll see a box that lists properties of that point label. In a moment, we’ll go through these one by one, but first, a crucial point (no pun intended) must be made. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop In generating this dataset, the creators set about asking yes/no questions about whether a given point corresponded to a given class label. This means that these point labels are _not_ labels for the classes they specify. Rather, the point labels are the class labels about which the yes/no questions were concerned. If you are familiar with previous versions of the Open Images dataset, this notion may be familiar to you, as the classification labels follow a similar pattern, with negative labels and positive labels. Point labels, however, are slightly more nuanced than classification labels. First off, different points received different numbers of total votes. Second, annotators were allowed to cast votes as “yes”, “no”, or “unsure” votes for the same class label. As a result, some point labels are estimated as unsure! One last difference is that these point labels were generated via two different methods. Some of the yes/no questions were answered by human annotators, while others were “answered” by a model which predicted a _candidate class_. **Properties** - `label`: the things/stuff class about which the yes/no question was asked - `yes_votes`: the number of “yes” votes cast by annotators - `no_votes`: the number of “no” votes cast by annotators - `unsure_votes`: the number of “unsure” votes cast by annotators - `estimated_yes_no`: the best guess for whether the class label fits the point, given the votes cast - `source`: the method used to evaluate (cast votes for) the class label, with “ih” for human annotators and “cc” for model-predicted candidate class - `id`: unique identifier within FiftyOne With the exception of the `id`, all of these properties are extracted directly from the Open Images raw data. For more details on these properties, see the [Open Images V7 paper](https://arxiv.org/pdf/2210.14142.pdf). In FiftyOne, these point labels are represented by [Keypoint Labels](https://voxel51.com/docs/fiftyone/user_guide/using_datasets.html#keypoints), which enable us to easily access, filter, and perform operations on these labels. For an individual sample, it is easy to read out this data: \[, , , \] We can also use FiftyOne’s [Aggregation](https://voxel51.com/docs/fiftyone/user_guide/using_aggregations.html) class and [filtering capabilities](https://voxel51.com/docs/fiftyone/user_guide/using_views.html#filtering) to quickly get some summary statistics about the dataset (the validation split that we downloaded in the first section). We can compute the total number of “yes”, “no” and “unsure” votes across all points: total YES votes: 1055849 total NO votes: 1700581 total UNSURE votes: 114893 And the number of times a point had a certain number of votes for these three options: YES vote counts: {0: 804909, 1: 450783, 2: 78768, 4: 132, 3: 148409, 5: 151, 6: 170} NO vote counts: {4: 113, 1: 376188, 3: 394390, 0: 642460, 6: 75, 5: 43, 2: 70053} UNSURE vote counts: {1: 114855, 2: 19, 0: 1368448} Using FiftyOne’s ViewField with `from fiftyone import ViewField as F`, we can also do things like count the total number of point labels: And efficiently compute the number of point labels that received at least one “yes” vote and at least one “no” vote: ### **Point(s) of interest** Now that we understand what these point labels are and how to access their basic properties, let’s discuss how to massage the data into a form that is potentially more useful for downstream processing. One thing we might want to do is extract a subset of the dataset with point labels for particular classes. As an example, suppose we are working on a wildlife conservation project and we want to train a model to identify turtles and tortoises. In FiftyOne there are multiple ways to accomplish this! If we already have the entire dataset loaded, we can filter the point labels for instances of “Turtle” or “Tortoise”, and either use `select_labels()` to get only the point labels and perhaps the positive classification labels, or we can toggle the rest of the labels off in the FiftyOne App, as we’ve done here: Alternatively, if you know from the start that you are only interested in certain label classes, you can pass this information directly into the `load_zoo_dataset()` method: Another thing we might want to do is turn the raw “yes”, “no”, and “unsure” votes that were cast for these point labels into something more concrete that we can use to train models. For simplicity’s sake, let’s say that rather than use the `estimated_yes_no` values given by the dataset’s authors, we want to generate “positive” and “negative” point labels in direct analogy with the classification labels. As an example workflow, we could imagine that we are only interested in point label votes cast by human annotators — not the candidate classes generated by models. And let’s further suppose that to ensure with high probability that our ground truth labels are accurate, we will only label points as “positive” or “negative” if there are at least two human votes cast, and all votes agree. Here’s one way to do this in FiftyOne: First, we will filter the point labels for points with multiple positive votes, no negative or unsure votes, and `source="ih"`, and clone this view into a new dataset `positive_dataset`, only keeping the `points` field: Then rename the embedded field `points` to `positive_points`, so that we can merge this into the original dataset shortly. We can then get rid of a bunch of the embedded fields within the Keypoint label, because we already know the source and number of “no” and “unsure” votes. For similar reasons, we can also rename the `yes_votes` field to `votes`: Finally, we can merge this `positive_dataset` into the original dataset: After going through an analogous procedure for the negative point labels, we can use `select_fields` one more time to create a view containing only the positive and negative point labels, and the positive and negative classification labels: \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ### **Conclusion** In this article, we’ve only scratched the surface of what you can do with FiftyOne and Google’s latest version of the Open Images dataset. If you want to dive deeper into features of the Open Images dataset, check out this [Medium article](https://medium.com/voxel51/loading-open-images-v6-and-custom-datasets-with-fiftyone-18b5334851c3) and [tutorial](https://voxel51.com/docs/fiftyone/tutorials/open_images.html), and if you want to know more about the pointillism-based approach to semantic segmentation, I encourage you to check out [Google’s paper on the topic](https://arxiv.org/pdf/2210.14142.pdf). I hope this article showed you how easy it is to get started working with Open Images V7 using FiftyOne! [classification](https://voxel51.com/blog/tag/classification) [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [Google](https://voxel51.com/blog/tag/google) [Keypoints](https://voxel51.com/blog/tag/keypoints) [object detection](https://voxel51.com/blog/tag/object-detection) [Open Images](https://voxel51.com/blog/tag/open-images) [Open Images V7](https://voxel51.com/blog/tag/open-images-v7) [point labels](https://voxel51.com/blog/tag/point-labels) [segmentations](https://voxel51.com/blog/tag/segmentations) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/40381f5f37fa5fcd70eddca2f63b6710568f5d2c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual Kinship Recognition with the Families in the Wild Computer Vision Dataset\\ \\ Datasets\\ \\ • \\ \\ Dec 7, 2022](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/33d08c7b16ab5bfa0e4c5a4f936be4592b8e0a90-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the Berkeley Deep Drive Autonomous Vehicle Dataset\\ \\ Datasets\\ \\ • \\ \\ Jan 11, 2023](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/129c3574861e6c307549e14106753af13ecfa2bf-1308x1044.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How to Download ActivityNet and Evaluate Video Understanding Models\\ \\ Datasets\\ \\ • \\ \\ Feb 8, 2022](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-264-lllmstxt|> ## Computer Vision in Manufacturing [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Industry Solutions](https://voxel51.com/blog/category/industry-solutions), [Product & News](https://voxel51.com/blog/category/product-news) How Computer Vision Is Changing Manufacturing Mar 9, 2023 • 16 min read Article content In this article [Industry overview](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#c6e3e31250eb) [Key industry challenges in manufacturing](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#55f08640700a) [Applications of computer vision in manufacturing](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#00bb7b773027) [Companies at the cutting edge of computer vision in manufacturing](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#d506c6742f9f) [Datasets](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#5005f36fd29b) [Join the FiftyOne community!](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#d23f165e52d4) In this article [Industry overview](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#c6e3e31250eb) [Key industry challenges in manufacturing](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#55f08640700a) [Applications of computer vision in manufacturing](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#00bb7b773027) [Companies at the cutting edge of computer vision in manufacturing](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#d506c6742f9f) [Datasets](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#5005f36fd29b) [Join the FiftyOne community!](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing#d23f165e52d4) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the second installment of [Voxel51](https://voxel51.com/)’s computer vision industry spotlight blog series. Each month, we highlight how different industries – from construction to climate tech, from retail to robotics, and more – are using computer vision, machine learning, and artificial intelligence to drive innovation. We’ll dive deep into the main computer vision tasks being put to use, current and future challenges, and companies at the forefront. In this edition, we’ll focus on _manufacturing_! Read on to learn about computer vision in manufacturing and industrial automation. ## Industry overview Key facts and figures: - [14.7 million Americans are employed in manufacturing](https://www.nist.gov/el/applied-economics-office/manufacturing/total-us-manufacturing#:~:text=In%202021%2C%20Manufacturing%20contributed%20%242.3,an%20estimated%2024%20%25%20of%20GDP.) - The industry contributed [$2.3 trillion to US GDP in 2021, accounting for 24%](https://www.nist.gov/el/applied-economics-office/manufacturing/total-us-manufacturing#:~:text=In%202021%2C%20Manufacturing%20contributed%20%242.3,an%20estimated%2024%20%25%20of%20GDP.) of the national total - [Industrial automation is expected to reach $359 billion in the United States by 2029](https://www.fortunebusinessinsights.com/industry-reports/industrial-automation-market-101589) - [Global manufacturing production is estimated to be north of $40 trillion](https://www.powermotiontech.com/news/article/21242721/global-manufacturing-production-to-reach-value-of-445-trillion-in-2022) - Manufacturing accounts for [16 percent of global GDP and 14 percent of employment](https://www.mckinsey.com/capabilities/operations/our-insights/the-future-of-manufacturing) Manufacturing is integral in the production of most of the physical goods that make up our modern world, from food and beverages to laptops, toys, and tools. The sector is also a critical component for both developing and developed economies, driving technological advancements and offering employment to millions. The manufacturing industry is currently undergoing a massive transformation referred to as the [fourth industrial revolution](https://en.wikipedia.org/wiki/Fourth_Industrial_Revolution) (4IR), headlined by the adoption of computer vision, artificial intelligence, robotics, and the [industrial internet of things](https://www.ptc.com/en/technologies/iiot) (IIoT). 4IR technologies present an estimated [multi-trillion dollar opportunity](https://www.mckinsey.com/capabilities/operations/our-insights/industrys-fast-mover-advantage-enterprise-value-from-digital-factories) and will enable factories to operate more accurately, efficiently, and safely. Computer vision is already playing a central role in this transformation, with companies using machine vision techniques to automate strenuous tasks, identify issues in both product and machinery, and improve safety conditions for workers. Before we dive into several popular applications of computer vision-based AI technologies in manufacturing, here are some of the industry’s key challenges. ## Key industry challenges in manufacturing - Labor shortages: [According to projections by Deloitte](https://www2.deloitte.com/us/en/insights/industry/manufacturing/manufacturing-industry-diversity.html), there will be more than two million unfilled manufacturing jobs in the United States by 2030. Globally, the industry is expected to reach a deficit of up to 7.9 million jobs over the same period, [according to a study by Korn-Ferry](https://www.kornferry.com/content/dam/kornferry/docs/pdfs/KF-Future-of-Work-Talent-Crunch-Report.pdf). - Inflation: [Rising costs for materials, energy, and transportation](https://www.macny.org/manufacturers-feeling-the-impacts-of-inflation-and-supply-chain-challenges/) put pressure on manufacturers to streamline their operations. - Scope and diversity: By its very nature, manufacturing touches almost every product we encounter, from screws, to shoes, to cars. Each product has its own manufacturing requirements and challenges, and each factory uses its own combination of cameras and sensors, so there are no one-size-fits all solutions. Continue reading for some ways computer vision is enabling industrial automation and improving safety in manufacturing. ## Applications of computer vision in manufacturing ### Bin picking A common industrial robotics application, [_bin picking_](https://en.wikipedia.org/wiki/Bin_picking) is the action of selecting an object from a bin, picking it up, and placing it in another location. For a robot to be successful with this task, it needs to precisely navigate and interact with objects of different shapes, sizes, and materials in a potentially cluttered, occluded, and poorly lit environment. Machine vision systems make this possible by mapping the environment and guiding the robotic arm’s motion. In the simplest cases, camera images are passed into object detection routines. In many cases, however, grabbing an object from a bin requires mapping depth information and performing 3D object detection to situate the object in three dimensional space. Some computer vision systems use point clouds from LiDAR sensors to generate these 3D representations. One additional complication is that some objects can only be easily grabbed and held in certain orientations. To overcome this, some machine vision systems [estimate the “pose” of objects](https://paperswithcode.com/task/pose-estimation) and use this information to orient the robotic arm for picking. Here are a few papers on computer vision in bin picking: - [Bin-Picking for Planar Objects Based on a Deep Learning Network: A Case Study of USB Packs](https://www.mdpi.com/1424-8220/19/16/3602) - [Occlusion, Clutter, and Illumination Invariant Object Recognition](https://citeseerx.ist.psu.edu/document?repid=rep1&type=pdf&doi=3436fffa61474e5fbfc5db547367d303c7cc02ee) - [Fast Object Localization and Pose Estimation in Heavy Clutter for Robotic Bin Picking](https://www.merl.com/publications/docs/TR2012-007.pdf) - [Large-scale 6D Object Pose Estimation Dataset for Industrial Bin-Picking](https://ieeexplore.ieee.org/document/8967594) ### Palletizing and depalletizing Pallets have been called the unsung heroes of our modern age. These flat, typically wooden, platforms are essential to the scale and economy of global logistics and transportation. Before pallets make their way to shipping containers, they are loaded up with goods. These stacks of goods can reach [up to 15 feet tall](https://igps.net/blog/2018/09/18/how-to-stack-empty-pallets-safely/) and [weigh more than two tons](https://www.gigacalculator.com/calculators/pallet-calculator.php). The process of loading and stacking products onto a pallet is known as _palletizing_. Similarly, once products have reached their destination, they must be unloaded from the pallet. This unloading process is referred to as _depalletizing_. To reduce injury and error, manufacturers have been automating palletizing and depalletizing tasks through the use of robotic arms equipped with computer vision systems. Computer vision is a great fit for this problem, as the items being loaded and unloaded are typically nearly similar, so object detection models can be trained with very high accuracy. Another crucial enabling element is calibration. As the robot arm loads or unloads items, it takes note of the disparity between where it estimated the object to be located and where it was actually located as feedback to refine its future predictions. A few resources to get you started: - [3D-Computer Vision for Automation of Logistic Processes](https://link.springer.com/chapter/10.1007/978-3-319-01378-7_5) - [Automated detection of euro pallet loads by interpreting PMD camera depth images](https://link.springer.com/article/10.1007/s12159-012-0095-8) - [Toward Future Automatic Warehouses: An Autonomous Depalletizing System Based on Mobile Manipulation and 3D Perception](https://www.mdpi.com/2076-3417/11/13/5959) - [Robotic de-palletizing using uncalibrated vision and 3D laser-assisted image analysis](https://ieeexplore.ieee.org/abstract/document/5354054) - [Robust Pallet Detection for Automated Logistics Operations](https://www.scitepress.org/papers/2016/56747/56747.pdf) ### Machine tending _Machine tending_ is the process of automatically loading raw material or components for input into a machine. This includes placing parts on a conveyor belt, as well as preparing materials for welding, grinding, milling, or injection molds. Automated machine tending has multiple advantages, from reducing injury risk to improving consistency. Computer vision-enabled automation in machine tending also allows for higher precision in these applications. With real-time monitoring, [object localization](https://paperswithcode.com/task/object-localization) can be used to precisely situate the input materials relative to the machine being tended, and a robotic arm can adjust positioning accordingly. As with many computer vision applications in manufacturing, data availability and quality in machine tending is a significant challenge. Machine vision systems for machine tending often need to be built with limited labeled training data. This means that data cleaning and curation are essential, and techniques like data augmentation and transfer learning can be incredibly important. Here are a few preliminary resources: - [An Intelligent Manufacturing Approach Based on a Novel Deep Learning Method for Automatic Machine and Working Status Recognition](https://www.mdpi.com/2076-3417/12/11/5697) - [Vision-Based Associative Robotic Recognition of Working Status in Autonomous Manufacturing Environment](https://pdf.sciencedirectassets.com/282173/1-s2.0-S2212827121X0011X/1-s2.0-S2212827121011574/main.pdf?X-Amz-Security-Token=IQoJb3JpZ2luX2VjEEMaCXVzLWVhc3QtMSJHMEUCIGXhlAGFj%2FyW9h%2Ft4mMovZwJvicNyjOMDIjx1ZALpK9UAiEA%2BsnTuSCVGtgt1J6WHgns4HOEfPW7rE0y%2BynOsA%2FmFjQq1QQI3P%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FARAFGgwwNTkwMDM1NDY4NjUiDObBXkVbYODaiabeqSqpBOdK9Z%2BqgDYHD%2FfN2Qdxtwonr2mJQqGTu%2FASwcvYyKnUodRM%2FhpH1PmLOkY4PDLy6SVXWiZdCavelQvDR%2FOQu1cLjzGdo7x1NGaJzjFUfrTTDf2%2FVezda%2Fio9BS%2F2Nm9OdqTu0HhJmjBDbTE0vXguJO7y6iiuMeAbBfGfNRLljSKKgQNISTB%2FtVvHMsRkcfcmY4izzTCCkPtpI302IWEU3SP87CBy5hEBcT%2FoYABhCkA0mSDcS0JvrZWu3jHgDWTaRg9uopQRrO%2BLlZMFbggIjT87f%2BDgBRryfTMBeh0HY5VcCd4YoTCthibtvLZ7FHd4dOUaOfSqdpRxaa1ljI1OdY3t0hAgEweWk%2B5%2BWwRBkvYqt62dy33yTMS6aJOt8n8CjUQDEli%2FsykLg8gRkjIhcliispYAh6Z6ctuwz6HCBxUPBxRr%2F%2FrTMdq5b99JuS2KlanaZxiTMuW7x2qDK38yEnByKwSBy7VDBD91GCi%2FVX5Gu%2Fzbt5IUt2VU9fac4I5PBQvLFmg%2FepVFow7xZQ3di%2BY6X74LAU%2FxqWcJWK5kWNwyYHaMSnvBAL0Zkxd1BEwzxCxEV7L56lIPS0o0p%2BB3THd9f3OohTBeiTr8n59vsOS8sgQCqq3dZP%2BraZdQ3ikzSQEEXDWmzRW2yBD4i7LulIwZo3%2F5Gbw1JkltEXaFdYMfsLNwIxcmkexYUYYc%2BmJFR0i8lPVJjo4odrrGqSwQAO2nha4Uzk7sX8wxr7ZnwY6qQG2r%2BObK%2FXl6kBscM3s3DTClW%2BygFSpAJ8KuI1M8SQNQTGTjJ1aq64Q52W8zjWpFGB7E1qkhflSPbSGKQhfybzrIOzikCRReHgAdFwwav4ny9cZ0Ifr%2FDtf6xq9qU4JQ9jCGox9utCUT5VbWM4Wjy3%2BF9LinfS6nstL%2BgQhQhO6EhTj%2BftG3lXrewOnFJ60nBwUS%2FwArl5AnIFGGqlco481vZyv2CSJ2cQf&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Date=20230222T191740Z&X-Amz-SignedHeaders=host&X-Amz-Expires=300&X-Amz-Credential=ASIAQ3PHCVTYTCU2DNOK%2F20230222%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Signature=ac9f1097e9e14d4fa12ac25e547d1fb8af096f45156b06bd6a21774715c860e4&hash=29f90f7f80bc78d0793432626e73f52cdf5eaab757fa9d69d370c12303b41dc0&host=68042c943591013ac2b2430a89b270f6af2c76d8dfd086a07176afe7c76c2c61&pii=S2212827121011574&tid=spdf-0dce9604-1711-4baa-af20-c380786f2573&sid=a24d03dd5bf8b3414d6ae3646965868a94degxrqa&type=client&tsoh=d3d3LnNjaWVuY2VkaXJlY3QuY29t&ua=0f1650580a065250530402&rr=79da102cfb09ce98&cc=us) ### Defect detection Computer vision has already become indispensable in ensuring quality control in industrial processes. Instance segmentation, for instance, is used in conjunction with high-resolution sensor data, to check if a manufactured part has the desired spatial dimensions, within an allowed tolerance. One area of quality control where computer vision features prominently is defect detection. Manufacturers want to identify defective parts and products as early in their journey as possible. In some applications, where the possible varieties of defects are known, object detection and classification are used to identify problems. In other cases, the full variety of possible defects is either unknown, or too complicated to easily categorize. In manufacturing, defects can range from minutiae like small scratches to entirely missing components, like a missing screw. This can also be exacerbated by class imbalance, where there are many more examples of “normal” products than “defective” products. [Anomaly detection](https://en.wikipedia.org/wiki/Anomaly_detection) provides an alternative, [unsupervised approach](https://en.wikipedia.org/wiki/Unsupervised_learning) that takes normal and defective examples as input and predicts a new product as normal, or “nominal”, if it is likely to have come from the same _distribution_ as the previously seen normal samples. To make this determination, anomaly detection models learn an approximate representation of this nominal distribution, which may involve using density-based models like [DBSCAN](https://en.wikipedia.org/wiki/DBSCAN), [support vector machines](https://towardsdatascience.com/support-vector-machine-introduction-to-machine-learning-algorithms-934a444fca47), or deep learning models like [autoencoders](https://en.wikipedia.org/wiki/Autoencoder#:~:text=An%20autoencoder%20is%20a%20type,data%20from%20the%20encoded%20representation.) and generative adversarial networks ( [GANs](https://blog.paperspace.com/complete-guide-to-gans/)). Libraries like Intel’s [Anomalib](https://github.com/openvinotoolkit/anomalib) provide tools for implementing and benchmarking anomaly detection algorithms. Here are a few papers on defect detection and anomaly detection in manufacturing: - [Using Deep Learning to Detect Defects in Manufacturing: A Comprehensive Survey and Current Challenges](https://www.mdpi.com/1996-1944/13/24/5755) - [Automatic Detection and Classification of Manufacturing Defects in Metal Boxes using Deep Neural Networks](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0203192) - [Anomaly detection with convolutional neural networks for industrial surface inspection](https://reader.elsevier.com/reader/sd/pii/S2212827119302409?token=0AE42B55604D8922A34AE64006FF4185B10B8F72E4441EBE7B7E96F25163A51070ABEB86ABD1271204ADFE2B2DC391CB&originRegion=us-east-1&originCreation=20230214052745) - [Towards Total Recall in Industrial Anomaly Detection](https://openaccess.thecvf.com/content/CVPR2022/papers/Roth_Towards_Total_Recall_in_Industrial_Anomaly_Detection_CVPR_2022_paper.pdf) - [The MVTec Anomaly Detection Dataset: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection](https://ieeexplore.ieee.org/document/8954181) ### Predictive maintenance Tools and machinery in manufacturing plants gradually accumulate wear and tear which, if left untreated, could lead to complete breakdown, as well as potential injuries and lost productivity. To avoid such failure, manufacturers have historically performed _routine maintenance_, cleaning, refurbishing, and changing out old parts on a regular basis. But routine maintenance has two predominant downsides: first, it introduces downtime and costly upkeep when maintenance may not be necessary; second, it is not able to pick up on issues that arise and escalate between maintenance intervals. Preventive maintenance (PdM) seeks to address these issues with more active monitoring. Computer vision and predictive analytics help manufacturers to save on maintenance costs and avert catastrophe. PdM is already applied to a wide variety of machines and mechanical parts, from blades and bearings to gears and gaskets. In some cases, as with the saw blade pictured above, computer vision techniques like segmentation, object detection, and classification suffice to identify and predict all likely failure modes. Common signs include cracks, corrosion, and leaks. As with defect detection, when the spectrum of failure modes is complex or unknown, anomaly detection is applied to the state of the machines. - [Review of tool condition monitoring in machining and opportunities for deep learning](https://link.springer.com/article/10.1007/s00170-020-05449-w) - [A computer vision system for saw blade condition monitoring](https://www.sciencedirect.com/science/article/pii/S2212827121010842) - [A machine vision system for micro-milling tool condition monitoring](https://www.sciencedirect.com/science/article/abs/pii/S0141635917302817) - [Deep learning models for predictive maintenance: a survey, comparison, challenges and prospect](https://arxiv.org/abs/2010.03207) ## Companies at the cutting edge of computer vision in manufacturing ### Mech-Mind Robotics With more than 700 employees, 1000+ customers, and more than $200M in funding, Mech-Mind Robotics is the largest 3D vision company in China, and one of the world’s largest providers of 3D vision cameras and machine vision software for robotic automation. Mech-Mind’s integrated hardware and software solutions are used in a wide range of manufacturing applications, including bin picking, machine tending, palletizing and depalletizing, assembly, and gluing. The [Mech-Eye](https://www.mech-mind.com/product/mech-eye-industrial-3d-camera.html) industrial 3D camera uses [structured light technology](https://en.wikipedia.org/wiki/Structured-light_3D_scanner) to generate high resolution, high accuracy point clouds. Mech-Mind’s [Mech-Vision machine vision software](https://www.mech-mind.com/product/mech-vision-machine-vision-software.html) provides a platform for customers to build industrial computer vision applications with Mech-Eye cameras and customers’ robots. Mech-Vision has built-in support for common computer vision tasks like pose adjustment, and 2D and 3D matching, wherein the customer generates a point cloud model of the object to be recognized, either from a CAD file or directly from a camera image, and this model is recognized in the scene. Built-in support for common industrial robots means that the [robot calibration](https://en.wikipedia.org/wiki/Robot_calibration) process, [while traditionally time consuming](https://www.sciencedirect.com/science/article/abs/pii/S0736584506000743) (on the scale of hours), can be completed in less than 20 minutes. To round things out, [Mech-Mind's deep learning software](https://www.mech-mind.com/product/mech-dlk-deep-learning-software.html) allows customers to fine-tune computer vision models for their specific use cases. Customers load in their own data, which are automatically pre-labeled, and can then be rapidly edited and revised. Typically, Mech-Mind’s deep learning software only needs 20-50 images of an object to train a model to recognize it in a scene. ### Instrumental Founded by two Stanford and MIT grads and former Apple employees in 2015 and based in Palo Alto, CA, [Instrumental](https://instrumental.com/) is leading the way in ensuring product quality in electronics manufacturing. They use computer vision in conjunction with predictive analytics to provide real-time monitoring and alerts, as well as root cause analysis for prior failures. Instrumental's AI-based computer vision suite supports both new product introduction (NPI) manufacturing, which is characterized by low volume, and mass production (MP) manufacturing, which is high volume. Even within electronics manufacturing, the wide variety of products and substantial variation from product to product means that general purpose computer vision models have seen very little success. Nevertheless, manufacturers want to detect defects and issues on their particular use case as quickly as possible. Instrumental's suite of computer vision tools is designed to achieve high performance application-specific defect detection given as few samples as possible. To do this, they use techniques like data augmentation, transfer learning, and active learning to build a robust dataset that they use to train an anomaly detection model. Their models are built easily with no coding. Once deployed, these models run real-time inference on the edge in the factory and create a record that can be shared, inspected, and evaluated. ### Protex AI Founded in 2020 and backed by YCombinator, Notion Capital, and Playfair Capital, Irish startup [Protex AI](https://www.protex.ai/) is helping enterprise safety teams to revolutionize how they make proactive safety decisions that contribute to a safer work environment. Their AI-powered technology is enabling businesses to gain greater visibility of unsafe behaviors in their facilities. The privacy-preserving platform plugs into existing CCTV infrastructure and uses its computer vision technologies to capture unsafe events autonomously in settings such as warehouses, manufacturing facilities, and ports. Protex AI provides a simple interface so that each user can create their own “rules”, including setting exclusion zones, speed limits for forklifts, or even minimum distances workers must maintain between themselves and machines. Protex then uses computer vision techniques, including object detection, object tracking, and pose estimation, in order to check these rules. For rules involving speeds or distances, the vision system employs calibration. Typically, calibration is performed using inputs from multiple cameras, but Protex uses special routines to estimate calibration from a sole CCTV camera. Due to privacy concerns surrounding customer image and video data, Protex AI runs all of their models on the edge on Nvidia powered devices. For power and compute efficiency, they [quantize](https://towardsdatascience.com/how-to-accelerate-and-compress-neural-networks-with-quantization-edfbbabb6af7) their model weights. As use cases can differ greatly, Protex AI deploys custom models for each customer. Their base model is trained on hundreds of thousands of images, and then a unique version is fine-tuned on a given customer’s data. In their line of work, data quantity is not an issue. The most important factor in model performance is having a clean, high quality dataset. ### Cognex More than forty years old but still on the cutting edge, Nasdaq-listed (CGNX) [Cognex](https://www.cognex.com/) is a world leader in machine vision for industrial automation. Their 2000+ employee team has a hand in almost every step of industrial automation processes, from sensors and barcode scanners to industrial cameras and fully integrated vision systems. Cognex has machine vision tools for rule-based applications, such as monitoring object location and detecting edges, as well as deep learning tools for cloud connected and edge devices. Their [VisionPro Deep Learning](https://www.cognex.com/products/machine-vision/vision-software/visionpro-deep-learning) software supports standard tasks like defect detection and segmentation, and assembly verification, as well as burgeoning tasks like [material classification](https://link.springer.com/article/10.1007/s10845-019-01508-6). Beyond specific tasks, Cognex’s VisionPro software expedites time to deployment with [AutoML](https://en.wikipedia.org/wiki/Automated_machine_learning) capabilities. The _label checker_ automatically verifies the vast majority of labels and flags the remaining images for manual review, minimizing the number of samples a user needs to assess. During training, _parameter autotune_ will use input example images to determine the optimal set of hyperparameters. In [optical character recognition](https://en.wikipedia.org/wiki/Optical_character_recognition) (OCR), for instance, it can be difficult to recognize text due to the wide spectrum of fonts and potential distortions. Traditional OCR systems require that users specify segmentation hyperparameters to achieve high precision and recall. Cognex [Blue Read](https://www.cognex.com/products/machine-vision/vision-tools/ai-tools/deep-learning-tools/blue-read) eliminates this requirement by comparing an input image to the library of hundreds of fonts on which it was trained, and automatically selecting the best hyperparameters. ### RIOS Intelligent Machines [RIOS Intelligent Machines](https://www.rios.ai/) is on a mission to transform labor-intensive factories into smart factories powered by robotics and AI. The company helps its global customers automate their factories, warehouses, and supply chain operations by deploying AI-powered end-to-end robotic workcells that integrate within existing workflows. The Menlo Park, CA-based company was founded by former Xerox PARC engineers who saw a massive failure of traditional robots and predicted that factories over reliance on labor would soon reach a breaking point. RIOS has developed some of the most advanced hardware and AI/software platforms in robotics, including human-like tactile sensors for robots, haptics intelligence platform, and highest performance end-of-arm tooling and food-grade grippers. Their AI Controlled Robotics platform delivers fixed, programmable, flexible and integrated automation. They also offer palletizing robots to load and unload products on or off of pallets, plus robotic packaging systems. ### A few more It’s impossible to highlight every company doing amazing work at the intersection of computer vision and manufacturing and industrial automation. Here are a few more companies that are pushing the boundaries: - [PreML GmBH](https://www.preml.io/): German startup founded in 2020 focused on automated visual quality inspection. - [Prophesee](https://www.prophesee.ai/about-prophesee/): French series C startup pioneering [event-based vision](https://www.prophesee.ai/2019/07/28/event-based-vision-2/). - [Datalogic](https://www.datalogic.com/eng/index.html): Italy-based leader in automated data capture, barcode readers, sensors, and vision systems. - [Stemmer Imaging](https://www.stemmer-imaging.com/): Based in Germany, S9I is Europe’s largest imaging technology provider, with a hand in everything from photography to factory floor vision systems. - [Pickit 3D](https://www.pickit3d.com/en/): 2016 spinout of NASA robotics software provider [Intermodalics](https://www.intermodalics.eu/), focused on 3D vision systems for robotic guidance. - [Matroid](https://www.matroid.com/): End-to-end no-code computer vision solutions for quality assurance, assembly verification, and safety and compliance founded by Stanford adjunct professor [Reza Zadeh](https://www.linkedin.com/in/rezab/). ## Datasets Due to the highly proprietary nature of manufacturing and industrial automation processes, public computer vision datasets are few and far between. Hopefully these datasets will help you get started: - [MVTec Datasets for anomaly detection and industrial object detection](https://www.mvtec.com/company/research/datasets) (explore [MVTec AD and its embeddings](https://try.fiftyone.ai/datasets/mvtec-ad/samples) instantly in your browser in FiftyOne!) - [CPPE-5: Medical Personal Protective Equipment Dataset](https://github.com/Rishit-dagli/CPPE-Dataset/) - [Roboflow Forklift and Pallet Detection Dataset](https://universe.roboflow.com/phantom/forklift-1) - [MetaGraspNet: A Large-Scale Benchmark Dataset for Scene-Aware Ambidextrous Bin Picking via Physics-based Metaverse Synthesis](https://github.com/maximiliangilles/MetaGraspNet) - [Kolektor Surface Defect Dataset (KSSD)](https://www.vicos.si/resources/kolektorsdd/) - [beanTech Anomaly Detection Dataset (BTAD)](https://paperswithcode.com/dataset/btad) If you would like to see any of these, or other computer vision manufacturing datasets added to the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html), get in touch and we can work together to make this happen! ## Join the FiftyOne community! Developers of manufacturing and industrial automation applications can benefit from FiftyOne’s ability to easily filter through the huge amounts of visual data collected daily from farms and other sources. Using FiftyOne, this data can be curated into datasets for model training, or to share with experts for annotation or analysis of CV models. Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 2,700+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 6,700+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 23,000+ [Meetup members](https://voxel51.com/computer-vision-ai-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 540+ repositories - 90+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [anomaly detection](https://voxel51.com/blog/tag/anomaly-detection) [bin picking](https://voxel51.com/blog/tag/bin-picking) [Cognex](https://voxel51.com/blog/tag/cognex) [datalogic](https://voxel51.com/blog/tag/datalogic) [defect detection](https://voxel51.com/blog/tag/defect-detection) [depalletizing](https://voxel51.com/blog/tag/depalletizing) [industrial automation](https://voxel51.com/blog/tag/industrial-automation) [industry spotlight](https://voxel51.com/blog/tag/industry-spotlight) [instrumental AI](https://voxel51.com/blog/tag/instrumental-ai) [machine tending](https://voxel51.com/blog/tag/machine-tending) [machine vision](https://voxel51.com/blog/tag/machine-vision) [manufacturing](https://voxel51.com/blog/tag/manufacturing) [matroid](https://voxel51.com/blog/tag/matroid) [mech-mind robotics](https://voxel51.com/blog/tag/mech-mind-robotics) [palletizing](https://voxel51.com/blog/tag/palletizing) [pickit 3D](https://voxel51.com/blog/tag/pickit-3d) [predictive maintenance](https://voxel51.com/blog/tag/predictive-maintenance) [preML GmBH](https://voxel51.com/blog/tag/preml-gmbh) [prophesee](https://voxel51.com/blog/tag/prophesee) [protex AI](https://voxel51.com/blog/tag/protex-ai) [RIOS Intelligent Machines](https://voxel51.com/blog/tag/rios-intelligent-machines) [stemmer imaging](https://voxel51.com/blog/tag/stemmer-imaging) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/55292ec1fada0552df3d72fb684759407b67f5bc-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Visual AI in Manufacturing: 2025 Landscape\\ \\ Industry Solutions\\ \\ • \\ \\ Jul 16, 2025](https://voxel51.com/blog/visual-ai-in-manufacturing-2025-landscape) [![](https://cdn.sanity.io/images/h6toihm1/production/63f2810f51764c0187b30ef9f0640ba49b96d70c-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Computer Vision in Agriculture is the Future\\ \\ Industry Solutions, Product & News\\ \\ • \\ \\ Jan 31, 2023](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/99dc871d875865fc200f14931a3cfb3118ded624-1020x1007.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Tunnel vision in computer vision: can ChatGPT see?\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 16, 2022](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-265-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – Mar 10, 2023 Mar 11, 2023 • 6 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#f4f3f5a639b0) [Adding metadata to a FiftyOne dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#9c99e53b3b1e) [Changing tags when loading CVAT annotations into FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#e3e840a3bf5f) [Previewing video frames in FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#a07ba565841a) [Frame-level aggregations](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#d61dc7b1c7ea) [Deleting duplicate samples](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#3d6d41834c1f) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#ace8b17b18b9) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#f4f3f5a639b0) [Adding metadata to a FiftyOne dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#9c99e53b3b1e) [Changing tags when loading CVAT annotations into FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#e3e840a3bf5f) [Previewing video frames in FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#a07ba565841a) [Frame-level aggregations](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#d61dc7b1c7ea) [Deleting duplicate samples](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#3d6d41834c1f) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023#ace8b17b18b9) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **Adding metadata to a FiftyOne dataset** Community Slack member Immanuel Weber asked, _“Hi everyone, first let me say that I really love FiftyOne! I want to add metadata to my samples, and using `set_values()` I was able to add metadata, but with this approach the fields do not show up in the FiftyOne App. Is my approach a valid way to add new metadata? Is there a way to have the FiftyOne App display these new fields? Thanks!”_ Great question, Immanuel. In FiftyOne, what fields are visible in the FiftyOne App is determined by the [_schema_](https://docs.voxel51.com/user_guide/using_datasets.html#field-schemas) of your dataset. Only fields that are part of the schema show up in the App. You can print out `dataset.get_field_schema()` to see your field schema. As of [FiftyOne 0.19.0](https://docs.voxel51.com/release-notes.html#fiftyone-0-19-0), when you use `set_values()` to add a field to the samples in your dataset, it has an argument `dynamic` which you can use to control whether or not the field is added to your schema. By default, `dynamic=False`, so the field is _not_ added. If you pass in `dynamic=True`, then it will be added to the schema, and so it will show up in the FiftyOne App. Additionally, while it is possible to add to the `metadata` field, we strongly recommend creating a separate field on your samples for whatever attribute you want to store. This is because the `compute_metadata()` method, which computes height and width for each image in a dataset, will not function as desired if the metadata is not empty, which could lead to issues downstream. You can still create more complicated, nested objects, in new `EmbeddedDocumentField` fields, and have them show up in the FiftyOne App. For instance, if you wanted to create a new `custom_metadata` field with an embedded field that stores the sample’s [uniqueness](https://docs.voxel51.com/tutorials/uniqueness.html), you could do so as follows: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4import fiftyone.core.odm as foo 5 6# load dataset 7dataset = foz.load_zoo_dataset("quickstart") 8 9# compute uniqueness 10fob.compute_uniqueness(dataset) 11 12# Add a generic embedded document field to which you can add any fields you want 13dataset.add_sample_field( 14 "custom_metadata", 15 fo.EmbeddedDocumentField, 16 embedded_doc_type=foo.DynamicEmbeddedDocument, 17) 18 19# set values and add to schema 20dataset.set_values( 21 "custom_metadata.uniqueness", 22 dataset.values("uniqueness"), 23 dynamic=True, 24) 25 26# visualize 27session = fo.launch_app(dataset) ``` \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Learn more about [embedded documents](https://docs.voxel51.com/user_guide/using_datasets.html#custom-embedded-documents) and [dynamic attributes](https://docs.voxel51.com/user_guide/using_datasets.html#dynamic-attributes) in the FiftyOne Docs. ## **Changing tags when loading CVAT annotations into FiftyOne** Community Slack member Daniel Fortunato asked, _“Hi all! I want to load samples from a CVAT annotation run that have the tag “to\_annotate”, and then change these to “being\_annotated” in FiftyOne, so that I can keep track of what samples still need to be loaded. I tried using `load_annotation_view()` to load the view from a specific annotation run, but this does not seem to work with changing the tags. How would you recommend I do this?”_ Hey, Daniel! When you change the tags, the reason `load_annotation_view()` no longer works is that internally, the method is using a [MatchTags](https://docs.voxel51.com/api/fiftyone.core.stages.html?highlight=matchtags#fiftyone.core.stages.MatchTags) view stage, which is defined by finding all samples that have certain tags. If you create the view by passing “ _to\_annotate”_ into _`match_tags()`_, and then change the tags on your samples to _“being\_annotated”_, these samples will no longer match the condition. An alternative approach that bypasses this problem is to redefine the `DatasetView` with a `select()` operation after `match_tags()` and before you change the tags. The [Select](https://docs.voxel51.com/api/fiftyone.core.stages.html?highlight=matchtags#fiftyone.core.stages.Select) view stage is defined by a set of sample IDs, so it will not be impacted by changes in tags. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# load dataset 5dataset = foz.load_zoo_dataset("quickstart") 6 7# match tags 8view = dataset.match_tags("to_annotate") 9 10# redefine the view by sample ID 11view = dataset.select(view) 12 13# change tags 14view.untag_samples("to_annotate") 15view.tag_samples("being_annotated") ``` Learn more about [view stages](https://docs.voxel51.com/cheat_sheets/views_cheat_sheet.html), the [FiftyOne Annotation API](https://docs.voxel51.com/user_guide/annotation.html), and our [CVAT integration](https://docs.voxel51.com/integrations/cvat.html) in the FiftyOne Docs. ## **Previewing video frames in FiftyOne** Community Slack member Thrisha Ramkumar asked, _“I have a video dataset with one thousand frames. How do I sample images at regular intervals and preview this in the FiftyOne App?”_ Hi Thrisha! Depending on the length of your videos, you may be able to natively “preview” them in the FiftyOne App with its built-in [video visualizer](https://docs.voxel51.com/user_guide/app.html#using-the-video-visualizer). With the video visualizer, you can play the video by hovering over the sample’s thumbnail, as well as scan frame-by-frame, or jump to specific timestamps. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop If you are working with very large videos, however, it might be the case that you only want to look at one out of every 100 or 1000 frames. One way you could “preview” the video frames in the FiftyOne App is by converting the video dataset to an image dataset, treating each frame as a new sample. You can then use FiftyOne’s [ViewField](https://docs.voxel51.com/api/fiftyone.core.expressions.html?highlight=viewfield#fiftyone.core.expressions.ViewField) and the `match()` method to filter by frame number, in field `frame_number`, of the frames-turned-image samples. For example, if you were working with the [Quickstart Video Dataset](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#quickstart-video), and you wanted to sample frames at a rate of one image per every ten frames in the original videos, you could run the following: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5# load dataset 6dataset = foz.load_zoo_dataset("quickstart-video") 7 8# convert from videos to frames 9frames = dataset.to_frames(sample_frames=True) 10 11# sample every tenth frame 12view = frames.match(F("frame_number") % 10 == 0) 13 14# display, or "preview" the results 15session = fo.launch_app(view) ``` \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Learn more about [working with videos](https://docs.voxel51.com/user_guide/dataset_creation/index.html#loading-videos), [video views](https://docs.voxel51.com/user_guide/using_views.html#video-views), and [frames](https://docs.voxel51.com/user_guide/using_datasets.html#video-frame-labels) in the FiftyOne Docs. ## **Frame-level aggregations** Community Slack member Joy Timmermans asked, _“How do I compute the height distribution for bounding boxes in my video dataset without looping over each detection?”_ Great question, Joy! Depending on what type of statistics or information about the distribution you want to extract from the dataset, there are a variety of [Aggregation](https://docs.voxel51.com/user_guide/basics.html#aggregations) methods available in FiftyOne. These methods allow you to extract values, distinct values, means, or upper and lower bounds. In your case, the `histogram_values()` method might be especially useful, which allows you to compute the histogram over a field’s values. You can even set the number of bins and the range for the histogram! In FiftyOne, aggregations and many other operations work natively on the frames of videos via the `"."` syntax to access frame-level attributes. For instance, the following generates a histogram of frame height values across all frames and samples in the [Quickstart Video Dataset](https://docs.voxel51.com/api/fiftyone.core.dataset.html?highlight=histogram_values#fiftyone.core.dataset.Dataset.histogram_values). ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5# load dataset 6dataset = foz.load_zoo_dataset("quickstart-video") 7 8# compute frame width and height 9dataset.compute_metadata() 10 11# compute histogram counts and bin edges 12# height is last element of bounding box 13count, edges, _ = dataset.histogram_values( 14 F('frames.detections.detections.bounding_box')[3] 15) ``` Learn more about [histogram\_values()](https://docs.voxel51.com/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.histogram_values) and [Expressions](https://docs.voxel51.com/api/fiftyone.core.expressions.html) in the FiftyOne Docs. ## **Deleting duplicate samples** Community Slack member Dan Erez asked, _“I accidentally ended up with a bunch of duplicate samples in my dataset. Is there a quick way to drop the duplicates and keep only one of each?”_ Hey Dan! Accidental duplication is a common problem when dealing with computer vision data. In FiftyOne, for instance, if you try to add a sample that already exists in your dataset, rather than doing nothing or overwriting the original sample, the dataset will create a new copy of the sample and add _that_ to the dataset. For instance, the following code adds fifty duplicate samples to the [Quickstart Dataset](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#quickstart): ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# load dataset 5dataset = foz.load_zoo_dataset("quickstart") 6 7print(dataset.count()) 8# 200 9 10# randomly select 50 samples 11samples_to_duplicate = dataset.take(50) 12 13# add these as duplicates 14dataset.add_samples(samples_to_duplicate) 15 16print(dataset.count()) 17# 250 ``` In many cases, the ability to have multiple samples with the same filepath can be useful, but in some cases, this behavior can have unintended consequences. Assuming that every sample in your original dataset had a unique filepath, it is easy to “undo” this duplication and get back to your initial dataset. The key to identifying the duplicates is finding which filepaths occur more than once in the dataset, and then deleting all but one sample with each. To find the multiply-occuring filepaths, you can use FiftyOne’s `count_values()` aggregation: ```python 1fp_counts = dataset.count_values("filepath") 2dup_fps = [key for key in list(fp_counts.keys()) if fp_counts[key] > 1] ``` In the above example, where samples were duplicated at most once, we can deduplicate our dataset by getting the sample IDs of all duplicate filepaths, and passing these into `delete_samples()`: ```python 1## IDs of first sample in dataset for each fp 2sids = [dataset[dup_fp].id for dup_fp in dup_fps] 3dataset.delete_samples(sids) ``` If you have more than one duplicate per filepath, check out our [recipe for image deduplication](https://docs.voxel51.com/recipes/image_deduplication.html). Learn more about [aggregations](https://docs.voxel51.com/user_guide/basics.html#aggregations) and [deduplication](https://docs.voxel51.com/recipes/remove_duplicate_annos.html) in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,400+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,600+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,300+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 254+ repositories - 56+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [aggregations](https://voxel51.com/blog/tag/aggregations) [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [quickstart dataset](https://voxel51.com/blog/tag/quickstart-dataset) [quickstart video dataset](https://voxel51.com/blog/tag/quickstart-video-dataset) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/dc8a2e7a894316856af5a109ae43f8787959f179-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Embeddings Tips and Tricks – Mar 31, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 31, 2023](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/ecb6afb20436d0f0e68fbb25bcfc7443657d7b91-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Feb 10, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Feb 10, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/a3a918e30b0553723b9392ea90763379f98480a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – April 7, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Apr 7, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-266-lllmstxt|> ## March 2023 Computer Vision Meetup [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup — March 2023 Mar 14, 2023 • 16 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#1f175d39454d) [Computer Vision Meetup Recap at a Glance](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#34f7e86fe3a2) [Why Discard if You can Recycle?: A Recycling Max Pooling Module for 3D Point Cloud Analysis](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#11ccc3d0327f) [Lighting Up Images in the Deep Learning Era](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#2c2e8d269b19) [Taking Computer Vision Models in Notebooks to Production](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#ac8da253877a) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#9d671ad89fa3) [Upcoming Computer Vision Meetup Speakers & Schedule](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#ff350d91cbaf) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#4486833edfe6) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#1f175d39454d) [Computer Vision Meetup Recap at a Glance](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#34f7e86fe3a2) [Why Discard if You can Recycle?: A Recycling Max Pooling Module for 3D Point Cloud Analysis](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#11ccc3d0327f) [Lighting Up Images in the Deep Learning Era](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#2c2e8d269b19) [Taking Computer Vision Models in Notebooks to Production](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#ac8da253877a) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#9d671ad89fa3) [Upcoming Computer Vision Meetup Speakers & Schedule](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#ff350d91cbaf) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#4486833edfe6) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Last Thursday, Voxel51 hosted the March 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/). In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup and events schedule so that you can join us in the future. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes this month was [The Foundation Fighting Blindness](https://www.fightingblindness.org/). We are sending this month’s charitable donation of $200 to them on behalf of the computer vision community! ![](https://cdn.sanity.io/images/h6toihm1/production/c1b81a240bde49092b376fffc3fdfd0717ca336b-600x278.png?auto=format&dpr=2&fit=max&q=75&w=600) ## Computer Vision Meetup Recap at a Glance **Jiajing Chen // Why Discard if You can Recycle?: A Recycling Max Pooling Module for 3D Point Cloud Analysis** - [Video replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#talk1-video) - [Presentation summary](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#talk1-summary) - [Additional resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#talk1-resources) **Soumik Rakshit // Lighting Up Images in the Deep Learning Era** - [Video replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#talk2-video) - [Presentation summary](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#talk2-summary) - [Q&A recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#talk2-qanda) - [Additional resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#talk2-resources) **Sumanth P // Taking Computer Vision Models in Notebooks to Production** - [Video replay](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#talk3-video) - [Presentation summary](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#talk3-summary) - [Q&A recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#talk3-qanda) - [Additional resources](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#talk3-resources) **Next steps** - [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#locations) - [Upcoming Computer Vision Meetup Speakers — April and May](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#schedule) - [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023#get-involved) ## Why Discard if You can Recycle?: A Recycling Max Pooling Module for 3D Point Cloud Analysis ### **Video replay** https://www.youtube.com/watch?v=0kxvRLQgZ0Q ### **Presentation Summary** [Jiajing Chen](https://www.linkedin.com/in/jiajing-chen-560189193/), a PhD student from Syracuse University, presents on this topic and research: Why Discard If You Can Recycle: A Recycling Max Pooling Module for 3D Point Cloud Analysis. Jiajing opens up by explaining what 3D point cloud data is and where it comes from, along with ways to apply machine learning techniques to point clouds: classification, object detection, and segmentation. ![](https://cdn.sanity.io/images/h6toihm1/production/6b73d26631016447cb92f1f451f8bdf90477eb48-1999x1129.png?auto=format&dpr=2&fit=max&q=75&w=1600) Next, Jiajing shows a general network architecture of most point-based models and then explains how the limitations of traditional approaches motivated his work. ![](https://cdn.sanity.io/images/h6toihm1/production/6a8361af12e8105655dd878faea1169551b180d6-1308x730.png?auto=format&dpr=2&fit=max&q=75&w=1308) - Most 3D point cloud analysis models have focused on developing different point feature learning and neighbor feature aggregation modules - One common theme with most of the existing approaches is their use of max-pooling to obtain permutation invariant features - Yet traditional max-pooling causes only a fraction of 3D points (red boxes) to contribute to the permutation invariant features and discards the rest (all purple rows) Jiajing, along with colleagues, set out on a research journey to see if they could find a way to recycle those discarded points to make the original network's performance better. And indeed, they did! The research ( [available here](https://openaccess.thecvf.com/content/CVPR2022/papers/Chen_Why_Discard_if_You_Can_Recycle_A_Recycling_Max_Pooling_CVPR_2022_paper.pdf)) has shown that recycling the still useful discarded points can improve the original network’s performance. In the rest of the talk, Jiajing explains the series of experiments that were conducted, the key observations they made along the way, and the new proposed method for a recycling max pooling module. In the first experiment, Jiajing and his colleagues set out to verify whether or not the number of utilized points has a correlation with the final accuracy. They selected three milestone networks for the experiments, PointNet, PointNet++, and DGCNN, and studied point utilization during prediction, both before training and after training. What they found was that at the end of the training the number of points kept after the max-pooling increased. Therefore the first key observation was: prediction accuracy is positively correlated with the point utilization percentage, indicating that recycling the discarded points has promise. (You can watch this part of the presentation from [~4:11](https://www.youtube.com/watch?v=0kxvRLQgZ0Q&t=251s) \- 6:05 in the video playback.) ![](https://cdn.sanity.io/images/h6toihm1/production/6699ef8242229ab721c97c3ea78c0cec2fb6affd-1999x1128.png?auto=format&dpr=2&fit=max&q=75&w=1600) Moving onto the second observation, analysis of the potential of discarded points, Jiajing presents the information to us NFL-play-by-play style (which you can watch from [~6:05](https://www.youtube.com/watch?v=0kxvRLQgZ0Q&t=365s) \- 8:49 in the video playback). The research team ultimately observes that applying recycling to max pooling (shown in the slide below as F2 and F3) does result in a drop in accuracy, but it’s not really that much, which means these discarded points indeed do have useful features that should be recycled. ![](https://cdn.sanity.io/images/h6toihm1/production/de2a0c7fd300ae4be5be34e2458d28f83287da57-1999x1123.png?auto=format&dpr=2&fit=max&q=75&w=1600) Next, Jiajing gives us an in-depth tour of the proposed method: a new module, referred to as the Recycling Max-Pooling (RMP) module, to recycle and utilize the features of some of the discarded points (you can watch the replay from [~8:59](https://www.youtube.com/watch?v=0kxvRLQgZ0Q&t=539s) to 12:48 in the video). Finally, Jiajing shares the experiment results to show us how the RMP module performed (from [~12:48](https://www.youtube.com/watch?v=0kxvRLQgZ0Q&t=768s) \- 14:15 in the video playback). The first round of experiments were performed on ScanObjectNN and ModelNet40 datasets for the point cloud classification task across a variety of milestone networks and state-of-the-art networks (PointNet, PointNet++, DGCNN, GDANet, DPFA, and CurveNet). **Applying the RMP module resulted in performance improvements for all networks.** Experiments were also performed for the point cloud segmentation task using the S3DIS dataset. This dataset is an indoor dataset that contains six areas covering about 271 rooms and each point belongs to one of 13 classes. The experiments were performed on the PointNet, DGCNN, and DPFA networks by applying the RMP module to them (and comparing the results to the experiments without the RMP module applied). **The RMP module brought improvements in overall accuracy and the mean IoU!** ### **Additional Resources** Check out the additional resources on the presentation: - [Talk transcript](https://www.rev.com/transcript-editor/shared/MAk-Sz71V3w6kqgQAdKmeTo6sO7UE4-WdAAOLQqngjfkMZMe4qui-bVLQD9wlqoLqiqEXpx9O-YwXLx3QAAiDjKorBo?loadFrom=SharedLink) - [Presentation slides](https://voxel51.com/wp-content/uploads/2023/03/RMP-Recycling-Max-Pooling-Presentation.pdf) Thank you Jiajing for sharing your research and information about the recycling max pooling module for 3D point clouds with us! ## Lighting Up Images in the Deep Learning Era ### **Video Replay** https://www.youtube.com/watch?v=FV1VZFXPk8s ### **Presentation Summary** [Soumik Rakshit](https://www.linkedin.com/in/soumikrakshit/), Machine Learning Engineer at Weights & Biases, gives a talk on image restoration, with a focus on how low light enhancement is addressed in the deep learning era. Why do we need to lighten up images at all? Soumik answers: images are often taken under sub-optimal lighting conditions, in uneven light, dim light, and in scenarios where the light is shining from behind the subject (backlit subjects). He further explains: “it turns out that such poorly lit images suffer a lot from not just compromised aesthetic quality, but also from diminished performance on high level computer tasks like object detection, object recognition, image segmentation, and other operations.” Soumik provides examples of how low light enhancements can be used in computer vision. First he shows how applying YOLOv8-large on a lightened image resulted in more detected objects than applying YOLOv8-large on the original low light image. Other applications for low light image enhancements are: visual surveillance, autonomous driving, and computational photography – in particular, smartphone photography, where it’s possible to turn your device into a night-vision system. Traditional methods for low light enhancement can be broadly categorized as Histogram Equalization and Retinex Models. But these traditional methods have some limitations: they are noisy, have high computational complexity, and require manual tweaking. Deep learning offers promise for improvements over traditional methods. Since the publication of [LLNet](https://arxiv.org/abs/1511.03995) in 2017, recent years have witnessed the compelling success of deep learning-based approaches for low light image enhancement. ![](https://cdn.sanity.io/images/h6toihm1/production/a5527c7b635e6d5b6c71d7b162cd3f400255c6f9-1999x939.png?auto=format&dpr=2&fit=max&q=75&w=1600) In the rest of talk, Soumik explores three of the models listed in the image above (MIRNet-v2, NAFNet, and Zero-DCE), including their architectures, as well as how to train them, evaluate them, and see how they perform on a few real-world images. But first, Soumik notes that the process starts with a dataset: the LoL dataset or LOw Light paired dataset \[ [Original Source](https://daooshee.github.io/BMVC2018website/) \| [Kaggle Dataset](https://www.kaggle.com/datasets/soumikrakshit/lol-dataset)\] which was created for training supervised models for low-light image enhancements. This dataset provides 485 images for training and 15 for testing. Each image pair in the dataset consists of a low light input image and its corresponding well-exposed reference image. Soumik builds out an input pipeline on the LoL dataset using [restorers](https://github.com/soumik12345/restorers) and [wandb](https://github.com/wandb/wandb). Restorers is an open source tool written using TensorFlow and Keras that provides out-of-the-box TensorFlow implementations of state-of-the-art (SOTA) image and video restoration models for tasks such as low light enhancement, denoising, deblurring, super-resolution, and more. Weights & Biases ( [wandb](https://github.com/wandb/wandb)) is an open source tool for visualizing and tracking your machine learning experiments. Now, Soumik dives into three models. “MIRNet-v2, as proposed by the paper [Learning Enriched Features for Fast Image Restoration and Enhancement](https://www.waqaszamir.com/publication/zamir-2022-mirnetv2/zamir-2022-mirnetv2.pdf) is a fully convolutional architecture that learns enriched feature representations for image restoration and enhancement. It is based on a recursive residual design with the multi-scale residual block or MRB at its core.”\[1\] ![](https://cdn.sanity.io/images/h6toihm1/production/cf8d88db7201aa07765dcae24f9dcaf8d653c111-2048x984.png?auto=format&dpr=2&fit=max&q=75&w=1600) To watch Soumik train the MIRNet-v2 model for low-light enhancement model using restorers and wandb, visit this portion of the video playback: [~14:29](https://www.youtube.com/watch?v=FV1VZFXPk8s&t=869s) \- 24:07. For the code, you can refer to this [Colab notebook](https://colab.research.google.com/github/wandb/examples/blob/restorers/colabs/keras/restorers/Train_MirNetv2_Restorers.ipynb). Next, Soumik explains what the NAFNet model is. “NAFNet or the Nonlinear Activation Free Network is a simple baseline model for all kinds of image restoration tasks as proposed by the paper [Simple Baselines for Image Restoration](https://arxiv.org/abs/2204.04676). The aim of the authors was to create a simple baseline that exceeds the then SOTA methods in terms of performance and is also computationally efficient. To further simplify the baseline, the authors reveal that nonlinear activation functions, e.g. Sigmoid, ReLU, GELU, Softmax, etc. are not necessary: they could be replaced by multiplication or removed.”\[1\] ![](https://cdn.sanity.io/images/h6toihm1/production/d23f363991a39b733d9a48c52065bb1ecbc93ec2-1178x578.png?auto=format&dpr=2&fit=max&q=75&w=1178) Learn more about NAFNet, including how to train it using restorers and wandb from [~24:29](https://www.youtube.com/watch?v=FV1VZFXPk8s&t=1469s) \- 31:33 in the video replay. For the code, you can refer to this [Colab notebook](https://colab.research.google.com/github/wandb/examples/blob/restorers/colabs/keras/restorers/Train_Nafnet_Restorers.ipynb). Soumik explains the last of the three models that he will feature in this presentation. “ [Zero-Reference Deep Curve Estimation](https://arxiv.org/abs/2001.06826) or Zero-DCE formulates low-light image enhancement as the task of estimating an image-specific tonal curve) with a deep neural network. Instead of performing image-to-image mapping, in the case of Zero-DCE, the problem is reformulated as an image-specific curve estimation problem. In particular, the proposed method takes a low-light image as input and produces high-order curves as its output. These curves are then used for pixel-wise adjustment on the dynamic range of the input to obtain an enhanced image. A unique advantage of this approach is that it is zero-reference, i.e., it does not require any paired or even unpaired data in the training process as in existing CNN-based (such as [MIRNet](https://arxiv.org/abs/2003.06792)) and GAN-based methods (such as [EnlightenGAN](https://arxiv.org/abs/1906.06972)).”\[1\] ![](https://cdn.sanity.io/images/h6toihm1/production/7b7a02576298668853116003feca11332cbcbbe7-891x375.png?auto=format&dpr=2&fit=max&q=75&w=891) Learn more about Zero-DCE, including how to train it using restorers and wandb from [~31:33](https://www.youtube.com/watch?v=FV1VZFXPk8s&t=1893s) \- 37:02 in the video replay. For the code, check out this [Colab notebook](https://colab.research.google.com/github/wandb/examples/blob/restorers/colabs/keras/restorers/Train_Zero_DCE_Restorers.ipynb). Now that Soumik has shown us how to train low light enhancement models, he shows us how to evaluate them on the LoL dataset using the [restorers.evaluation](https://github.com/soumik12345/restorers/tree/d8a39973ebb374c466e00277610cc962f535ca54/restorers/evaluation) API, and how to log the results in a Weights & Biases dashboard. For the code, check out this [Colab notebook](https://colab.research.google.com/github/wandb/examples/blob/restorers/colabs/keras/restorers/Evaluation_low_light.ipynb). While upon first glance, it seems like NAFNet is the most performant model according to the number of trainable parameters, and especially compared to MIRNet-v2, Soumik next runs inference on the models and shows examples of poorly lit images so we can see how the models perform. (You can follow along with this part from [~40:09](https://www.youtube.com/watch?v=FV1VZFXPk8s&t=2409s) \- 44:35 in the video playback.) What do we see? NAFNet produces some great images, but it can also introduce weird artifacts in your image. This won’t work for a scenario where a low light enhancement model is used as a pre-processing step for object detection models. Soumik has already dismissed MIRNet-v2 because it was too slow and too large of a model. But Zero-DCE is looking really good. ![](https://cdn.sanity.io/images/h6toihm1/production/57965f168f705c8a20fd6ebe605b7c759c5c49c8-1844x1340.png?auto=format&dpr=2&fit=max&q=75&w=1600) Zero-DCE is a lightweight model (less than 1MB) which makes it the best candidate to be applied for a low light enhancement model as a pre-processing step for computer vision tasks. Zero-DCE is quick and easy to train and it produces images that are visually pleasing. Soumik ends the presentation with this conclusion: all of these factors make Zero-DCE one of the most desirable models to be taken into production for low light image enhancement. ### **Q&A Recap** **The unsupervised model is cool…but what would be the benefit over just applying the contrast and pixel distribution ops to the images “manually”?** In the presentation, we briefly touched on the pitfalls of traditional auto contrast algorithms, which are [summarized here](https://wandb.ai/ml-colabs/low-light-enhancement/reports/Lighting-up-Images-in-the-Deep-Learning-Era--VmlldzozNzE4Njkz#%E2%98%8E%EF%B8%8F-traditional-methods-for-low-light-enhancement). In addition to that, they do not perform as well as Zero-DCE. When I created this tutorial for Zero-DCE over on keras.io, François Chollet, the creator of Keras, on a PR thread performed an interesting benchmark. You can see the original low light image compared to the result by a traditional auto contrast algorithm, and then the result that was spewed out by Zero-DCE. You can clearly see that the auto contrast models do not actually perform very well. ![](https://cdn.sanity.io/images/h6toihm1/production/435610af3ded9480c363a20165aea18dab0f9514-1450x1284.png?auto=format&dpr=2&fit=max&q=75&w=1450) **Are the test models from a totally new scene? Or related to trained data?** The models are not being tested on trained data. If we look at the results for evaluation, you can see the distinction. The top panels show the results on the training dataset. And the panels in the second row show the results on completely new data, which the model hasn't seen at all. ![](https://cdn.sanity.io/images/h6toihm1/production/4164a960ae3ff0a5b79fef2ece4bd6b0b570ef92-1496x824.png?auto=format&dpr=2&fit=max&q=75&w=1496) In terms of qualitative analysis, Zero-DCE has already shown that it's capable of holding its own against the larger models like NAFNet or MIRNet-v2. Also the images that we show in the “usage section” of the [online tutorial](https://wandb.ai/ml-colabs/low-light-enhancement/reports/Lighting-up-Images-in-the-Deep-Learning-Era--VmlldzozNzE4Njkz) are completely out of distribution images from different datasets (some are from the Dark Face dataset, which is actually for object detection in incredibly low light conditions; and I tossed in some other ones as well). ### **Additional Resources** Check out these additional resources: - [Talk transcript](https://www.rev.com/transcript-editor/shared/bmZn4cPEoosvNtLt64zFsDyhcLNsouth3Q5KhHPlUMx9vG41EmOHzKKzehCAEqyuf3P0xzqDxVp4ae7uvT8-kQFMpkg?loadFrom=SharedLink) - [Presentation slides](https://docs.google.com/presentation/d/1nmNwymrP0ExoF7W08Vo5Os2E91i7IXuIajqNMOpwYS0/edit) - \[1\] [Online tutorial](https://wandb.ai/ml-colabs/low-light-enhancement/reports/Lighting-up-Images-in-the-Deep-Learning-Era--VmlldzozNzE4Njkz) - Weights & Biases is hosting [Fully Connected](https://www.fullyconnected.com/), The ML Conference for practitioners, by practitioners. [Register](https://www.fullyconnected.com/) to hear from the teams building the most impactful large models & production-ready ML models with speakers from Stability AI, Spotify, NVIDIA, [fast.ai](http://fast.ai/) and many more. - Would you like to learn more about Weights & Biases? Request a demo [here](https://wandb.ai/site/contact). A big thank you to Soumik on behalf of the entire Computer Vision Meetup community for enlightening us on how to light up low light images in the deep learning era! ## Taking Computer Vision Models in Notebooks to Production ### **Video Replay** https://www.youtube.com/watch?v=HLTTrqHNnz0 ### **Presentation Summary** Sumanth’s presentation is all about taking models to production. As a machine learning engineer or data scientist, it’s not the case that your job is done when you create a model and deploy it. The machine learning lifecycle is continuous. You need to continuously monitor your model and iterate on it, whether it is retraining the model, collecting more data, or other important tasks with the goal of creating an end-to-end model. To be successful creating a process that works end-to-end all the way through production, you need to keep track of your data, your steps, and your environment. Sumanth explores some common problems after deployment, and ways to address them, including, with the biggest issue being data drift. Data drift is a common problem faced in machine learning that causes the model quality to decrease over time. This is because there are gaps in your data. Maybe there are cases your data didn’t cover when you were building the model, but you are now seeing when your model is in production. Ex: maybe the data you initially collected was from “summer” but now your model is operating in “winter”. Different types of data drift include covariate shift, label shift, and domain shift. To avoid pitfalls when taking models into production, there are some best practices: data versioning, data & model validation, experiment tracking, scalable pipelines, and continuous monitoring. The good news is that there are many open source tools that will help you address the common problems across the machine learning lifecycle. Sumanth explores some examples: ![](https://cdn.sanity.io/images/h6toihm1/production/27c1159308f88788e079a6017badef870958de56-1896x916.png?auto=format&dpr=2&fit=max&q=75&w=1600) But after you have all these disparate open source tools running, how do you manage them? And what if you switch a tool? You don’t want to have to rewrite your code. So how can we create reproducible, maintainable, and modular data science code? Enter ZenML. Sumanth loves this tool and demonstrates it in the presentation (from [~25:26](https://www.youtube.com/watch?v=HLTTrqHNnz0&t=1526s) \- 35:01). (He also notes that alternatives include Kedro and Flyte in case you want to check them out). ZenML has a notion of stacks, which represent a set of configurations for your ML Ops tools and infrastructure. For examples, you can do all this with ZenML stacks: - Orchestrate your ML workflows with Kubeflow - Save ML artifacts in an Amazon S3 bucket - Track your experiments with Weights & Biases - Deploy models on Kubernetes with BentoML Sumanth explains that ZenML makes all this easy. Plus you can change flavors, which means easily change and switch between the tools you want to use, easily in ZenML too. In summary: there are common challenges in the ML lifecycle that give rise to best practices, and there are a bunch of open source tools to help you along the way. Simply select the right one for your use case, and if you need to switch tools, there are tools to help you do that easily, too. ### **Q&A Recap** **Are there tools to automate the measurement of label shift?** Yes, you’re able to track this by tracking the model after you deploy it. The tool I use for this is [Deepchecks](https://deepchecks.com/) for continuous data and model validation. **How important is the evaluation of "Neural network verification" in model evaluation?** It’s important. Once you create a model, it’s general practice to evaluate the model. You have to find the model that performs the best for your use case, so it should always be done. **One of the toughest “drift” problems I’ve had in the past in production I had to investigate - it turned out that some training data started to leak into the test set (both were evolving). What methods or tools would you recommend for preventing this issue?** I’ve seen problems with data leakage as well, mostly due to trying to divide a dataset for train, test, and validation. One suggestion here is to follow the best practices for splitting data into the different sets; these little things can also have an impact on reducing or stopping data leaks. **What’s the best way to deploy a CV model?** There are really cool tools for this based on your specific use case, and some examples are: BentoML, KServe, and others like MLflow and Seldon Core. Each of these has its own advantages so I encourage you to check them out, as well as the open source communities around them. ### **Additional Resources** Check out the [talk transcript](https://www.rev.com/transcript-editor/shared/VaDI8jlu7d3yT4Ks6G_OR2Vlph8vEK1IZqWZEWfqLkD1Hy8ejnE448eh85cuslcfeTp6GXD49YrxIN0RHnFijyxvaaI?loadFrom=SharedLink). Also, thank you to Sumanth for all the great information on taking computer vision models to production! ## Computer Vision Meetup Locations Computer Vision Meetup membership has grown to nearly [3,300+ members](https://www.meetup.com/pro/computer-vision-meetups/) in just a few months! The goal of the meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. We invite you to join one of the 13 Computer Vision Meetup locations closest to your timezone: - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) ## Upcoming Computer Vision Meetup Speakers & Schedule We have other exciting speakers already signed up for the next several Computer Vision Meetups! Become a member of the [Computer Vision Meetup closest to you](https://www.meetup.com/pro/computer-vision-meetups/) to get details about the meetups already scheduled and be one of the first to receive new meetup details as they become available. ### April 13 @ 10AM PT - Generating Diverse and Natural 3D Human Motions from Texts - Chuan Guo (University of Alberta) - Emergence of Maps in the Memories of Blind Navigation Agents - Dhruv Batra (Meta & Georgia Tech) - Using Computer Vision to Understand Biological Vision - Benjamin Lahner (MIT) - Here’s the [Zoom Link](https://us02web.zoom.us/webinar/register/3416762523346/WN_oXTCgqQiQT6ouNxygQkpsg) to register ### April 26 @ 9:30 PM PT (APAC) - Leveraging Attention for Improved Accuracy and Robustness - Hila Chefer (Tel Aviv University) - AI Deployments at the Edge with OpenVINO - Zhuo Wu (Intel) - Here’s the [Zoom Link](https://us02web.zoom.us/webinar/register/1116762530249/WN_pXqu-h8cQHeNKExjQtiLAA) to register ### May 11 @ 10AM PT - The Role of Symmetry in Human and Computer Vision - [Sven Dickinson](https://www.linkedin.com/in/sven-dickinson-1091b73/) (University of Toronto & Samsung) - Machine Learning for Fast, Motion-Robust MRI - [Nalini Singh](https://www.linkedin.com/in/nalinimsingh/) (MIT) - Here’s the [Zoom Link](https://us02web.zoom.us/webinar/register/2516761651886/WN_D0MKbn3eTcSRtyHW1A6z1Q) to register ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [image restoration](https://voxel51.com/blog/tag/image-restoration) [low light image enhancement](https://voxel51.com/blog/tag/low-light-image-enhancement) [models in production](https://voxel51.com/blog/tag/models-in-production) [point clouds](https://voxel51.com/blog/tag/point-clouds) [recycling max pooling](https://voxel51.com/blog/tag/recycling-max-pooling) [RMP](https://voxel51.com/blog/tag/rmp) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/308698a5aece1d5b1b95ee1bf52811b24448458c-1200x672.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne 0.18 for Computer Vision\\ \\ Event Recaps\\ \\ • \\ \\ Dec 6, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ed751b5c0fbc3d2cb74f0b30e6418d3319564c-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — December 2022\\ \\ Event Recaps\\ \\ • \\ \\ Dec 13, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/bbb1d9add0b0b9aa12682acac795df7c2ba760a9-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — November 2022\\ \\ Event Recaps\\ \\ • \\ \\ Nov 16, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-267-lllmstxt|> ## Cityscapes Dataset Exploration [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) Exploring the Cityscapes Dataset for Semantic Urban Scene Understanding Mar 14, 2023 • 6 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#332faf9fdc57) [About the Cityscapes Dataset](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#497e944be4d6) [What Is Visual Scene Understanding?](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#59d57327d873) [Design Choices](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#0a6ac050eadc) [Features](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#86fec6dd8976) [Labeling Policy](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#8672b3f6c01f) [Class Definitions](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#6f377f59b7be) [Dataset Quick Facts](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#fe8770e721f9) [Step 1: Download the Dataset](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#30588b37fdee) [Step 2: Install FiftyOne](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#1f6309963a43) [Step 3: Import the Dataset](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#92bd05d01783) [Sample Details](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#c4c7591254e5) [Filtering by ID](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#9b0aef96be35) [Filtering by Label](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#45b6c4425ee5) [Start Working with the Cityscapes Dataset](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#9cc5d096e279) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#332faf9fdc57) [About the Cityscapes Dataset](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#497e944be4d6) [What Is Visual Scene Understanding?](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#59d57327d873) [Design Choices](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#0a6ac050eadc) [Features](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#86fec6dd8976) [Labeling Policy](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#8672b3f6c01f) [Class Definitions](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#6f377f59b7be) [Dataset Quick Facts](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#fe8770e721f9) [Step 1: Download the Dataset](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#30588b37fdee) [Step 2: Install FiftyOne](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#1f6309963a43) [Step 3: Import the Dataset](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#92bd05d01783) [Sample Details](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#c4c7591254e5) [Filtering by ID](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#9b0aef96be35) [Filtering by Label](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#45b6c4425ee5) [Start Working with the Cityscapes Dataset](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding#9cc5d096e279) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Welcome to the latest installment of our ongoing blog series where we highlight datasets from the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/datasets.html)! FiftyOne provides a Dataset Zoo that contains a collection of common datasets that you can download and load into FiftyOne via a few simple commands. In this post, we explore the [Cityscapes](https://www.cityscapes-dataset.com/) dataset. ## Wait, What’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop The FiftyOne Dataset Zoo comprises more than 30 datasets, with new datasets being added all the time! They cover a variety of computer vision use cases including: - Video - Images - Location - Point-cloud - Action-recognition - Classification - Detection - Segmentation - Relationships - And more! ## About the Cityscapes Dataset The Cityscapes Dataset is a large-scale dataset that contains a diverse set of stereo video sequences recorded in street scenes from 50 different cities, with high quality pixel-level annotations of 5,000 frames in addition to a larger set of 20 000 weakly annotated frames. At the time of its release it was an order of magnitude larger than similar previous attempts. Its primary use case is for assessing the performance of vision algorithms for major tasks of semantic urban scene understanding: pixel-level, instance-level, and panoptic semantic labeling; supporting research that aims to exploit large volumes of (weakly) annotated data, e.g. for training deep neural networks. ## What Is Visual Scene Understanding? Scene understanding is the process of perceiving, analyzing and elaborating an interpretation of a 3D dynamic scene observed through a network of sensors. This usually involves matching signal information coming from the sensors observing the scene, with machine learning models humans are using to understand the scene. As a result, scene understanding both adds and extracts semantic information from the sensor data characterizing a scene. The type of sensors usually involved in visual scene understanding are cameras. But, you may also have scenarios where additional data is being captured by microphones, radar or other sensors. Object-wise, the scene can contain a variety of physical objects of various types (for example cars and people) interacting with each other or with their environment. The scene itself can be just a few seconds long or a multi-day time lapse. It can also be confined to a microscopic view or involve an entire cityscape. ## Design Choices Here’s an overview of the design choices that were made in regards to the dataset’s focus. ## Features Polygonal annotations - Dense semantic segmentation - Instance segmentation for vehicle and people Complexity - 30 classes - See Class Definitions for a list of all classes and have a look at the applied labeling policy. Diversity - 50 cities - Several months (Spring, Summer, Fall) - Daytime - Good/medium weather conditions - Manually selected frames - Large number of dynamic objects - Varying scene layout - Varying background Volume - 5 000 annotated images with fine annotations (examples) - 20 000 annotated images with coarse annotations (examples) Metadata - Preceding and trailing video frames. Each annotated image is the 20th image from a 30 frame video snippets (1.8s) - Corresponding right stereo views - GPS coordinates - Ego-motion data from vehicle odometry - Outside temperature from vehicle sensor Extensions by other researchers - Bounding box annotations of people - Images augmented with fog and rain Benchmark suite and evaluation server - Pixel-level semantic labeling - Instance-level semantic labeling - Panoptic semantic labeling ## Labeling Policy Labeled foreground objects must never have holes. For example if there is some background visible ‘through’ some foreground object, it is considered to be part of the foreground. This also applies to regions that are highly mixed with two or more classes: they are labeled with the foreground class. Some examples would include: - tree leaves in front of house or sky (everything tree) - transparent car windows (everything car) ## Class Definitions **Group** **Classes** flat - road - sidewalk - parking+ - rail track+ human - person\* - rider\* vehicle - car\* - truck\* - bus\* - on rails\* - motorcycle\* - bicycle\* - caravan\* - trailer\*+ construction - building - wall - fence - guard rail+ - bridge+ - tunnel+ object - pole - pole group+ - traffic sign - traffic light nature - vegetation - terrain sky - sky void - ground+ - dynamic+ - static+ \\* Single instance annotations are available. However, if the boundary between such instances cannot be clearly seen, the whole crowd/group is labeled together and annotated as group, e.g. car group. \+ This label is not included in any evaluation and treated as void (or in the case of license plate as the vehicle mounted on). ## Dataset Quick Facts - **Research Paper:** [The Cityscapes Dataset for Semantic Urban Scene Understanding](https://www.cityscapes-dataset.com/wordpress/wp-content/papercite-data/pdf/cordts2016cityscapes.pdf) - **Authors:** M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele - **Download Dataset:** [Register and download](https://www.cityscapes-dataset.com/downloads/) - **License:** [Free](https://www.cityscapes-dataset.com/license/), but registration is required - **Dataset Size:** 11.8 GB - **Last Release:** 2016 - **FiftyOne Dataset Name:** `cityscapes` - **Tags:** `image`, `multilabel`, `automotive`, `manual` - **Supported Splits:** `train`, `validation`, `test` - **Zoo Dataset class:** [`CityscapesDataset`](https://docs.voxel51.com/api/fiftyone.zoo.datasets.base.html#fiftyone.zoo.datasets.base.CityscapesDataset) ## Step 1: Download the Dataset In order to load the Cityscape Dataset into FiftyOne, you must [download the source data manually](https://www.cityscapes-dataset.com/register/) with your `source_dir` organized in the following manner: Note that the `gtFine_trainvaltest`, `gtCoarse`, and `gtBbox_cityPersons_trainval` are optional directories. ## Step 2: Install FiftyOne If you don’t already have FiftyOne installed on your laptop, it takes just a few minutes! For example on macOS: - [Verify your version](https://docs.voxel51.com/getting_started/virtualenv.html#creating-a-virtual-environment-using-venv) of Python - Create and activate a [virtual environment](https://docs.voxel51.com/getting_started/virtualenv.html#creating-a-virtual-environment-using-venv) - [Install IPython](https://docs.voxel51.com/getting_started/troubleshooting.html#ipython-installation) (optional) - [Upgrade](https://docs.voxel51.com/getting_started/virtualenv.html#creating-a-virtual-environment-using-venv) your `Setuptools` - [Install FiftyOne](https://docs.voxel51.com/getting_started/install.html#installing-fiftyone) Learn more about how to [get up and running with FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html) in the Docs. ## Step 3: Import the Dataset Now that you have the dataset downloaded and FiftyOne installed, let’s import the dataset into FiftyOne and launch the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html). This should take just a few minutes and a few more lines of code. The last line in the code snippet will launch the FiftyOne App in your default browser. You should see the following initial view of the `cityscapes-validation` dataset in the FiftyOne App: **Tip:** If you want to persist the dataset, add the following to your initial load command: Ok, let’s do a quick exploration of the Cityscape Dataset! ## Sample Details Click on any of the samples to get additional detail like tags, metadata, labels, and primitives. ## Filtering by ID FiftyOne makes it very easy to filter the samples to find the ones that meet your specific criteria. For example we can filter by a specific `id`: ## Filtering by Label In this example we filter the samples by the `gt_person` label selecting only those with `pedestrian`: In this example we filter the samples by the `gt_coarse` label: ## Start Working with the Cityscapes Dataset Now that you have a general idea of what the dataset contains, you can start using FiftyOne to perform a variety tasks including: - [Creating dataset views](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) - [Creating aggregations](https://voxel51.com/docs/fiftyone/user_guide/using_aggregations.html) - [Creating interactive plots](https://voxel51.com/docs/fiftyone/user_guide/plots.html) - [Annotating datasets](https://voxel51.com/docs/fiftyone/user_guide/annotation.html) - [Evaluating models](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html) You can also start making use of the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html) which provides powerful machine learning techniques you can apply to your workflows like visualizing embeddings, finding similarity, uniqueness and mistakenness. [Cityscapes](https://voxel51.com/blog/tag/cityscapes) [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [semantic urban scene understanding](https://voxel51.com/blog/tag/semantic-urban-scene-understanding) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/40381f5f37fa5fcd70eddca2f63b6710568f5d2c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual Kinship Recognition with the Families in the Wild Computer Vision Dataset\\ \\ Datasets\\ \\ • \\ \\ Dec 7, 2022](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/33d08c7b16ab5bfa0e4c5a4f936be4592b8e0a90-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the Berkeley Deep Drive Autonomous Vehicle Dataset\\ \\ Datasets\\ \\ • \\ \\ Jan 11, 2023](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/129c3574861e6c307549e14106753af13ecfa2bf-1308x1044.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How to Download ActivityNet and Evaluate Video Understanding Models\\ \\ Datasets\\ \\ • \\ \\ Feb 8, 2022](https://voxel51.com/blog/how-to-download-activitynet-and-evaluate-video-understanding-models) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-268-lllmstxt|> ## FiftyOne Workshop Series [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Announcing the FiftyOne Computer Vision Educational Workshop Series Mar 17, 2023 • 3 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series#e167b5579ad2) [About the workshop](https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series#1a49f954c060) [Prerequisites for the workshop](https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series#038e603dee3a) [About the FiftyOne computer vision workshop series](https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series#7737a723d505) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series#e167b5579ad2) [About the workshop](https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series#1a49f954c060) [Prerequisites for the workshop](https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series#038e603dee3a) [About the FiftyOne computer vision workshop series](https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series#7737a723d505) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Voxel51 is excited to announce the first in a series of hands-on, educational workshops focused on showing you step-by-step how to use FiftyOne, the open source toolkit that enables you to build better computer vision workflows by improving the quality of your datasets and delivering insights about your models. Make sure to register to reserve your spot for the first virtual workshop, [Getting Started with FiftyOne](https://voxel51.com/computer-vision-events/getting-started-with-fiftyone-workshop/?utm_source=blog), happening on March 29 at 10 AM Pacific. Your instructor will be [Jacob Marks, PhD](https://www.linkedin.com/in/jacob-marks/), machine learning engineer, prolific blogger, and member of the Developer Relations team here at Voxel51. ## Wait, what’s FiftyOne? For the uninitiated, [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## About the workshop Delivering production-grade AI requires high-quality datasets and high-performing models. To get there, data engineers and scientists need the right tools to visualize datasets and interpret models faster and more effectively. We created the “Getting Started with FiftyOne” workshop to help you gain greater visibility into the quality of your computer vision datasets and models. This 90 minute workshop is broken up into two main sections, lecture and lab. All attendees will get access to the tutorials, videos, and code examples used in the workshop. In the interactive “lecture” portion of the workshop we’ll cover the following topics: - FiftyOne Basics (terms, architecture, installation, and general usage) - An overview of useful workflows to explore, understand, and curate your data - How FiftyOne represents and semantically slices unstructured computer vision data ![](https://cdn.sanity.io/images/h6toihm1/production/5d8b3d04cae5e0a1c12b2d83da647c7d62f44dc7-1523x1046.png?auto=format&dpr=2&fit=max&q=75&w=1523) In the hands-on “lab” portion of the workshop we’ll take what we learned in the lecture and put it into action: - Loading datasets from the FiftyOne [Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html) - How to easily navigate the FiftyOne App’s features - Programmatically inspecting attributes of a dataset - Adding new sample and custom attributes to a dataset - Generating and evaluating model predictions - How to save insightful views into the data ![](https://cdn.sanity.io/images/h6toihm1/production/7a96144a1be2b0b2d7dc10994507899a69888b96-917x605.gif?auto=format&dpr=2&fit=max&q=75&w=917) Throughout the workshop there will be “pop quizzes” to test your knowledge. ![](https://cdn.sanity.io/images/h6toihm1/production/4e6182c345ded4f8eb4be1c8b7f3242aa3a399a9-1041x546.png?auto=format&dpr=2&fit=max&q=75&w=1041) At the end of the workshop you’ll have learned how to install FiftyOne, work with the Python SDK and App, plus perform basic tasks like importing datasets, creating views, and drawing out insights from your data and models. ## Prerequisites for the workshop To get the most out of the workshop we recommend you have a working knowledge of Python and a basic understanding of computer vision concepts. Familiarity with FiftyOne is a bonus, but have no fear, getting you proficient with FiftyOne basics is what this workshop is all about! We do highly recommend installing FiftyOne prior to the workshop so that you can make the most of the lab portion. Installing FiftyOne is simple and takes just a few minutes! Visit the [FiftyOne installation](https://docs.voxel51.com/getting_started/install.html) docs to get up and running in your preferred environment. ![](https://cdn.sanity.io/images/h6toihm1/production/0b42c9111ab2e6b9cf510251eff787b957c7f675-1728x1080.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## About the FiftyOne computer vision workshop series Look for additional virtual and in-person workshops to be added to the [Voxel51 events page](https://voxel51.com/computer-vision-events/) in the coming weeks. Lectures and labs will include following topics: - Bring Your Own Data: Load, import, and export your custom data in FiftyOne - How to send annotation tasks and load LabelBox annotations in FiftyOne - Exploring embeddings to uncover dataset insights - Curating better datasets by finding duplicate images and objects - Improving the quality of datasets by finding annotation mistakes - Tips and tricks for working with video data - Conducting semantic search at scale with vector search engines - Enter the third dimension: working with point clouds and more - And more! [events](https://voxel51.com/blog/tag/events) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [getting started with FiftyOne](https://voxel51.com/blog/tag/getting-started-with-fiftyone) [training](https://voxel51.com/blog/tag/training) [workshop](https://voxel51.com/blog/tag/workshop) [workshop series](https://voxel51.com/blog/tag/workshop-series) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/83d7ece0c635d6204dd333b0175cc55ec02c8c4a-1199x675.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Getting Started with FiftyOne Workshop – March 29 Recap\\ \\ Event Recaps\\ \\ • \\ \\ Apr 4, 2023](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap) [![](https://cdn.sanity.io/images/h6toihm1/production/622b7369c791083b44e3034b2b8772d3ecada8bb-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – November 2023\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Nov 1, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-november-2023) [FiftyOne Computer Vision Community Update – February 2024\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Feb 9, 2024](https://voxel51.com/blog/fiftyone-computer-vision-community-update-february-2024) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-269-lllmstxt|> ## Visualizing 3D Point Clouds [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tutorials](https://voxel51.com/blog/category/tutorials) A Better Way to Visualize 3D Point Clouds and Work with OpenAI’s Point-E Mar 22, 2023 • 6 min read Article content In this article [How to visualize point clouds, create orthographic projections, and evaluate detections with the latest release of FiftyOne](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#7e933c79b2f1) [Preview 3D data with orthographic projections](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#517ec0052fdf) [Group point cloud slices](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#20aace6a16ec) [Evaluate 3D object detections](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#36a73634e281) [Create point cloud-only datasets](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#6b064dbec629) [(3D point cloud) synthesis](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#8c9626fc3030) [Join the FiftyOne community!](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#3fb036f3db9d) In this article [How to visualize point clouds, create orthographic projections, and evaluate detections with the latest release of FiftyOne](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#7e933c79b2f1) [Preview 3D data with orthographic projections](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#517ec0052fdf) [Group point cloud slices](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#20aace6a16ec) [Evaluate 3D object detections](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#36a73634e281) [Create point cloud-only datasets](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#6b064dbec629) [(3D point cloud) synthesis](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#8c9626fc3030) [Join the FiftyOne community!](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e#3fb036f3db9d) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## How to visualize point clouds, create orthographic projections, and evaluate detections with the latest release of FiftyOne ![](https://cdn.sanity.io/images/h6toihm1/production/d888c320e21c5b809f89e7b22069a65f91e10f0f-3446x1810.gif?auto=format&dpr=2&fit=max&q=75&w=1600) 3D perception in computer vision enables computers and machines to understand the depth and structure of the 3D world around us, just as we do. Work in this field is exciting, with limitless potential for applications to revolutionize the way we live and work across all industries, from automotive to virtual reality. At the heart of 3D imaging applications, point clouds are used to efficiently represent three-dimensional spatial data. That’s why recent years have seen a flood of algorithms for processing, understanding, and making predictions using _point clouds_. These point clouds can be generated via either laser scanning techniques, such as lidar, or photogrammetry, or even via generative techniques, such as OpenAI’s recently released [Point-E](https://github.com/openai/point-e). The latest release of the FiftyOne computer vision toolset, 0.20, includes enhanced point cloud support to deliver unprecedented access to and control over your 3D data. [FiftyOne 0.20](https://voxel51.com/blog/announcing-fiftyone-0-20/) ships with the following 3D functionality: - [Orthographic projections](https://en.wikipedia.org/wiki/Orthographic_projection), including bird’s eye view (BEV) - Support for point cloud-only datasets in the FiftyOne App - Multiple point cloud slices in grouped datasets - Enhanced rendering and customization - Support for evaluating 3D object detection predictions These features add to FiftyOne’s existing 3D capabilities for working with and visualizing point clouds. Read on and learn how to harness FiftyOne to inspect, explore, and interact with your 3D data! ## Preview 3D data with orthographic projections Do you ever have a bunch of 3D samples that you want to rapidly peruse? Perhaps you want a [bird’s eye view](https://medium.com/syncedreview/learning-to-map-vehicles-into-birds-eye-view-dbf4d3de513) (BEV) of autonomous driving scenes from the [KITTI Vision Benchmark Suite](https://www.cvlibs.net/datasets/kitti/), [nuScenes](https://www.nuscenes.org/), or [Waymo Open Dataset](https://waymo.com/open/)? Or perhaps you’re working with a dataset of indoor scenes such as the [Stanford Large-Scale 3D Indoor Spaces Dataset](http://buildingparser.stanford.edu/dataset.html#Download), and you want an _elevation view_ into the scene. Our new [3D utils](https://docs.voxel51.com/api/fiftyone.utils.utils3d.html) integrate this functionality into the FiftyOne library and the FiftyOne App via the `compute_orthographic_projection_images()` method. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3import fiftyone.utils.utils3d as fou3d 4 5dataset = foz.load_zoo_dataset("quickstart-groups") 6 7min_bound = (0, -15, -2.73) 8max_bound = (20, 15, 1.27) 9size = (608, -1) 10 11fou3d.compute_orthographic_projection_images( 12 dataset, 13 size, 14 "bev_images", 15 shading_mode="height", 16 bounds=(min_bound, max_bound) 17) 18 19session = fo.launch_app(dataset) ``` By default, this method generates bird’s eye view projections of your point clouds, which then show up in the FiftyOne App as previews of point cloud samples (with filterable projections of polylines and bounding boxes as well). Also note that we’ve passed in `bounds`, telling FiftyOne where to crop the generated images, as well as a `shading_mode`, specifying that the point cloud’s intensity should be used to color the projection (as opposed to the height values, or colors of the individual points). ![](https://cdn.sanity.io/images/h6toihm1/production/67c55a341d29967fec93200618be7db3792156b9-1728x876.gif?auto=format&dpr=2&fit=max&q=75&w=1600) If you’d like, you can also pass in a normal vector to specify the plane with respect to which the routine should perform the projection, for instance, try `projection_normal=(0.5, 0.5, 0.)`! Inside of the [3D visualizer](https://docs.voxel51.com/user_guide/groups.html#using-the-3d-visualizer), you can also control a variety of characteristics of the look and feel, including setting point size and turning grid lines on or off. ![](https://cdn.sanity.io/images/h6toihm1/production/3d1b0b8deaf51596a41bd4419695b745645f9166-2044x1080.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## Group point cloud slices In FiftyOne, [grouped datasets](https://docs.voxel51.com/user_guide/groups.html) allow you to combine samples - potentially with varied media types (image, video, and point cloud) - in _groups_, with samples occupying different _slices_. New in this release, FiftyOne has revamped grouped datasets so that groups can have multiple point cloud samples. This can come in handy in a variety of scenarios, including: - Multiple point clouds for the same scene, coming from different sensors - Examining the effect of subsampling point clouds with millions of points - Transforming point clouds by rotation or scaling operations - Coloring points by cluster index or semantic segmentation label Let’s see this in action, clustering our point clouds with [DBSCAN](https://en.wikipedia.org/wiki/DBSCAN). We’ll create a new group slice, and add a new sample to each group in the `pcd_cluster` group slice. See [this gist](https://gist.github.com/jacobmarks/99a991659ed30434d55d514184e3eca6) for the corresponding code. Then we compute the orthographic projection images for these new point clouds - but this time, we pass in `shading_mode=rgb`, because we’ve used the point cloud’s RGB channels to encode cluster numbers. We’ll also use slightly different bounds, so it is easier to see the clusters. ```python 1min_bound = (0, -10, -2.73) 2max_bound = (20, 10, 1.27) 3 4fou3d.compute_orthographic_projection_images( 5 dataset, 6 size, 7 "/tmp/bev_cluster_images", 8 in_group_slice="pcd_cluster", 9 shading_mode="rgb", 10 bounds=(min_bound, max_bound) 11) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/31fafddbb9461ed0a3d8c7b3b5403b9a7df1bdee-1728x796.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## Evaluate 3D object detections If you’re familiar with FiftyOne’s [Evaluation API](https://docs.voxel51.com/user_guide/evaluation.html), you’ll know that the `evaluate_detections()` already supported the 2D object detection bounding boxes in image and video datasets. Today’s release extends these capabilities to 3D bounding boxes, with arbitrary rotation angles. FiftyOne automatically recognizes when the bounding box is three dimensional, and applies the appropriate method to compute 3D intersection over union (IoU) scores, which are used to determine whether a prediction agrees with a ground truth object. As an example, here is some sample code to generate a point cloud only dataset with 50 samples, and 10 ground truth bounding boxes per sample. To generate predictions, we randomly perturb some of the ground truth bounding boxes, and omit others from our predicted detections. The code can be found in [this gist](https://gist.github.com/jacobmarks/6b0c26b3771bf1d50fe19ca7c19de921). ![](https://cdn.sanity.io/images/h6toihm1/production/bed52614c96455cbf9163a72f66df926bef4a5fd-3438x1794.png?auto=format&dpr=2&fit=max&q=75&w=1600) We can then evaluate our detection predictions with `evaluate_detections()`: ```python 1results = dataset.evaluate_detections( 2 "predictions", 3 eval_key="eval" 4) ``` And print out a report on dataset-level evaluation metrics: ```python 1results.print_report() ``` precision recall f1-score support dog 0.83 0.66 0.74 500 micro avg 0.83 0.66 0.74 500 macro avg 0.83 0.66 0.74 500 weighted avg 0.83 0.66 0.74 500 As with 2D detection evaluations, you can specify what IoU threshold to use during evaluation by passing in the `iou` argument: ```python 1results_high_iou = dataset.evaluate_detections( 2 "predictions", 3 iou=0.75 4) 5 6results_high_iou.print_report() ``` precision recall f1-score support dog 0.12 0.09 0.10 500 micro avg 0.12 0.09 0.10 500 macro avg 0.12 0.09 0.10 500 weighted avg 0.12 0.09 0.10 500 Once you have evaluated your object detection predictions, you can also isolate _evaluation_ _patches_ containing, for instance, false positive predictions: ```python 1eval_patches = dataset.to_evaluation_patches("eval") 2fp_patches = eval_patches.match(F("type") == "fp") ``` You can then sort by prediction confidence to identify your highest confidence false positive predictions: ```python 1high_conf_fp_view = fp_patches.sort_by("predictions.confidence") ``` ## Create point cloud-only datasets Previously, FiftyOne supported point clouds in a grouped dataset along with other media. However, point clouds are first class citizens. As such, the FiftyOne App now supports point cloud only datasets! One situation in which this might be useful, for instance, is if you’re generating point clouds from scratch. Let’s see this with an example, using OpenAI’s Point-E to turn text prompts into three dimensional point cloud models. We use the sampler from the Point-E [text2pointcloud](https://github.com/openai/point-e/blob/main/point_e/examples/text2pointcloud.ipynb) example notebook, and convert the resulting point clouds using [Open3d](http://www.open3d.org/). ```python 1def generate_pcd_from_text(prompt): 2 samples = None 3 for x in sampler.sample_batch_progressive( 4 batch_size=1, 5 model_kwargs=dict(texts=[prompt]) 6 ): 7 8 samples = x 9 10 pointe_pcd = sampler.output_to_point_clouds(samples)[0] 11 channels = pointe_pcd.channels 12 r, g, b = channels["R"], channels["G"], channels["B"] 13 colors = np.vstack((r, g, b)).T 14 points = pointe_pcd.coords 15 16 pcd = o3d.geometry.PointCloud() 17 pcd.points = o3d.utility.Vector3dVector(points) 18 pcd.colors = o3d.utility.Vector3dVector(colors) 19 return pcd ``` Then we generate an example dataset in FiftyOne, assigning each point cloud a random filename: ```python 1def generate_random_filename(): 2 rand_str = str(uuid.uuid1()).split('-')[0] 3 return "pointe_vehicles/" + rand_str + ".pcd" 4 5def generate_dataset(num_samples = 100): 6 vehicles = ["car", "bus", "bike", "motorcycle"] 7 colors = ["red", "blue", "green", "yellow", "white"] 8 9 samples = [] 10 for i in tqdm(range(num_samples)): 11 vehicle = random.choice(vehicles) 12 cols = random.choices(colors, k=2) 13 prompt = f"a {cols[0]} {vehicle} with {cols[1]} wheels" 14 pcd = generate_pcd_from_text(prompt) 15 ofile = generate_random_filename() 16 o3d.io.write_point_cloud(ofile, pcd) 17 18 sample = fo.Sample( 19 filepath = ofile, 20 tags = cols, 21 vehicle_type = fo.Classification(label = vehicle) 22 ) 23 samples.append(sample) 24 25 dataset = fo.Dataset("point-e-vehicles") 26 dataset.add_samples(samples) 27 return dataset ``` All that is left to do is compute the orthographic projections. Here we will use a non-default `projection_normal` so that our preview image is not a bird’s eye view: ```python 1import fiftyone.utils.utils3d as fou3d 2fou3d.compute_orthographic_projection_images( 3 dataset, 4 (-1, 608), 5 "/tmp/side_images", 6 shading_mode="rgb", 7 projection_normal = (0, -1, 0) 8) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/9a4d0b067757a99c767f977be724c22cf9bf469b-2066x1080.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## (3D point cloud) synthesis If you aren’t working with your 3D point clouds in FiftyOne, you’re missing out. Visualize your point clouds in the same place that you visualize your images, videos, geo data, and more. Dive deeper and learn how to [use Point-E point cloud synthesis with FiftyOne to generate your own 3D self-driving dataset](https://docs.voxel51.com/tutorials/pointe.html)! ![](https://cdn.sanity.io/images/h6toihm1/production/2291806e685455a3e7ccaff1f215e65c28ec3a28-1726x862.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,400+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,700+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,500+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 258+ repositories - 58+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [3D point cloud](https://voxel51.com/blog/tag/3d-point-cloud) [birds eye view](https://voxel51.com/blog/tag/birds-eye-view) [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [DBSCAN](https://voxel51.com/blog/tag/dbscan) [Diffusion models](https://voxel51.com/blog/tag/diffusion-models) [Evaluation](https://voxel51.com/blog/tag/evaluation) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [OpenAI](https://voxel51.com/blog/tag/openai) [point cloud](https://voxel51.com/blog/tag/point-cloud) [point cloud synthesis](https://voxel51.com/blog/tag/point-cloud-synthesis) [Point-E](https://voxel51.com/blog/tag/point-e) [quickstart dataset](https://voxel51.com/blog/tag/quickstart-dataset) [quickstart groups dataset](https://voxel51.com/blog/tag/quickstart-groups-dataset) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/803cb935ddffbb5b29b6d3c73104b3a1221ddbfc-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 3)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3) [![](https://cdn.sanity.io/images/h6toihm1/production/3f54d0a45faa06a04b5d0244dd7c092603150cf0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Mar 10, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 11, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/2867ac2853fae5362ca6bd2d208358dc94556344-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Google Search Experience for Computer Vision Data\\ \\ Tutorials, Vector Search\\ \\ • \\ \\ Mar 22, 2023](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-270-lllmstxt|> ## Google Search for Computer Vision [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tutorials](https://voxel51.com/blog/category/tutorials), [Vector Search](https://voxel51.com/blog/category/vector-search) A Google Search Experience for Computer Vision Data Mar 22, 2023 • 8 min read Article content In this article [How to Use Vector Search Engines, NLP, and OpenAI's CLIP in FiftyOne](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#2f054d6dc447) [Search for similar images](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#c7e8577503e1) [Natural language search](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#5a2cf42c0350) [Search at scale with Pinecone and Qdrant](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#6a123555aaa3) [Mix, match, frame, and patch](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#0e7d5e6c8b86) [Conclusion](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#2284cd44fce6) [Join the FiftyOne community!](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#17c0c87774f9) In this article [How to Use Vector Search Engines, NLP, and OpenAI's CLIP in FiftyOne](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#2f054d6dc447) [Search for similar images](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#c7e8577503e1) [Natural language search](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#5a2cf42c0350) [Search at scale with Pinecone and Qdrant](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#6a123555aaa3) [Mix, match, frame, and patch](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#0e7d5e6c8b86) [Conclusion](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#2284cd44fce6) [Join the FiftyOne community!](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#17c0c87774f9) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## How to Use Vector Search Engines, NLP, and OpenAI's CLIP in FiftyOne Have you ever wanted to find the images most similar to an image in your dataset? What if you haven’t picked out an illustrative image yet, but you can describe what you are looking for using natural language? And what if your dataset contains millions, or tens of millions of images? With today’s release of [FiftyOne .20,](https://voxel51.com/blog/announcing-fiftyone-0-20/) you can now natively search through your computer vision data, at scale, with either images or text! What’s more, these new querying capabilities interface seamlessly with FiftyOne’s existing semantic slicing operations. Restrict your queries to filtered and matched views of your data with ease, and compose queries however you’d like. FiftyOne .20 brings these exciting vector search capabilities: - [Native integration with Pinecone](https://docs.voxel51.com/integrations/pinecone.html) vector search database - [Native integration with Qdrant](https://docs.voxel51.com/integrations/qdrant.html) vector search engine - SDK support for querying by sample ID, query vector, or text prompt - UI support for querying with images extended to use vector search engines - UI support for querying with natural language with vector search engines In this blog post we show you how to apply the new vector search functionality across these workflows: - [Search for similar images](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#image-search) - [Search with natural language](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#language-search) - [Scale up your search with Pinecone and Qdrant](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#vector-index-search) - [What’s possible with vector search?](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data#mix-and-match) Read on to learn how you can take your dataset exploration, understanding, and curation to the next level with vector search! ## Search for similar images FiftyOne has long supported tasks that help you understand relationships between images in your dataset. With FiftyOne, you can - Find both near and exact [duplicate images](https://docs.voxel51.com/recipes/image_deduplication.html), and [duplicate objects](https://docs.voxel51.com/recipes/remove_duplicate_annos.html) in your dataset - [Identify the most “unique” samples](https://docs.voxel51.com/tutorials/uniqueness.html) in your dataset - [Visualize and explore your data](https://docs.voxel51.com/tutorials/image_embeddings.html) with embeddings and dimensionality reduction techniques At the heart of these methods lie _embeddings_. Because embeddings are foundational to all of the new vector search features in FiftyOne .20, it’s worth taking a moment to summarize what embeddings are and how you can work with them in FiftyOne. Embeddings are vectors (typically between 512 and 2048 entries) that are generated one-to-one as a [smaller representation of your data](https://towardsdatascience.com/neural-network-embeddings-explained-4d028e6f0526). For example, a 512 dimensional vector embedding can represent a 12MP image. To compute the similarity between images - and perform the subsequent similarity search - you must specify how to embed your samples. In FiftyOne, you can do so by passing either the embedding vectors themselves, or the model with which to perform the embedding, into `compute_similarity()` which creates the “similarity index”, or vector index, containing embeddings for your images. Any of the following options work: - **Store embeddings in a field**: compute embeddings in a field on your samples, and point to that field. - **BYO embeddings**: compute embeddings however you like, and pass these in as a two-dimensional array. - **Pass in a model from the FiftyOne Model Zoo** - **Pass in a model name** Once you have generated your similarity index, you can query your dataset by sample ID in Python with `sort_by_similarity()`. To find the 25 most similar images to the first sample in the dataset, for instance: Alternatively, we can perform the same operation in the FiftyOne App. When we tick the checkbox in the upper left corner of the first sample, the menu bar above it expands with new options, including a new `image` icon. Click on this icon and hit enter. It really is that easy! If we wanted a view containing just the ten most similar images, we could set this by clicking on the gear icon and replacing 25 (the default) with 10. When creating our similarity index, we can also choose the metric for determining closeness with the `metric` keyword. Additionally, what we specify as the `brain_key` can then be used to retrieve the similarity index. If we have used: We could instantiate the similarity index with: When you instantiate a similarity index, you can then use that index to [find unique samples](https://docs.voxel51.com/api/fiftyone.brain.similarity.html#fiftyone.brain.similarity.SimilarityResults.find_unique) with `sim_index.find_unique()`, or [find duplicate samples](https://docs.voxel51.com/api/fiftyone.brain.similarity.html#fiftyone.brain.similarity.SimilarityResults.find_duplicates), with `sim_index.find_duplicates()`. ## Natural language search Building on our [recent blog post](https://medium.com/voxel51/finding-images-with-words-92b078314ed1), where we showed how you can “find images with words” with CLIP embeddings, natural language queries are now natively supported, both in the FiftyOne software development kit and in the FiftyOne App. Any model in the [FiftyOne Model Zoo](https://voxel51.com/docs/fiftyone/user_guide/model_zoo/index.html) that can embed both language and images (such as CLIP) can be used to search through an image dataset with text prompts. Simply pass the name of the model into your call to `compute_similarity()`. Here, we use cosine similarity as our metric for determining closeness, and we store the similarity index with `brain_key=”clip”`. In Python, we can now query our dataset with a text prompt by passing a string into `sort_by_similarity()` as a query. The following Python query creates a `DatasetView` with the 25 images that most resemble the prompt “a piece of pie”, as determined by our similarity metric: Alternatively, we can perform the same search in the FiftyOne App without writing a single line of code. If we reset the session ( `session = fo.launch_app(dataset)`), we can recreate `pie_view` by clicking on the magnifying glass, typing our query directly into the search field that appears, and hitting enter. We can then click the bookmark icon to convert this into a view. If you reset the view (by clicking on the magnifying glass and then clicking the `reset` button), and instead click on the gear icon after the magnifying glass, you will see a radio button with a single option: “clip”. This is telling you that FiftyOne is performing the natural language query using the similarity index you stored with `brain_key=”clip”`. In this example, we only have one similarity index on the dataset with a model that supports prompts, but we have a great deal of freedom in how we name and populate these similarity indexes. We can have different indexes for the same model with different metrics, multiple models that support text prompts, or both. If you click on the gear icon after the image icon, you will see that the radio button has two options, because there are now two similarity indexes that support image similarity searches. You can also load a model from the FiftyOne Model Zoo with custom weights, or even [add your own model to the zoo](https://docs.voxel51.com/user_guide/model_zoo/index.html#using-custom-models)! ## Search at scale with Pinecone and Qdrant By default, FiftyOne uses Scikit-learn (Sklearn) to compute the nearest neighbors of embedding vectors and construct the similarity index. However, the larger a dataset gets, the harder it becomes to find the _best_ matches to a particular vector search query. It can become so difficult, in fact, that there are no exact methods known with [favorable scaling](https://en.wikipedia.org/wiki/Nearest_neighbor_search) in preprocessing time, search time, and memory consumption. When your datasets start to reach millions of samples, exact nearest neighbor search can become infeasible. But often - especially when the datasets become that large - retrieving _approximately the best_ matches is all you need. That’s where vector search engines like Pinecone and Qdrant come in with approximate nearest neighbor (ANN) search. In particular, these vector search engines implement the [hierarchical navigable small world](https://arxiv.org/abs/1603.09320) (HNSW) algorithm. In [Finding Images with Words](https://medium.com/voxel51/finding-images-with-words-92b078314ed1) and [Nearest Neighbor Embeddings Classification with Qdrant](https://docs.voxel51.com/tutorials/qdrant.html), we just scratched the surface of what FiftyOne plus Pinecone or Qdrant can do for your computer vision workflows. With today’s FiftyOne .20 release, you can now use these _backends_ to generate your similarity indexes in FiftyOne, accelerating your dataset exploration and evaluation. You can now pass a `backend` argument into `compute_similarity()` to tell FiftyOne what vector search engine to use to create the similarity index: `backend=”qdrant”` for Qdrant, and `backend=”pinecone”` for Pinecone. Depending on the backend, there are a variety of optional arguments you can pass in, which allow you to customize the replication factor, sharding, and the degree of approximation in _approximate_ nearest neighbor search. For a complete discussion of these arguments, see our [Qdrant integration](https://docs.voxel51.com/integrations/qdrant.html) and [Pinecone integration](https://docs.voxel51.com/integrations/pinecone.html) docs. You can also write these kwargs - including your credentials - to your FiftyOne Brain config file, instead of passing them in to `compute_similarity()` each time you want to create a new similarity index. ### Qdrant backend To get started using Qdrant in FiftyOne, pull the pre-built Docker image from DockerHub and run the container: You’ll also need to install the Qdrant Python client: Then create a similarity index, passing in the name of the Qdrant collection to be created: ### Pinecone backend To get started using Pinecone in FiftyOne, set up an account [here](https://www.pinecone.io/) if you don’t have one already, and copy an API key from [here](https://app.pinecone.io/organizations). Then install the Pinecone Python client: Just like that, you can use Pinecone vector search as a backend, specifying an `index_name` if desired: **Note**: Pinecone’s free tier only allows users a single vector index at a time. You may need to delete your existing vector index before creating one for your FiftyOne dataset. ## Mix, match, frame, and patch Native vector search on your computer vision data becomes even more valuable when you can interleave these queries with other semantic slicing operations like matching, filtering, sorting, and shuffling your data. In FiftyOne, all of these logical operations - including vector search queries - are represented as view stages, and can be combined in whatever order you’d like! Here are just a few examples of what you can do with vector search in FiftyOne: ### Find race car tires in a general purpose dataset There are many workflows you could use for this, for instance: 1. Filter the dataset for images containing positive `Car` labels 2. Visually identify a race car and use image similarity search to find other images with race cars 3. Convert to object patches 4. Use natural language search to query for object patches with the text prompt “tire” With the random subset of Open Images, all we need to do to run this workflow is run `compute_similarity()` on the `detection` object patches with a model that supports prompts, such as CLIP: ### Find video frames with cars in an intersection Again, there are many ways to do this. Perhaps the simplest is to use FiftyOne’s `ToFrames` ViewStage to generate a view containing one image per frame across all video samples. You can then construct a similarity index for this view using `compute_similarity()`, and query with the text prompt “cars in an intersection”: ## Conclusion In this blog post, we’ve introduced just a few of the myriad ways you can leverage vector search natively in FiftyOne to accelerate your computer vision workflows. Whether you want to query your computer vision data with images, natural language, or raw numerical vectors, FiftyOne has you covered. We’ve made it easy to search at scale with Qdrant or Pinecone, and you can query images, video frames, or object patches, intertwining vector search queries with filtering and matching operations to your heart’s desire. Explore your computer vision data like never before! ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,400+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,700+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,500+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 259+ repositories - 58+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [CLIP](https://voxel51.com/blog/tag/clip) [embeddings](https://voxel51.com/blog/tag/embeddings) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [filtering](https://voxel51.com/blog/tag/filtering) [natural language search](https://voxel51.com/blog/tag/natural-language-search) [OpenAI](https://voxel51.com/blog/tag/openai) [Pinecone](https://voxel51.com/blog/tag/pinecone) [Qdrant](https://voxel51.com/blog/tag/qdrant) [semantic search](https://voxel51.com/blog/tag/semantic-search) [similarity search](https://voxel51.com/blog/tag/similarity-search) [vector database](https://voxel51.com/blog/tag/vector-database) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/19eaad3d85784642bb3629627ea6081d1b5c56bc-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Computer Vision Interface for Vector Search\\ \\ Product & News, Vector Search\\ \\ • \\ \\ Jul 12, 2023](https://voxel51.com/blog/the-computer-vision-interface-for-vector-search) [![](https://cdn.sanity.io/images/h6toihm1/production/25ed1525cb4a1797abb9edcdf5ac3d7b8239eaf3-2805x1581.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Concept Traversal Plugin for FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Oct 19, 2023](https://voxel51.com/blog/computer-vision-concept-traversal-plugin-for-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/c332c478d66b51893447f19eb71d84a940b94a09-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Finding Images with Words\\ \\ Computer Vision, Vector Search\\ \\ • \\ \\ Jan 11, 2023](https://voxel51.com/blog/finding-images-with-words) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-271-lllmstxt|> ## FiftyOne Workshop Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Getting Started with FiftyOne Workshop – March 29 Recap Apr 4, 2023 • 9 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap#a4917b42ff47) [New workshops announced!](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap#919925a826de) [Workshop summary](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap#949a2a82cc64) [Q&A recap](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap#4782a5443f89) [Additional resources](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap#bee86b4b5e4c) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap#a4917b42ff47) [New workshops announced!](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap#919925a826de) [Workshop summary](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap#949a2a82cc64) [Q&A recap](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap#4782a5443f89) [Additional resources](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap#bee86b4b5e4c) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Earlier this week, [Jacob Marks](https://www.linkedin.com/in/jacob-marks/), PhD and Machine Learning Engineer at Voxel51, presented the workshop: Getting Started with FiftyOne. This workshop was the first in a series of hands-on, educational events focused on showing you step-by-step how to use FiftyOne. In this post, we summarize the workshop, recap the questions and their answers that came up during the event, and share upcoming dates for the workshop in case you want to join or share them with colleagues. ## Wait, what’s FiftyOne? The Getting Started with FiftyOne Workshop is all about the FiftyOne toolset. But if you’re new to FiftyOne – you may be wondering, what is it? Data engineers and scientists need the right tools to visualize datasets and interpret models faster and more effectively. [FiftyOne](https://voxel51.com/fiftyone/) does just that – it is the open source machine learning toolkit that enables you to build better computer vision workflows by improving the quality of your datasets and delivering insights about your models, so that you can get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## New workshops announced! We are excited to announce the dates and times for three more [Getting Started with FiftyOne Workshops](https://voxel51.com/computer-vision-events/)! - April 26 @ 9:00 AM IST \[1:30 PM AEST / 03:30 UTC\] - May 31 @ 4 PM BST \[11 AM EDT / 15:00 UTC\] - June 28 @ 10 AM PDT \[1 PM EDT / 17:00 UTC\] In addition to the Getting Started with FiftyOne Workshops, we are also building out a catalog of advanced workshops to take you beyond the basics of getting started and into deeper ways FiftyOne can enhance and streamline your computer vision workflows. Stay tuned for future announcements with the schedule of advanced topics. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## Workshop summary We created the “Getting Started with FiftyOne” workshop to help you gain greater visibility into the quality of your computer vision datasets and models. The workshop was half lecture and half lab so that you would walk away with a solid understanding of the basics of the FiftyOne toolset, architecture, and popular workflows, as well as learn how to install FiftyOne, work with the Python SDK and App, and perform basic tasks like importing datasets, creating views, and drawing out insights from your data and models. ### **Lecture: Get up-to-speed on the basics** Jacob kicked off the workshop promptly at 51 o’clock 😂 and explained FiftyOne at a high level: “It helps you to visualize, clean, and curate your data, find hidden structure in that data, evaluate model predictions on your datasets, as well as different subsets of your datasets. And its design philosophy is all about flexibility and customizability. So FiftyOne is all about giving you the power to explore and understand your data, regardless of your specific workflow or machine learning pipeline.” Jacob then thoroughly walked us through the basic concepts summarized below. ### **Curate data** In the workshop, Jacob demonstrated some of the ways FiftyOne helps you to curate data: - Find: filter, match, sort, select - Remove: duplicates - Add: tags, metadata, predictions - Correct: annotation mistakes - Save: interesting “views” ![](https://cdn.sanity.io/images/h6toihm1/production/7b3da0c91f58f14aca255076428b7117463a408b-966x798.gif?auto=format&dpr=2&fit=max&q=75&w=966) ### **Understand data** Jacob also covered how FiftyOne helps you understand your data with: - Aggregate statistics: FiftyOne supports a variety of histograms and all of the traditional aggregations for numerical quantities you would expect: min, max, mean, standard aviation, and more. - Embeddings: Embeddings are numerical vector representations of certain aspects of the properties of our data. And those help us to understand our data in a lot of different ways. In the workshop, Jacob explored the Berkeley Deep Drive (BDD) dataset to show embeddings in action, including clusters of daytime images and nighttime images. Embeddings can help identify hidden structures in datasets that we may not otherwise be aware of. - Interactive visualization: Jacob noted that all these visualizations have been and are interactive. Examples: if you lasso points in an embeddings plot, then you’ll see just those samples; if you explore a cell in a confusion matrix, you can see just those samples. ![](https://cdn.sanity.io/images/h6toihm1/production/94c168b1408e66a8d61a767d02e7ba541cd34cd6-1235x770.gif?auto=format&dpr=2&fit=max&q=75&w=1235) ### **Evaluate data** Here Jacob explained: “Evaluating is a key component in many computer vision workflows. So FiftyOne has support for tons of one-number metrics: precision, recall, F1 score, intersection over union, you name it. There’s support for all of your favorite plots including PR curves and confusion matrices. You can also perform analysis on samples, labels, and entire datasets.” ![](https://cdn.sanity.io/images/h6toihm1/production/e0343288f9074357318ae0c4a84783864329fb9f-966x798.gif?auto=format&dpr=2&fit=max&q=75&w=966) ### **Tap into the flexibility of FiftyOne** Jacob noted: “FiftyOne's design philosophy is all about flexibility and customizability. Everything I've mentioned so far has flexibility surrounding that because we know that computer vision is not a one-size-fits-all solution field.” Jacob described FiftyOne’s flexibility with regards to: - Datasets - Models - Media types - Labels - Plugins - More! ![](https://cdn.sanity.io/images/h6toihm1/production/abad3ea7f3410ff83d15ccab154f7154384deea5-1324x710.gif?auto=format&dpr=2&fit=max&q=75&w=1324) ### **Key components of FiftyOne** Before the hands-on lab portion, Jacob gave a primer on the core components of FiftyOne that in the lab part we will see firsthand! \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop He also primed us on some additional basic concepts: - A comparison of tabular data (structured data) and computer vision data (unstructured data); FiftyOne is the pandas of computer vision - A look under the hood of a schema - including a dataset, samples, fields, metadata, filepath, labels, media type, and more ### **Lab: fire up FiftyOne and experience it for yourself!** The second half of the workshop was a hands-on lab, so you could put what you learned in the lecture into action. By the end of the workshop, attendees fired up FiftyOne and explored datasets and models firsthand. The lab portion of the workshop focused on enabling you to perform all the steps needed to achieve all of this: - Install FiftyOne - Load datasets and models from the FiftyOne [Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html) and FiftyOne Model Zoo - Easily navigate the FiftyOne App’s features - Programmatically inspect attributes of a dataset - Add new samples and custom attributes to a dataset - Evaluate model predictions - Save insightful views into the data - More! ## Q&A recap **If we manually create a view in the GUI, can this be exported somehow, to be used later on a different machine?** With open source FiftyOne, any views that you save in the GUI can be pulled up in Python on the same machine in the future. If you want to share those views with others, you can serialize views to JSON, transfer them, then rebuild them: ```python 1stages = view._serialize() 2still_view = fo.DatasetView._build(dataset, stages) ``` If you expect to be frequently sharing views with other people, and/or working on multiple machines, then you may want to consider FiftyOne Teams, which has built-in support for sharing datasets and views. **Can FiftyOne App be rendered inside JupyterLab?** Yep! It gets rendered in an output cell that you can pop out and move into different tabs within JupyterLab. FiftyOne also supports Colab notebooks, Databricks notebooks, and more. **Does the MongoDB connection work inside a JupyterLab environment?** Yes! On the backend, FiftyOne uses a non-relational database structure with MongoDB. The database gets launched in a separate process, even when launched inside of the JupyterLab environment. You can specify your MongoDB configuration within a Jupyter notebook. Learn [how to configure MongoDB manually](https://docs.voxel51.com/user_guide/config.html#configuring-a-mongodb-connection). **Do you have any docs to find out more about the Colab or Databricks integration?** Yes! Visit the [docs on notebook environments](https://docs.voxel51.com/environments/index.html#notebooks). Additionally, you can try [FiftyOne in Colab right in your browser](https://colab.research.google.com/github/voxel51/fiftyone-examples/blob/master/examples/quickstart.ipynb). **What’s the typical cadence of new releases?** Major releases are primarily made available around the completion of new features. However, between major releases, there are more frequent minor releases to handle bugs. You can check out the release notes for all versions here: [https://docs.voxel51.com/release-notes.html](https://docs.voxel51.com/release-notes.html) **I want to use the FiftyOne App to quickly look at my YOLOv5 dataset and annotations. I'd like to quickly analyze my data and do some very basic tasks like remove duplicates, show distributions, etc. Can you explain what can be done natively in the App vs what requires the Python SDK?** Today, you need to first load your dataset into FiftyOne either through the Python SDK or the command-line interface. But no need to be afraid, [loading a YOLOv5 dataset takes only a few lines of code](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#yolov5dataset). Once your dataset is in FiftyOne, you can then visualize it in the App, view distributions, filter on your labels, etc. Noting that finding duplicate samples will require [a couple more lines of code to compute similarity](https://docs.voxel51.com/user_guide/brain.html#finding-near-duplicate-images). **Is there a tutorial for integration on FiftyOne and Label Studio?** Yes! Check out the [Label Studio integration and examples](https://docs.voxel51.com/integrations/labelstudio.html) in the docs. And if you’re curious about integrations with other annotation tools, FiftyOne also integrates with [CVAT](https://docs.voxel51.com/integrations/cvat.html), [Labelbox](https://docs.voxel51.com/integrations/labelbox.html), and [Scale](https://docs.voxel51.com/api/fiftyone.utils.scale.html). **To get uniqueness measures and those cool scatterplots, do we compute the embeddings once? Or do we have to compute embeddings each time for uniqueness, similarity, visualizations, etc.?** Great question! You can compute the embeddings once and then reuse them for uniqueness, similarity, visualizations, etc. All of these FiftyOne Brain methods allow you to specify embeddings in multiple ways. You could provide a FiftyOne Zoo Model, in which case the embeddings would be generated, but you can also provide a NumPy array of precomputed embeddings, or a field of your dataset which contains the embeddings for each sample. For example, for similarity: ```python 1# Compute embeddings each time 2results = fob.compute_similarity(dataset, model=foz.load_zoo_model(...), ...) 3 4# Compute embeddings once and reuse them 5results = fob.compute_similarity(dataset, embeddings=np.array(...), ...) ``` Learn more about this [similarity example in the docs](https://docs.voxel51.com/api/fiftyone.brain.html#fiftyone.brain.compute_similarity). **Can the underlying data in those histograms be extracted using Python?** Yes! The underlying data in these histograms can absolutely be extracted using Python. Visit the docs on [using aggregations](https://docs.voxel51.com/user_guide/using_aggregations.html#histogram-values) to learn how. **Is it possible to convert an image dataset to a video dataset?** Yes, you can convert an image dataset to a video dataset. There are many different formats for videos, so you would need to be thoughtful about the way that you did that, but it is absolutely possible. Additionally, a video dataset in FiftyOne does require a video media file (ex: an .mp4 file) today. So if you have a dataset that is a collection of images, you could convert those to videos with something like FFmpeg, then load those videos into FiftyOne. Stay tuned for updates on this in the near future. **How many images can I browse with FiftyOne? Is there an upper limit?** There's no limit to the number of images you can browse in FiftyOne. We frequently see users with 10+ million samples. Though there are two axes to consider in terms of performance: number of samples and number of fields. When you have a billion detections on a dataset, a filter query that touches each of them will take some time. The sweet spot for snappy FiftyOne usage is on the order of hundreds of thousands of samples and dozens of fields. A common use case when datasets are larger than this is to have a data lake dataset with all samples in it, and then smaller working dataset clones that you actively work with. See [dataset cloning docs here](https://docs.voxel51.com/user_guide/using_datasets.html#cloning-datasets). **If I run an evaluation over a complete dataset, is it possible to obtain metrics for a filtered DatasetView without having to rerun the complete evaluation? (For example, running evaluations for a dataset containing data from all countries, and later obtaining/extracting per-country metrics (precision, recall, mAP).)** Currently, if you want to evaluate a subset into a dataset, then you need to perform that evaluation separately. The evaluation of detections can change depending on the ground truth/predicted labels that exist. So two bounding boxes that were matched in one view may be matched differently in another. We plan to add more flexibility around this in the future. **Is there an easy way to import YOLO predictions into FiftyOne?** Yes! There is a [one-line way to do it here](https://docs.voxel51.com/user_guide/dataset_creation/index.html#model-predictions). And you can learn even more about working with YOLOv8 (and therefore also v5 because they share the same format) [in the YOLO tutorial](https://docs.voxel51.com/tutorials/yolov8.html). For other YOLO versions, reach out in [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) and we can assist. ## Additional resources If you missed the workshop or would like to revisit it, here are some additional resources for you: - The [Colab notebook](https://colab.research.google.com/drive/1dzvfCdza5cmlFX5fIA452TkoZVNV4QZ8?usp=sharing) with the exercises from the hands-on lab - [Join us for a future workshop](https://voxel51.com/computer-vision-events/) Stay tuned for the video recap of the Getting Started with FiftyOne Workshop coming soon. [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [getting started with FiftyOne](https://voxel51.com/blog/tag/getting-started-with-fiftyone) [Getting Started with FiftyOne Workshop](https://voxel51.com/blog/tag/getting-started-with-fiftyone-workshop) [training](https://voxel51.com/blog/tag/training) [workshop](https://voxel51.com/blog/tag/workshop) [workshop series](https://voxel51.com/blog/tag/workshop-series) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/17737c0400f6f42ba0a33f0b809d3707a5df971c-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Announcing the FiftyOne Computer Vision Educational Workshop Series\\ \\ Product & News\\ \\ • \\ \\ Mar 17, 2023](https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series) [![](https://cdn.sanity.io/images/h6toihm1/production/308698a5aece1d5b1b95ee1bf52811b24448458c-1200x672.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne 0.18 for Computer Vision\\ \\ Event Recaps\\ \\ • \\ \\ Dec 6, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/17422a76c76f14096dce21e43da51945f498a811-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne & FiftyOne Teams\\ \\ Event Recaps\\ \\ • \\ \\ Oct 8, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-272-lllmstxt|> ## FiftyOne Community Update [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) FiftyOne Computer Vision Community Update – April ‘23 Apr 6, 2023 • 8 min read Article content In this article [Community Spotlights](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#46c2efcf40d4) [Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#98be7476744e) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#202159c11aae) [FiftyOne Community News](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#0d463ffa6936) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#853520ed7e42) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#e11e6a209aee) [Industry Spotlight](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#bbca7d82e25a) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#a326d8773050) [Community-Powered Charitable Contributions](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#afb8b5364dba) In this article [Community Spotlights](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#46c2efcf40d4) [Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#98be7476744e) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#202159c11aae) [FiftyOne Community News](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#0d463ffa6936) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#853520ed7e42) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#e11e6a209aee) [Industry Spotlight](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#bbca7d82e25a) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#a326d8773050) [Community-Powered Charitable Contributions](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023#afb8b5364dba) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the monthly blog series where we bring you up to speed on recent happenings in the FiftyOne community and celebrate noteworthy milestones. 🙌 🚀 ## Community Spotlights We love hearing how FiftyOne helps you solve challenges and reach new heights! Curious what sorts of use cases and computer vision workflows are possible with FiftyOne? Here are just a few highlights from what people in the community have to say. ![](https://cdn.sanity.io/images/h6toihm1/production/aba4ea3804ddde804885d3569426b8a3296527a5-294x55.png?auto=format&dpr=2&fit=max&q=75&w=294) [LiveReach](https://livereachmedia.com/) offers a premiere video intelligence platform featuring cloud-based enterprise video security and industry leading motion intelligence. > _"I’ve done extensive integration with FiftyOne and it is a game changer!"_ > > — Dean Webb, Head of Artificial Intelligence ![](https://cdn.sanity.io/images/h6toihm1/production/6830169cfb7c3e5351998872585a6dfc08e056d8-550x126.png?auto=format&dpr=2&fit=max&q=75&w=550) [Protex AI](https://www.protex.ai/) enables businesses to gain greater visibility of unsafe behaviors in their facilities. Protex AI's privacy-preserving platform plugs into existing CCTV infrastructure and uses computer vision technologies to capture unsafe events autonomously in settings such as warehouses, manufacturing facilities, and ports. > _"At Protex AI, we develop a platform to monitor worker health and safety using existing CCTV infrastructure. Our core computer vision technologies are object detection, classification, and pose estimation. We use FiftyOne as a vital component in our pipeline to validate dataset annotations, intelligently subsample datasets to ensure balance, and also to visualize and debug model predictions to assess accuracy.”_ > > — Patrick Rowsome, Lead Computer Vision Engineer ![](https://cdn.sanity.io/images/h6toihm1/production/7f251fb004a59d1572e1bda704d6810b9cd0ad7d-1380x812.png?auto=format&dpr=2&fit=max&q=75&w=1380) [G42](https://www.g42cloud.com/#/) is a leading AI & Cloud Computing company based in Abu Dhabi, working on projects from molecular medicine to space travel and everything in between. > _“FiftyOne provides a superior interface for dealing with computer vision data. Its extensive Python package lets you do almost any data transformation, perform similarity search, and easily evaluate model predictions.“_ > > — Rustem Galiullin, Data Scientist ## Community Rewards ![](https://cdn.sanity.io/images/h6toihm1/production/e3ed352e497d0e500ff4a1484b8422b3c9bef5cb-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) Is your organization using FiftyOne to solve interesting computer vision problems? [Share your success story](https://voxel51.com/fiftyone-computer-vision-success-story-submission/?utm_source=blog-cu) and claim a box of community rewards as a thank you! ## **Product Releases** ### FiftyOne 0.20 Is Here! The latest product release, [FiftyOne 0.20](https://voxel51.com/blog/announcing-fiftyone-0-20/?utm_source=blog-cu), is here and it’s packed with new features including support for [natural language queries](https://docs.voxel51.com/user_guide/brain.html#brain-similarity-text) in the FiftyOne App, integrations with [Qdrant](https://docs.voxel51.com/integrations/qdrant.html#qdrant-integration) and [Pinecone](https://docs.voxel51.com/integrations/pinecone.html#pinecone-integration) for native text and image searches on FiftyOne datasets, and much more! Check out the new features in the [announcement blog post](https://voxel51.com/blog/announcing-fiftyone-0-20/?utm_source=blog-cu), the latest [release notes](https://docs.voxel51.com/release-notes.html), and in the [live demo & AMA](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision/) with Voxel51 CTO Brian Moore on April 20, 2023 @ 10AM PT \[1PM ET\]. Oh yea, you can also see it for yourself! It’s easy to [get up and running](https://docs.voxel51.com/index.html) in just a few minutes. ### Community Contribution Shoutouts Shoutout to the following community members who contributed to the latest project release! - [Joy Timmermans](https://github.com/timmermansjoy) contributed [#2716- Adding ability to pass CVAT organization for annotations](https://github.com/voxel51/fiftyone/pull/2716) - [Akshit Priyesh](https://github.com/akshitpriyesh) contributed [#2774 — fix app crash while filtering keypoints](https://github.com/voxel51/fiftyone/pull/2774) - [Kishan Savant](https://github.com/NeoKish) contributed [#2771 — fix broken torchvision dataset links](https://github.com/voxel51/fiftyone/pull/2771) and [#2844 - Updated URLs for CalTech Dataset](https://github.com/voxel51/fiftyone/pull/2841) ### FiftyOne Teams 1.2 Is Also Generally Available! Check out the [FiftyOne Teams 1.2 release blog post](https://voxel51.com/blog/announcing-fiftyone-teams-1-2/?utm_source=blog-cu) to explore what’s new (it’s fully compatible with your existing FiftyOne workflows). ## **FiftyOne Community News** ### 1M+ Downloads FiftyOne adoption continues to accelerate and the project recently crossed 1M+ downloads 🚀🚀🚀. Thank you to everyone who is using FiftyOne in your computer vision workflows. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ### FiftyOne on GitHub GitHub is home to the open source FiftyOne project. Here’s the latest snapshot of what’s happening in the [FiftyOne GitHub repo](https://github.com/voxel51/fiftyone): - **Total stars:** 2,750+ - **Total contributors:** 59 - **Total used by:** 268 repositories - **Total forks:** 328 - **Total issues closed so far**: 777 - **Total commits so far this year:** 792 ### FiftyOne Community Slack The FiftyOne Community [Slack channel](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ/?utm_source=blog-cu) is where you can join more than 1,500 machine learning engineers and data scientists using FiftyOne to improve the quality of their computer vision data and build better models. Ask questions, answer questions, or simply follow along with the discussion! To make it easy to catch the highlights, every Friday we recap interesting questions and answers from Slack in [Tips & Tricks blog series](https://voxel51.com/blog/category/tips-tricks/). Recent posts include: - [FiftyOne Computer Vision Embeddings Tips and Tricks – Mar 31, 2023](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023/?utm_source=blog-cu) - [FiftyOne Computer Vision Tips and Tricks – Mar 24, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-24-2023/?utm_source=blog-cu) - [FiftyOne Tips and Tricks for Accelerating Computer Vision Workflows – Mar 17, 2023](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023/?utm_source=blog-cu) - [FiftyOne Computer Vision Tips and Tricks – Mar 10, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023/?utm_source=blog-cu) - [FiftyOne Tips and Tricks for Customizing your Computer Vision Workflows – Mar 03, 2023](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023/?utm_source=blog-cu) ## Computer Vision Meetups \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Voxel51 sponsors 13 virtual [Computer Vision Meetups](https://www.meetup.com/pro/computer-vision-meetups/) around the world. (To join, visit the Meetup [link](https://www.meetup.com/pro/computer-vision-meetups/) and scroll down to find the location friendliest to your time zone.) The Computer Vision Meetups are geared towards data scientists, machine learning engineers, and open source enthusiasts who want to expand their knowledge of computer vision and complementary technologies. We put an emphasis on open source software, and speakers who are computer vision practitioners or academics doing research in the field. ### Upcoming Meetups: - [April ’23 Computer Vision Meetup](https://voxel51.com/computer-vision-events/), April 13, 2023 – 10AM PT / 17:00 UTC: Talks include: Generating Diverse and Natural 3D Human Motions from Texts – [Chuan Guo](https://www.linkedin.com/in/chuan-guo-59b6a810a/)(University of Alberta); Emergence of Maps in the Memories of Blind Navigation Agents – [Dhruv Batra](https://www.linkedin.com/in/dhruv-batra-dbatra/) (Meta & Assoc. Professor, Georgia Tech); and Using Computer Vision to Understand Biological Vision – [Benjamin Lahner](https://www.linkedin.com/in/benlahner/)(MIT) - [April '23 APAC Computer Vision Meetup,](https://voxel51.com/computer-vision-events/) April 27 – 10AM IST / 04:30 UTC: Talks include: Leveraging Attention for Improved Accuracy and Robustness - Hila Chefer (Tel Aviv University); and AI Deployments at the Edge with OpenVINO – Zhuo Wu (Intel) - [May ’23 Computer Vision Meetup](https://voxel51.com/computer-vision-events/), May 11, 2023 – 10AM PT / 17:00 UTC: Talks include: The Role of Symmetry in Human and Computer Vision – Sven Dickinson (University of Toronto & Samsung); and Machine Learning for Fast, Motion-Robust MRI – Nalini Singh (MIT) ### Recapping the March ‘23 Meetup We recently held the March ‘23 Computer Vision Meetup that showcased these topics and speakers: - Lightning Talk: A Recycling Max Pooling Module for 3D Point Cloud Analysis — [Jiajing Chen](https://www.linkedin.com/in/jiajing-chen-560189193/) (PhD Candidate, Syracuse University) - Lighting up Images in the Deep Learning Era — [Soumik Rakshit](https://www.linkedin.com/in/soumikrakshit/), ML Engineer (Weights & Biases) - Taking Computer Vision Models in Notebooks to Production — [Sumanth P](https://www.linkedin.com/in/sumanth-p-09b339173/) (ML Engineer) [Get the Meetup recap](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023/?utm_source=blog-cu), including video playbacks, executive summaries, slides, and Q&A. ### Even More Community Events In addition to meetups, we invite you to join us for one or more of these upcoming [events](https://voxel51.com/computer-vision-events/?utm_source=blog-cu): - What’s New [in FiftyOne 0.20 & AMA](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision/), April 20, 2023 – 10AM PT / 17:00 UTC: For each new release, we host a live demo of the new features, followed by an open Q&A where you can get answers to any questions you might have. Come seethe new features in FiftyOne 0.20 in action and get your questions answered! - Getting Started with FiftyOne Workshop - April, May, and June dates: We recently launched a free 90 minute, hands-on workshop to help you get up and running with FiftyOne. Building on the success of the [March workshop](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap/?utm_source=blog-cu), we’re excited to announce three new workshop dates! - [April 26 @ 9:00 AM IST \[1:30 PM AEST / 03:30 UTC\]](https://voxel51.com/computer-vision-events/) - [May 31 @ 4 PM BST \[11 AM EDT / 15:00 UTC\]](https://voxel51.com/computer-vision-events/) - [June 28 @ 10 AM PDT \[1 PM EDT / 17:00 UTC\]](https://voxel51.com/computer-vision-events/) ## New Docs, Blogs, Videos, and Tutorials We want everyone to be successful with FiftyOne, and one of the ways we try to enable that is by publishing resources that you might find helpful and handy. Here’s a list of some of the new [documentation](https://docs.voxel51.com/), [blogs](https://voxel51.com/blog/), [videos](https://www.youtube.com/@voxel51/videos), [tutorials](https://docs.voxel51.com/tutorials/index.html), [integrations](https://docs.voxel51.com/integrations/index.html), and [cheat sheets](https://docs.voxel51.com/cheat_sheets/index.html) that you may want to check out. ### Blogs - [Announcing FiftyOne 0.20 with Natural Language Search, Vector Database Integrations, and Point Cloud-Only Datasets](https://voxel51.com/blog/announcing-fiftyone-0-20/?utm_source=blog-cu) - [Exploring Google’s Open Images V7](https://voxel51.com/blog/exploring-google-open-images-v7/?utm_source=blog-cu) - [A Better Way to Visualize 3D Point Clouds and Work with OpenAI’s Point-E](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e/?utm_source=blog-cu) - [A Google Search Experience for Computer Vision Data](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data/?utm_source=blog-cu) - [Announcing the FiftyOne Computer Vision Educational Workshop Series](https://voxel51.com/blog/announcing-the-fiftyone-computer-vision-workshop-series/?utm_source=blog-cu) - [Exploring the Cityscapes Dataset for Semantic Urban Scene Understanding](https://voxel51.com/blog/exploring-the-cityscapes-dataset-for-semantic-urban-scene-understanding/?utm_source=blog-cu) - [How Computer Vision Is Changing Manufacturing in 2023](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing-in-2023/?utm_source=blog-cu) - [Exploring the UCF101 Dataset: A Large-Scale, YouTube-Based Action Recognition Dataset](https://voxel51.com/blog/exploring-ucf101-youtube-based-action-recognition-dataset/?utm_source=blog-cu) - Giving YOLOv8 a Second Look \[ [Part 1](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1/?utm_source=blog-cu) \| [Part 2](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2/?utm_source=blog-cu) \| [Part 3](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3/?utm_source=blog-cu)\] ### Videos - Computer Vision Meetup: [Why Discard if You can Recycle?: A Recycling Max Pooling Module for 3D Point Cloud Analysis](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023/#talk1-video) - Computer Vision Meetup: [Lighting Up Images in the Deep Learning Era](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023/#talk2-video) - Computer Vision Meetup: [Taking Computer Vision Models in Notebooks to Production](https://voxel51.com/blog/recapping-the-computer-vision-meetup-march-2023/#talk3-video) - Webinar Recap: [What’s New in FiftyOne 0.19 for Computer Vision](https://www.youtube.com/watch?v=6hlJO8wW0YU) - Video tutorial: [How to get started with the UCF101 action recognition video dataset](https://youtu.be/xArphgd_hVs) ### Tutorials and Cheat Sheets - [Point-E tutorial showcasing the 3D Visualizer’s capabilities in the context of building a 3D self-driving dataset](https://docs.voxel51.com/tutorials/pointe.html) - [Fine-tune YOLOv8 models for custom use cases with the help of FiftyOne](https://docs.voxel51.com/tutorials/yolov8.html) - [Downloading and evaluating Open Images](https://docs.voxel51.com/tutorials/open_images.html) ### New & Updated Documentation We published a bunch of new docs to help you make the most of your FiftyOne experience: - [Working with point cloud-only datasets](https://docs.voxel51.com/user_guide/using_datasets.html#point-cloud-datasets) - [On-the-fly custom embedded document creation](https://docs.voxel51.com/user_guide/using_datasets.html#custom-embedded-documents) - Viewing [Segmentation](https://docs.voxel51.com/user_guide/using_datasets.html#semantic-segmentation) and [Heatmap](https://docs.voxel51.com/user_guide/using_datasets.html#heatmaps) data stored as images in the cloud in the App - Search a Brain similarity index by [arbitrary natural language queries](https://docs.voxel51.com/user_guide/brain.html#brain-similarity-text) natively in the App - Using [Qdrant integration](https://docs.voxel51.com/integrations/qdrant.html#qdrant-integration) for native text and image searches on FiftyOne datasets - Using [Pinecone integration](https://docs.voxel51.com/integrations/pinecone.html#pinecone-integration) for native text and image searches on FiftyOne datasets If you’re using FiftyOne Teams, there’s new documentation on configuring the [default access level](https://docs.voxel51.com/teams/roles_and_permissions.html#teams-default-access) on newly created datasets. And, as usual, you can find the entire collection in the [FiftyOne](https://docs.voxel51.com/user_guide/index.html) and [FiftyOne Teams](https://docs.voxel51.com/teams/index.html#fiftyone-teams) documentation! ## Industry Spotlight We started a blog series to highlight how different industries – from construction to climate tech, from retail to robotics, and more – are using computer vision, machine learning, and artificial intelligence to drive innovation. This month’s featured industry is manufacturing. Learn more in the blog post, [How Computer Vision Is Changing Manufacturing in 2023](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing-in-2023/), and in the PDF below. [Download PDF: Manufacturing Computer Vision Spotlight](https://voxel51.com/wp-content/uploads/2023/03/02.23_Carosuel_Industry_Manu_AV_7.1-combined_1.pdf) [Download](https://voxel51.com/wp-content/uploads/2023/03/02.23_Carosuel_Industry_Manu_AV_7.1-combined_1.pdf) ## Voxel51’s Commitment to Open Source and Community If you’re new to Voxel51, open source, transparency, and giving back to the computer vision community are what we are all about! Whether it’s developing the open source [FiftyOne computer vision toolset](https://github.com/voxel51/fiftyone) to help engineers and data scientists build high-quality datasets and models, sponsoring [Meetups](https://www.meetup.com/pro/computer-vision-meetups/) to help members boost their computer vision knowledge, or [giving to charitable causes](https://voxel51.com/charitable-giving/) on behalf of the community, Voxel51 is committed to bringing transparency and clarity to the world's data. ## Community-Powered Charitable Contributions During each of our virtual Meetups and other virtual events, we ask attendees to vote for their favorite charity, and then we make a donation to the charity that receives the most votes. Since launching the charitable giving program, on behalf of the computer vision community, FiftyOne donated $1400 USD across these admirable causes: [Children International](https://www.children.org/), [Foundation Fighting Blindness](https://www.fightingblindness.org/), [Turkey-Syria Earthquake Relief](https://www.directrelief.org/emergency/turkey-syria-earthquake/) by Direct Relief, and [World Literacy Foundation](https://worldliteracyfoundation.org/). [Community Update](https://voxel51.com/blog/tag/community-update) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [computer vision events](https://voxel51.com/blog/tag/computer-vision-events) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [open source](https://voxel51.com/blog/tag/open-source) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/338b38d41e6072dd11af86f21f5309337c52f36b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – May ‘23\\ \\ Product & News\\ \\ • \\ \\ May 5, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/7685a2b9b8681c3c641a1118ad1c4685f1af21b7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – July 2023\\ \\ Product & News\\ \\ • \\ \\ Jul 6, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4a03c488b86f54b773a6dbe9d9fe21ca6c7da882-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – August 2023\\ \\ Product & News\\ \\ • \\ \\ Aug 3, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-273-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – April 7, 2023 Apr 7, 2023 • 5 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#8d49fa679e17) [Customizing the visualization of object embeddings](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#80a07d63c7a1) [Using FiftyOne to copy predictions over to ground truth for CVAT labels](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#3fbcfb6bbc2e) [Adding detections to a video dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#fb7d334cbb39) [How to split data into train, validation, and test sets in FiftyOne and export them separately](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#78c4fc425a5b) [Keeping a persistent view of dataset for multiple sessions](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#68b9b6b5808c) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#1b54b9e6fa15) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#8d49fa679e17) [Customizing the visualization of object embeddings](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#80a07d63c7a1) [Using FiftyOne to copy predictions over to ground truth for CVAT labels](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#3fbcfb6bbc2e) [Adding detections to a video dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#fb7d334cbb39) [How to split data into train, validation, and test sets in FiftyOne and export them separately](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#78c4fc425a5b) [Keeping a persistent view of dataset for multiple sessions](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#68b9b6b5808c) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023#1b54b9e6fa15) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## Customizing the visualization of object embeddings Community Slack member Joy asked, _“Could you please provide an example of how to restrict the visualization to only objects in a subset of classes using the compute\_visualization() method's default embeddings model and dimensionality method?”_ Hey Joy! To visualize patch embeddings and color them by their label, you need to choose a `brain_key` that corresponds to patch embeddings. You can use the `patches_field="ground_truth"` argument to embed the patches defined by the ground\_truth field instead of the entire images. The following example shows how to restrict the visualization to only objects in a subset of the classes. To restrict the visualization to only objects in a subset of the classes, each point in the scatter plot in the [Embeddings panel](https://docs.voxel51.com/user_guide/app.html#app-embeddings-panel) corresponds to an object, colored by its label class. When points are lassoed in the plot, the corresponding object patches are automatically selected in the [Samples panel](https://docs.voxel51.com/user_guide/app.html#app-samples-panel). ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4from fiftyone import ViewField as F 5 6dataset = foz.load_zoo_dataset("quickstart") 7 8# Generate visualization for `ground_truth` objects 9results = fob.compute_visualization( 10 dataset, patches_field="ground_truth", brain_key="vj_viz" 11) 12 13# Restrict to the 10 most common classes 14counts = dataset.count_values("ground_truth.detections.label") 15classes = sorted(counts, key=counts.get, reverse=True)[:10] 16view = dataset.filter_labels("ground_truth", F("label").is_in(classes)) 17 18session = fo.launch_app(view) ``` Once the session is launched, you can use the embeddings panel and filter down the ground truth labels by applying the brain key as demonstrated here: Learn more about [_object embeddings_](https://docs.voxel51.com/user_guide/brain.html?highlight=compute_visualization#object-embeddings-example) in the FiftyOne docs. ## Using FiftyOne to copy predictions over to ground truth for CVAT labels Community Slack member Kais asked, _“In my dataset, I have two fields, ground\_truth and predictions. For a number of samples, I want to copy the content of 'predictions' over to 'ground\_truth', so that the labels that I create in CVAT get saved under 'ground\_truth'. How can I approach it in FiftyOne?”_ Hey Kais, thanks for your question! It sounds like you want to copy the content of `predictions` over to `ground_truth` for some samples in your dataset, so that the labels you create in CVAT get saved under `ground_truth`. [There is a tight integration](https://voxel51.com/docs/fiftyone/integrations/cvat.html) between FiftyOne and CVAT that is designed to manage the full annotation workflow, from task creation to annotation import. However, if you have created CVAT tasks outside of FiftyOne, you can use the [import\_annotations()](https://docs.voxel51.com/api/fiftyone.utils.cvat.html#fiftyone.utils.cvat.import_annotations) utility to import individual task(s) or an entire project into a FiftyOne dataset. First, you need to create a pre-existing CVAT project outside of FiftyOne. Once you have that, you can use 'annotate()\` to add the project to your dataset. Then, you can use `import_annotations()` to import the annotations and specify the name of the field into which to load your labels: ```python 1fouc.import_annotations(dataset, project_name=project_name, label_types={"detections": "ground_truth"}) ``` Make sure to specify the correct `project_name` and `label_types`. You can also download both the annotations and the media from CVAT by setting `download_media=True`. Once you've imported the annotations, you can launch the FiftyOne App with `fo.launch_app(dataset)` to view and edit your labeled samples. The code below demonstrates the functionality discussed in the response. ```python 1import os 2import fiftyone as fo 3import fiftyone.utils.cvat as fouc 4import fiftyone.zoo as foz 5 6# Load a FiftyOne dataset with some samples 7dataset = foz.load_zoo_dataset("quickstart", max_samples=3).clone() 8 9# Create a pre-existing CVAT project with some annotations 10project_name = "my_cvat_project" # Replace with your project name 11cvat_annotations_path = "/path/to/cvat_annotations.xml" # Replace with the path to your CVAT annotations file 12# Add the project to your dataset using annotate() 13results = dataset.annotate(project_name, label_field="ground_truth", annotation_backend="cvat", annotation_filepath=cvat_annotations_path) 14 15# Import the annotations into your FiftyOne dataset using import_annotations() 16fouc.import_annotations(dataset, project_name=project_name, label_types={"detections": "ground_truth"}, data_path="/tmp/cvat_import", download_media=True) 17 18# Copy the content of 'predictions' over to 'ground_truth' for some samples 19for sample in dataset.take(2): # Replace with the samples you want to modify 20 sample.ground_truth = sample.predictions 21 sample.save() 22 23# Launch the FiftyOne app to view and edit your labeled samples 24session = fo.launch_app(dataset) ``` Learn more about [CVAT integration](https://docs.voxel51.com/integrations/cvat.html) in the FiftyOne Docs. ## Adding detections to a video dataset Community member Kevin asked, _“I have a video dataset and I am using a Python loop to perform the necessary edits. I would like to add detections to each sample (video). What is the right approach to add detections to video samples in FiftyOne?_" Hey Kevin! In FiftyOne, adding detections to video samples is done by using the [frames](https://docs.voxel51.com/api/fiftyone.core.frame.html#fiftyone.core.frame.Frame) attribute of each video sample. Video samples are recognized by their MIME type and have media type video in FiftyOne datasets. The frames attribute is a dictionary whose keys are frame numbers and values are instances of the frame class. Frame instances can hold all `label` instances and other primitive-type fields for the frame. To add, modify, or delete labels of any type as well as primitive fields such as integers, strings, and booleans, you can use the same dynamic attribute syntax that you use to interact with samples. Here's an example code snippet that demonstrates how to add a detection to a frame in a video sample: ```python 1import fiftyone as fo 2# Load the dataset containing video samples 3dataset = fo.load_dataset('path/to/dataset', 'video') 4 5# Iterate over the video samples 6for sample in dataset: 7 # Get the frame number and add a detection to it 8 frame_number = 10 9 detection = fo.Detection(label='person', bounding_box=[0, 0, 100, 100]) 10 sample.frames[frame_number].detections.append(detection) 11 12 # Save the changes to the database 13 sample.save() ``` In this example, we load the dataset containing video samples and iterate over each sample. For each sample, we add a detection to the frame with frame number 10. The detection consists of a label and a bounding box. Finally, we save the changes to the database by calling \`sample.save()\`. Learn more about [adding object detection to datasets](https://docs.voxel51.com/recipes/adding_detections.html?highlight=detecting) in the FiftyOne docs. ## How to split data into train, validation, and test sets in FiftyOne and export them separately Community member Eli asked, _“What is the process for splitting a dataset into train, validation, and test sets in FiftyOne, and how can each split be exported separately to its own directory?”_ Hey Eli, to split your dataset into train, validation, and test sets, you can use the `dataset.split()` method in FiftyOne. You can specify the name and percentage of the samples to include in each split, and then export them separately as shown below: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3import fiftyone.utils.random as four 4import fiftyone.utils.yolo as fouy 5 6dataset = foz.load_zoo_dataset("quickstart") 7 8four.random_split(dataset, {"val": 0.6, "test": 0.4}) 9val_view = dataset.match_tags("val") 10test_view = dataset.match_tags("test") 11 12val_view.export( 13 export_dir="val_export", 14 dataset_type=fo.types.YOLOv5Dataset, 15 classes=classes 16) 17 18test_view.export( 19 export_dir="test_export", 20 dataset_type=fo.types.YOLOv5Dataset, 21 classes=classes 22) ``` The following image shows the directory structure after running the code. And here’s the metadata information of the output. Learn more about [various export methods](https://docs.voxel51.com/user_guide/export_datasets.html) in the FiftyOne docs. ## Keeping a persistent view of dataset for multiple sessions Community member Dan asked, _“I used dataset.clone() to create a clone of a dataset. But after closing my session debugger and re-running it, the dataset was no longer present. What could be the reason for this?”_ Hey Dan! It looks like when you cloned the dataset in FiftyOne you did not make it persistent. This means it was not saved to the database and was deleted once the session was closed. By default, datasets are non-persistent. Non-persistent datasets are deleted from the database each time the database is shut down. Note that FiftyOne does not store the raw data in datasets directly (only the labels), so your source files on disk are untouched. You can make a cloned dataset persistent by setting its [persistent](https://docs.voxel51.com/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.persistent) property to `True` before saving it to the database. To do this, you can add the following code after cloning the dataset: ```python 1cloned_dataset.persistent = True 2cloned_dataset.save() ``` Finally, you can check to see what [datasets](https://docs.voxel51.com/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset) exist at any time via [list\_datasets()](https://docs.voxel51.com/api/fiftyone.core.dataset.html#fiftyone.core.dataset.list_datasets). ```python 1print(fo.list_datasets()) ``` To check whether a given dataset is persistent, load it with `dataset = fo.load_dataset(“my_dataset”)`, and then print out `dataset.persistent`. Learn more about [data persistence](https://docs.voxel51.com/user_guide/using_datasets.html) and [datasets](https://docs.voxel51.com/user_guide/using_datasets.html) in the FiftyOne Docs. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,500+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,750+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,500+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 265+ repositories - 55+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [CVAT](https://voxel51.com/blog/tag/cvat) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [integrations](https://voxel51.com/blog/tag/integrations) [Video Data Analysis](https://voxel51.com/blog/tag/video-data-analysis) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/4ac1a727dc192a21563cde51b6e345f620e09376-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/79d00d175a8098516cb2f4a7711131cbf322d01a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Finding and Correcting Mistakes – FiftyOne Tips and Tricks – Aug 18, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Aug 18, 2023](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/03107d477b7db4be03031293fa4fe15aaea806f0-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Sept 16, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 17, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-274-lllmstxt|> ## Controllable Diffusion Models [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) Towards Controllable Diffusion Models with GLIGEN Apr 12, 2023 • 7 min read Article content In this article [Background and Motivation](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen#939f0b6138b3) [Approach](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen#ab910cb93bf1) [GLIGEN in Action](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen#5783fded1879) [The Role of Data](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen#884b0b01f2c9) [Conclusion & Next Steps](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen#355ce37a3c80) In this article [Background and Motivation](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen#939f0b6138b3) [Approach](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen#ab910cb93bf1) [GLIGEN in Action](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen#5783fded1879) [The Role of Data](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen#884b0b01f2c9) [Conclusion & Next Steps](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen#355ce37a3c80) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) _Editor’s note: this is a guest post by [Yuheng Li](https://yuheng-li.github.io/), computer science Ph.D. student at University of Wisconsin-Madison_ ![](https://cdn.sanity.io/images/h6toihm1/production/b05bb8278dd99efec5be3d643b06dcf38f955900-1920x789.png?auto=format&dpr=2&fit=max&q=75&w=1600) AI-generated imagery is an extremely exciting area of computer vision, with a variety of impressive innovations and technologies available to help bring new images to life. But what if you want more control over the generation process? That’s the question that inspired recent research and ultimately led to the creation of GLIGEN (Grounded Language to Image Generation). In this blog post, I’ll share how GLIGEN came to be, how it works, and show you how you can gain control over the outputs of diffusion models by adding new trainable parameters! ## Background and Motivation Image generation has seen remarkable progress in recent years, with diffusion models being crucial to the field for [AI image generation](https://www.midjourney.com/showcase/recent/), [3D modeling](https://make-it-3d.github.io/), and more. Large-scale text-to-image (T2I) models like [DALL-E2](https://openai.com/product/dall-e-2) and [Stable Diffusion](https://stablediffusionweb.com/) can create complex images from text, but can only be conditioned on text input, not on input from other modalities. This can present challenges for users who want more control over the generation process. If you're an interior designer looking to visualize how furniture placement will look in a living room, for example, existing T2I diffusion models likely won't bring your imagination to life. Similarly, if you're hoping to create an image of yourself in the same pose as Michael Jackson, these models won't do the trick easily for you. The community has been trying to work on bringing more control over the diffusion generation process. These efforts can be broadly categorized into four groups based on the type of trainable parameters. #### 1\. Train a new model from scratch ![](https://cdn.sanity.io/images/h6toihm1/production/4ab47dbb88952d38787eccacfdbc7b618ff94688-1828x728.png?auto=format&dpr=2&fit=max&q=75&w=1600) One example is [Composer](https://damo-vilab.github.io/composer-page/), which defines representation elements of an image: caption, sketch, color, etc., and trains a new model from scratch conditioned on these elements. The advantage of this direction is: one can design its own model architecture to be better compatible with controllable elements. However, since it does not utilize existing foundation image generation models, training is costly each time. #### 2\. Fine-tune a pre-trained model ![](https://cdn.sanity.io/images/h6toihm1/production/b3dabd9a95d496265885e38aa2d9b003ed24e19b-1828x728.png?auto=format&dpr=2&fit=max&q=75&w=1600) Another paradigm is to fine-tune weights of an existing model. For example, [ReCo](https://arxiv.org/abs/2211.15518) appends new bounding box information into the caption and fine-tunes the pre-trained text encoder as well as diffusion models. #### 3\. Add new trainable parameters to a frozen pre-trained model ![](https://cdn.sanity.io/images/h6toihm1/production/2e250b075c41e5b36f47d5137bcdad2a959acb52-1999x728.png?auto=format&dpr=2&fit=max&q=75&w=1600) [GLIGEN](https://gligen.github.io/) and [ControlNet](https://github.com/lllyasviel/ControlNet) fall into this category. Instead of changing weights of the foundation models, they add new learnable parameters to adapt and modify intermediate features in existing models. #### 4\. Change sampling direction for a pre-trained model ![](https://cdn.sanity.io/images/h6toihm1/production/47a4c7b8ecfa5fe37cd849df1940552abf938ad9-1999x728.png?auto=format&dpr=2&fit=max&q=75&w=1600) What if I don’t want to train any parameters at all? Can I still control diffusion models? The answer is Yes! [Universal Guidance](https://arxiv.org/abs/2302.07121) proposes to use pre-trained discriminative models such as object detectors to control the sampling process. ## Approach #### **Diffusion Mo** del Before diving into the GLIGEN approach, let’s first get familiar with a pre-trained diffusion model architecture. Typically, a diffusion model is a Unet architecture as shown in the figure below (left), consisting of stacks of residual blocks and Transformer blocks (details on the right) where the real magic happens: the self-attention layer makes global visual feature processing possible; the cross-attention layer absorbs in caption features. ![](https://cdn.sanity.io/images/h6toihm1/production/015c11f500b042bf13df9327fbe8bd4a31c57a54-1546x836.png?auto=format&dpr=2&fit=max&q=75&w=1546) This is great if we want to condition our generation process on text alone, but what if we want more control over our generation process? Training models from scratch, conditioned on new control inputs can be quite costly, and it is unfeasible to do so whenever users ask for a new conditional input! #### GLIGEN GLIGEN bypasses this problem by **adding new control to a pre-trained model without changing its weights**. The core idea of the GLIGEN is: **modifying the original visual features with Gated Self-Attention in the Transformer blocks.** The figure below (left) shows where the gated self-attention layer is inserted in GLIGEN. **Input to Gated Self-Attention** - This layer takes in image visual features and extra conditional features called grounding tokens. - Grounding tokens represent the new conditional input users want. For example, if a user wants to control the generation process with bounding boxes then the grounding tokens contain information for bounding boxes. **Operation within Gated Self-Attention** - As shown on the right in the figure below, the visual tokens and grounding tokens are concatenated along the sequence dimension and fed into a self-attention layer. For the output, we discard the grounding tokens and treat the remaining ones as residual (light purple). - Instead of directly adding the residual, we first multiply the residual with a learnable constant γ which is initialized as 0. This γ acts like a gate, and that’s where the name for this layer comes from. This means that at the beginning of training, the new gated self-attention layer will not affect the original feature, leading to more stable training. ![](https://cdn.sanity.io/images/h6toihm1/production/6df1ddb8cb3d04f926dab306b4bebdffb8a4b66d-1999x1074.png?auto=format&dpr=2&fit=max&q=75&w=1600) **Optional Input to the Unet** The above design conceptually works for _any_ extra conditional input due to the generalizability of the Transformer. We also find that for conditions that are spatially aligned with the output image such as an edge map, depth map, etc., the training converges faster and more easily, if they are also given as input to the Unet as shown in the figure below. In this case, the first convolutional (“conv”) layer needs to be modified to take in extra channels and needs to be trainable. ![](https://cdn.sanity.io/images/h6toihm1/production/9b74c6c9b3a52314a73de5b2ea85707f136be9a6-1999x434.png?auto=format&dpr=2&fit=max&q=75&w=1600) **Scheduled Sampling** As we only modify intermediate features of pre-trained diffusion models, we can freely remove gated self-attention layers during the generation sampling process as shown below. Since the initial steps typically regulate the basic structure, this approach can effectively achieve a good trade-off between adhering to the conditions and ensuring image quality. ![](https://cdn.sanity.io/images/h6toihm1/production/7d8244718b6cc61721372c43ec210b0c8ed5a937-1958x1080.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## GLIGEN in Action Here we demonstrate GLIGEN results across three different modality use cases: generating new images given grounding input on bounding boxes, keypoints, and Canny maps. #### Grounding on Bounding Boxes ![](https://cdn.sanity.io/images/h6toihm1/production/12b82821ff26d7ab09c511e211a15083b0511c06-1999x500.png?auto=format&dpr=2&fit=max&q=75&w=1600) In this modality, users can specify the location of objects in their caption prompt. For example, you can control the position of your favorite celebrities to create a poster. On top of that, you can also control the style of the generated image by providing a reference image you’d like to use as inspiration in the new creation. #### Grounding on Keypoints ![](https://cdn.sanity.io/images/h6toihm1/production/a2854d510d3e67abfdbd73abe19e9ccf36facd6c-1999x500.png?auto=format&dpr=2&fit=max&q=75&w=1600) One can also control the pose of the generated object by providing a set of keypoints. Note that in this case the GLIGEN is only trained using human keypoint data, but the model can be generalized into other domains such as monkeys and cartoon characters due to the scheduled sampling technique. #### Grounding on Canny Maps ![](https://cdn.sanity.io/images/h6toihm1/production/b0b4a4f102275cf54a8fa69a574e43054c76f6ee-1999x500.png?auto=format&dpr=2&fit=max&q=75&w=1600) GLIGEN also enables users to easily generate various colorized versions of a canny drawing picture, allowing designers to quickly fill in colors and experiment with different styles. This feature makes GLIGEN a valuable tool for artists and designers who seek to explore multiple design options efficiently. You can refer to our [paper](https://arxiv.org/abs/2301.07093) for more details and quantitative analysis. ## The Role of Data #### **Data Used to Train GLIGEN** GLIGEN was trained mostly on data with bounding box grounding. Ideal data for this task would consist of image-text grounding pairs (see below). However, this type of data is rare (thousands). To overcome this shortage, the training data was augmented with a combination of three different data types. ![](https://cdn.sanity.io/images/h6toihm1/production/d8d87986f503e793592eb9f74fdbb9a36766a1ab-1999x787.png?auto=format&dpr=2&fit=max&q=75&w=1600) - **Grounding data** - Each image is associated with a caption describing the whole image; noun entities are extracted from the caption, and are labeled with bounding boxes. - Since the noun entities are taken directly from the natural language caption, they can cover a much richer vocabulary which will be beneficial for open-world vocabulary grounded generation. - **Detection data** - Noun entities are pre-defined closed-set categories (e.g., 80 object classes in COCO). In this case, we choose to use either a null caption or concatenate class name as a caption. - The detection data is of larger quantity (millions) than the grounding data (thousands), and can therefore greatly increase overall training data. - **Detection and caption data** - Noun entities are the same as those in the detection data, and the image is described separately with a text caption. - In this case, the noun entities may not exactly match those in the caption. For example, in the above example, the caption only gives a high-level description of the living room, whereas the detection annotation provides more fine-grained object-level details. #### GLIGEN for Other Tasks **Object Detection** Object detection is one of the most important perception tasks in vision, and often requires increasing amounts of labeled data for training. Automatic ways of generating training data have been explored in recent years. One common approach is to get pseudo-label from a pre-trained detector like [GLIP](https://github.com/microsoft/GLIP). GLIGEN potentially opens another way by generating an infinite number of training data. **Causal Inference** GLIGEN also demonstrates that it can generate counterfactual results. For example, an apple is the same size as a dog; a hen is smaller than an egg as shown below. These data can potentially help to train reasoning models to better understand spatial relationships between objects. ![](https://cdn.sanity.io/images/h6toihm1/production/350fa34b600d91a52d276b1d1c61e7180590e328-1999x449.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Conclusion & Next Steps Although T2I diffusion models have their limitations when it comes to controllable generation, the community has been actively working on addressing these concerns. One of the proposed solutions, GLIGEN, modifies the features of diffusion models, without altering their weights, which makes it a cost-effective and modulated training approach. The outcomes of GLIGEN highlight its proficiency in multiple grounding modalities, indicating that it has the potential to improve controllable diffusion models for generating images. Here are some resources to learn more about GLIGEN and give it a try: - [The GLIGEN paper](https://arxiv.org/abs/2301.07093) - [GLIGEN on GitHub](https://github.com/gligen/GLIGEN) - [GLIGEN on HuggingFace](https://huggingface.co/gligen) And if you really want to explore further, you may want to consider using GLIGEN to build your own high quality dataset, and then visualize it in [FiftyOne](https://github.com/voxel51/fiftyone), the open source computer vision toolset maintained by Voxel51. [Diffusion models](https://voxel51.com/blog/tag/diffusion-models) [GLIGEN](https://voxel51.com/blog/tag/gligen) [Grounded Language to Image Generation](https://voxel51.com/blog/tag/grounded-language-to-image-generation) [text-to-image](https://voxel51.com/blog/tag/text-to-image) [Text-to-Image Diffusion Models](https://voxel51.com/blog/tag/text-to-image-diffusion-models) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/ff87d65e5b4e5ef50c5732e905f3f16aff8b0a4e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 Survival Guide\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 25, 2023](https://voxel51.com/blog/cvpr-2023-survival-guide) [![](https://cdn.sanity.io/images/h6toihm1/production/0aa3f8dad8ae1464d05d81ac4a301bd92aea55e3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2024 Survival Guide: Five Vision-Language Papers You Don’t Want to Miss\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Apr 15, 2024](https://voxel51.com/blog/cvpr-2024-survival-guide-five-vision-language-papers-you-dont-want-to-miss) [![](https://cdn.sanity.io/images/h6toihm1/production/964c6084b2c194bb2816f2888b6d2e2d9a2ba447-1200x675.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2024 Datasets and Benchmarks – Part 1: Datasets\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Apr 23, 2024](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-275-lllmstxt|> ## Computer Vision Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup — April 13, 2023 Apr 14, 2023 • 5 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#7d338843daad) [Using Computer Vision to Understand Biological Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#d828d315fcdb) [Emergence of Maps in the Memories of Blind Navigation Agents](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#1213c7d10139) [Generating Diverse and Natural 3D Human Motions from Texts](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#5654fd297bc5) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#d0ddcbc8b07c) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#3d9643d89288) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#dc7184442da7) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#7d338843daad) [Using Computer Vision to Understand Biological Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#d828d315fcdb) [Emergence of Maps in the Memories of Blind Navigation Agents](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#1213c7d10139) [Generating Diverse and Natural 3D Human Motions from Texts](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#5654fd297bc5) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#d0ddcbc8b07c) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#3d9643d89288) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023#dc7184442da7) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Yesterday Voxel51 hosted the April 13, 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/). In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes this month was [Wildlife AI](https://www.wildlife.ai/). We were first introduced to Wildlife AI through the FiftyOne community! They are using FiftyOne to enable their users to easily analyze the camera data and create their own models. We are sending this month’s charitable donation of $200 to Wildlife AI on behalf of the computer vision community. ![](https://cdn.sanity.io/images/h6toihm1/production/390540dade0a0e2a6ad0411400d158eb6463bb6c-1024x196.png?auto=format&dpr=2&fit=max&q=75&w=1024) Missed the Meetup? No problem. Here are the playbacks and talk abstracts from the event. ## **Using Computer Vision to Understand Biological Vision** https://www.youtube.com/watch?v=p6gBO2HsIcI In the past decade, deep neural networks (DNNs) have become a leading choice among neuroscientists to model the visual brain. While DNNs are often celebrated for their biological inspiration, they are also criticized for their lack of interpretability. In this talk we will discuss how DNNs help us understand biological intelligence, their promises and pitfalls as models of the brain, and what may be in store for the future. [Benjamin Lahner](https://www.linkedin.com/in/benlahner/) is a PhD student at MIT studying computational neuroscience. ## Emergence of Maps in the Memories of Blind Navigation Agents https://www.youtube.com/watch?v=\_p3p-HxsIyI Animal navigation research posits that organisms build and maintain internal spatial representations, or maps, of their environment. We ask if machines — specifically, artificial intelligence (AI) navigation agents — also build implicit (or ‘mental’) maps. A positive answer to this question would (a) explain the surprising phenomenon in recent literature of ostensibly map-free neural-networks achieving strong performance, and (b) strengthen the evidence of mapping as a fundamental mechanism for navigation by intelligent embodied agents, whether they be biological or artificial. Learn more on [Arxiv.](https://arxiv.org/abs/2301.13261) [Dhruv Batra](https://www.linkedin.com/in/dhruv-batra-dbatra/) is an Associate Professor in the School of Interactive Computing at Georgia Tech (leading the machine Learning & perception lab) and a Research Director in the Fundamental AI Research (FAIR) team at Meta (leading the embodied AI and robotics efforts at FAIR.) Q&A from the talk included: - How does research perform in dynamic environments? - What were the common distance metrics that proved to be most useful? - Have you considered where language may interact with the memory and mapping as part of these experiments? - What is next for blind agents? - Where may blind agents be guaranteed to fail? - Any advice for people who work with non-traditional robots like those that require micro-nano sized sensors? You can jump straight to the Q&A [here.](https://youtu.be/_p3p-HxsIyI?t=1474) ## Generating Diverse and Natural 3D Human Motions from Texts https://www.youtube.com/watch?v=9toMPbHw8uE Automated generation of 3D human motions from text is an interesting yet challenging problem, which also owns a broad range of applications such as VR/AR, 3D content creation. Specifically, the generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem with a two-stage approach: text2length sampling and text2motion generation. Text2length involves sampling from the learned distribution function of motion lengths conditioned on the input text. This is followed by our text2motion module using a temporal variational autoencoder to synthesize a diverse set of human motions of the sampled lengths. Moreover, a large-scale dataset of scripted 3D Human motions, HumanML3D, is constructed, consisting of 14,616 motion clips and 44,970 text descriptions. You can get the data, code, paper and watch the demo video on the [Text-to-Motion](https://ericguo5513.github.io/text-to-motion/) website. [Chuan Guo](https://www.linkedin.com/in/chuan-guo-59b6a810a/) is a fourth-year ECE PhD student at the University of Alberta. Q&A from the talk included: - What was the loss error that was used? - Does it make any difference when using a transformer based language model? Such as BERT? You can jump straight to the Q&A [here.](https://youtu.be/9toMPbHw8uE?t=1388) ## **Computer Vision Meetup Locations** \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Computer Vision Meetup membership has grown to more than [3,700 members](https://www.meetup.com/pro/computer-vision-meetups/) in just under 9 months! The goal of the Meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. Join one of the 13 Meetup locations closest to your timezone. - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) ## **What’s Next?** We have exciting speakers already signed up over the next few months! Become a member of the [Computer Vision Meetup closest to you](https://www.meetup.com/pro/computer-vision-meetups/), then register for the Zoom. Up next on April 27 at 10 AM IST (04:30 UTC) we have the APAC friendly Computer Vision Meetup happening with talks including: - **Leveraging Attention for Improved Accuracy and Robustness** - Hila Chefer (PhD student and lecturer at Tel-Aviv University) - **Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO** - Zhuo Wu (AI software evangelist at Intel focusing on the OpenVINO toolkit) You can find a complete schedule of upcoming Meetups on [the Voxel51 Events page](https://voxel51.com/computer-vision-events/). ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ed751b5c0fbc3d2cb74f0b30e6418d3319564c-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — December 2022\\ \\ Event Recaps\\ \\ • \\ \\ Dec 13, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/bbb1d9add0b0b9aa12682acac795df7c2ba760a9-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — November 2022\\ \\ Event Recaps\\ \\ • \\ \\ Nov 16, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-november-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/09d19530030fd3df3cb247fdc77b8354b768f397-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup – January 2023\\ \\ Event Recaps\\ \\ • \\ \\ Jan 18, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-january-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-276-lllmstxt|> ## Kaggle Image Matching Challenge [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) Exploring Google Research’s Kaggle Image Matching Challenge 2023 Dataset Apr 20, 2023 • 5 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#a91e54a0876b) [About the competition](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#c497cd823f46) [About the dataset](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#d17571c1789f) [Dataset quick facts](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#2fee37c04686) [Step 1: Download the dataset](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#8edab8ffb491) [Step 2: Install FiftyOne](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#e1879a6c4ea4) [Step 3: Import the dataset](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#55333dbcbc32) [Step 4: Optional, but awesome! Add Embeddings](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#63af99b5acb3) [Step 5: Launch the App to visualize the dataset](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#5e0ed3d5ef3f) [Sample details](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#60121acab02d) [Filtering by location](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#cf700a68a9d2) [Filtering by type](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#c16801ef5a94) [Filtering by id or filepath](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#dd3e26c4c5c0) [Embeddings](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#ea58b2ca9c41) [Start working with the dataset](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#1c9ec3756688) [Sponsors and acknowledgements](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#bc0f57425f1d) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#a91e54a0876b) [About the competition](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#c497cd823f46) [About the dataset](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#d17571c1789f) [Dataset quick facts](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#2fee37c04686) [Step 1: Download the dataset](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#8edab8ffb491) [Step 2: Install FiftyOne](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#e1879a6c4ea4) [Step 3: Import the dataset](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#55333dbcbc32) [Step 4: Optional, but awesome! Add Embeddings](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#63af99b5acb3) [Step 5: Launch the App to visualize the dataset](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#5e0ed3d5ef3f) [Sample details](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#60121acab02d) [Filtering by location](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#cf700a68a9d2) [Filtering by type](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#c16801ef5a94) [Filtering by id or filepath](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#dd3e26c4c5c0) [Embeddings](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#ea58b2ca9c41) [Start working with the dataset](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#1c9ec3756688) [Sponsors and acknowledgements](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset#bc0f57425f1d) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the latest installment of our ongoing blog series where we explore computer vision related datasets, this one is from a new Kaggle competition! In this post we’ll use the open source FiftyOne computer vision toolset to explore Google Research’s [Image Matching Challenge 2023 - Reconstruct 3D Scenes from 2D Images](https://www.kaggle.com/competitions/image-matching-challenge-2023/overview) competition. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. ![](https://cdn.sanity.io/images/h6toihm1/production/136887bcc1e07d86237d0f79c0f2bf731805f6fb-1280x720.gif?auto=format&dpr=2&fit=max&q=75&w=1280) ## **About the competition** In this challenge, competitors will need to reconstruct 3D scenes from many different views and help uncover the best methods for building accurate 3D models that can be applied to photography, cultural heritage preservation, and more! It doesn’t hurt that there is $50,000 in prizes up for grabs as well! ![](https://cdn.sanity.io/images/h6toihm1/production/98b7f77f9f337bc53b65e951a099848477d8ba88-1624x662.png?auto=format&dpr=2&fit=max&q=75&w=1600) Let’s say we want to create a 3D scene of a famous landmark from images that tourists and professional photographers have taken from various angles and shared on the internet? What if we were able to combine all the photos to create a more complete, three-dimensional view of the landmark? To accomplish this, we’ll need to use [Structure from Motion](https://en.wikipedia.org/wiki/Structure_from_motion) (SfM), which is the process of reconstructing the 3D model of an environment from a collection of images. SfM normally deals with two types of data: - **Uniform, high-quality data:** This type of data is typically captured by trained operators or with additional sensor data, for example as the cars used by Google Maps. This results in homogeneous, high-quality data. - **Uneven, varying quality data:** This will include assorted images taken using cameras of varying quality, with a wide variety of viewpoints, along with lighting, weather, and other variables. ## About the dataset According to the competition organizers, the competition uses a hidden test. When your submitted notebook is scored, the actual test data (including a sample submission) will be made available to your notebook. You should expect to find roughly 1,100 images in the hidden test set. The number of images in a scene may vary from <10 to ~250. ## Dataset quick facts - **Download Dataset:** [Download from Kaggle](https://www.kaggle.com/competitions/image-matching-challenge-2023/data) - **License:** Varies. Check the LICENSE.txt files in the image directories for details - **Dataset Size:** 12.64 GB - **FiftyOne Dataset Name:** `image-matching-challenge-2023` Up next, let’s download the dataset, install FiftyOne and import the dataset into the App so we can visualize it! ## Step 1: Download the dataset In order to load the “Image Matching Challenge 2023” dataset into FiftyOne, you’ll need to first [download the source data from Kaggle](https://www.kaggle.com/competitions/image-matching-challenge-2023/data). (Don’t forget to accept the terms of the competition first!). After unzipping the download, your directory should look like this: ![](https://cdn.sanity.io/images/h6toihm1/production/a646968884d4c09d3e6ccb216534b97a998e7939-530x608.png?auto=format&dpr=2&fit=max&q=75&w=530) ## Step 2: Install FiftyOne ```python 1pip install fiftyone umap-learn ``` ![](https://cdn.sanity.io/images/h6toihm1/production/d307a75b42dd058cdc439a2d35149d0cdadca528-1280x720.gif?auto=format&dpr=2&fit=max&q=75&w=1280) If you don’t already have FiftyOne installed on your laptop, it takes less than a minute! For example on macOS: - [Verify your version](https://docs.voxel51.com/getting_started/virtualenv.html#creating-a-virtual-environment-using-venv) of Python - Create and activate a [virtual environment](https://docs.voxel51.com/getting_started/virtualenv.html#creating-a-virtual-environment-using-venv) - [Install IPython](https://docs.voxel51.com/getting_started/troubleshooting.html#ipython-installation) (optional) - [Upgrade](https://docs.voxel51.com/getting_started/virtualenv.html#creating-a-virtual-environment-using-venv) your setuptools - [Install FiftyOne](https://docs.voxel51.com/getting_started/install.html#installing-fiftyone) Learn more about how to [get up and running with FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html) in the Docs. ## Step 3: Import the dataset Now that you have the dataset downloaded and FiftyOne installed, let’s import the dataset and launch the FiftyOne App. ```python 1import glob 2import os 3import fiftyone as fo 4 5# Download and unzip `image-matching-challenge-2023.zip` and put path here 6dataset_dir = "/path/to/image-matching-challenge-2023" 7 8samples = [] 9for filepath in glob.glob(os.path.join(dataset_dir, "**"), recursive=True): 10 if filepath.endswith((".jpg", ".jpeg", ".png", ".JPG")): 11 folders = filepath[len(dataset_dir) + 1:].split("/")[:-2] 12 sample = fo.Sample( 13 filepath=filepath, 14 tags=[folders[0]], 15 type=folders[1], 16 location=folders[2], 17 ) 18 samples.append(sample) 19 20dataset = fo.Dataset( 21 "image-matching-challenge-2023", 22 persistent=True 23) 24dataset.add_samples(samples) 25dataset.compute_metadata() ``` ## Step 4: Optional, but awesome! Add Embeddings Visualizing datasets in a low-dimensional embedding space is a powerful workflow that can reveal patterns and clusters in your data that can answer important questions about the critical failure modes of your model and how to augment your dataset to address these failures. Using the [FiftyOne Brain’s embeddings visualization capability](https://docs.voxel51.com/tutorials/image_embeddings.html) can help uncover hidden patterns in the data, enabling us to take the required actions to improve the quality of the dataset and associated models. If you’d like to take advantage of embeddings with this dataset, just bolt the following snippet onto the previous Step 3 snippet. ```python 1import fiftyone.brain as fob 2 3fob.compute_visualization( 4 dataset, 5 model="clip-vit-base32-torch", 6 brain_key="img_viz", 7) ``` **Tips and Tricks:** If you opt to include the embeddings snippet above, depending on your hardware setup, it will take a few minutes to calculate them. So, this would be a great time to check out all the cool [upcoming computer vision events](https://voxel51.com/computer-vision-events/) sponsored by Voxel51! Also, if you are running Python 3.11, you might see an error related to `umap-learn` as one of its dependencies, `numba`, is not yet supported. If you encounter this, simply run a supported environment like 3.9. For example: ```python 1conda create -env39 python=3.9 2conda activate env39 3 4pip install fiftyone umap-learn ``` ## Step 5: Launch the App to visualize the dataset ```python 1session = fo.launch_app(dataset) ``` The code snippet above will launch the FiftyOne App in your default browser. You should see the following initial view of the `image-matching-challenge-2023` dataset by default in the App: ![](https://cdn.sanity.io/images/h6toihm1/production/92c0460bf9fd2d1ac59bc69037e40d8504581dad-1999x1059.png?auto=format&dpr=2&fit=max&q=75&w=1600) Ok, let’s do a quick exploration of the dataset! ## Sample details Click on any of the samples to get additional detail like tags, metadata, labels, and primitives. ![](https://cdn.sanity.io/images/h6toihm1/production/681003d68d2a4dd4c80c56e7bae53bdcae5e5002-1358x857.png?auto=format&dpr=2&fit=max&q=75&w=1358) ## Filtering by location The dataset includes 23 distinct locations including Notre Dame, Taj Mahal, Buckingham Palace, the Lincoln Memorial and more. For example, let’s filter the samples by those labeled pantheon\_exterior. ![](https://cdn.sanity.io/images/h6toihm1/production/5d589d4742bee8e4d6b5fe834964109294fc08ff-1999x1184.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Filtering by type The dataset includes four ‘types’, including `phototourism`, `haiper`, `heritage`, and `urban`. For example, let’s filter the samples by those of the `heritage` type. ![](https://cdn.sanity.io/images/h6toihm1/production/a48c00835230491f16bb6aeacacd408e40fbf086-1999x1448.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Filtering by id or filepath FiftyOne also makes it very easy to filter the samples to find the ones that meet your specific criteria. For example we can filter by a specific `id`: ![](https://cdn.sanity.io/images/h6toihm1/production/0081c327f9bc9facf3d7ae0428e053e997bdab99-1907x1999.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Embeddings If you opted for calculating embeddings in Step 4 you can view them by clicking on the “+” to the right of the “Samples” tab and selecting “Embeddings.” In this example we’ll select `img_vix` as our Brain Key, the `location` label, and lasso the green cluster which captures the images labeled `cyprus`. ![](https://cdn.sanity.io/images/h6toihm1/production/d51ee155f0249be1ba454bcba793474a6e7a8200-1564x700.png?auto=format&dpr=2&fit=max&q=75&w=1564) Cool! So, what can you do with embeddings? You can: - [Identify anomalous/incorrect image labels](https://docs.voxel51.com/tutorials/image_embeddings.html#Tagging-label-mistakes) - [Find examples of scenarios of interest](https://docs.voxel51.com/tutorials/image_embeddings.html#Investigating-outliers) - [Pre-annotate unlabeled data for training](https://docs.voxel51.com/tutorials/image_embeddings.html#Pre-annotation-of-samples) Learn more about how to work with embeddings in the [FiftyOne Docs.](https://docs.voxel51.com/tutorials/image_embeddings.html) ## Start working with the dataset Now that you have a general idea of what the dataset contains, you can start using FiftyOne to perform a variety tasks including: - [Creating dataset views](https://voxel51.com/docs/fiftyone/user_guide/using_views.html) - [Creating aggregations](https://voxel51.com/docs/fiftyone/user_guide/using_aggregations.html) - [Creating interactive plots](https://voxel51.com/docs/fiftyone/user_guide/plots.html) - [Annotating datasets](https://voxel51.com/docs/fiftyone/user_guide/annotation.html) - [Evaluating models](https://voxel51.com/docs/fiftyone/user_guide/evaluation.html) You can also start making use of the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html) which provides powerful machine learning techniques you can apply to your workflows like visualizing embeddings, finding similarity, uniqueness, and mistakenness. ## Sponsors and acknowledgements This competition is being sponsored by Haiper (Canada) Ltd., Google, and Kaggle.The individuals responsible for putting together the competition include: Eduard Trulls (Google), Dmytro Mishkin (Czech Technical University in Prague/HOVER Inc), Jiri Matas (Czech Technical University in Prague), Fabio Bellavia (University of Palermo), Luca Morelli (University of Trento/Bruno Kessler Foundation), Fabio Remondino (Bruno Kessler Foundation), Weiwei Sun (University of British Columbia), and Kwang Moo Yi (University of British Columbia/Haiper). [Image Matching Challenge 2023](https://voxel51.com/blog/tag/image-matching-challenge-2023) [Kaggle](https://voxel51.com/blog/tag/kaggle) [Kaggle competition](https://voxel51.com/blog/tag/kaggle-competition) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/40381f5f37fa5fcd70eddca2f63b6710568f5d2c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visual Kinship Recognition with the Families in the Wild Computer Vision Dataset\\ \\ Datasets\\ \\ • \\ \\ Dec 7, 2022](https://voxel51.com/blog/visual-kinship-recognition-with-the-families-in-the-wild-computer-vision-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/33d08c7b16ab5bfa0e4c5a4f936be4592b8e0a90-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the Berkeley Deep Drive Autonomous Vehicle Dataset\\ \\ Datasets\\ \\ • \\ \\ Jan 11, 2023](https://voxel51.com/blog/exploring-the-berkeley-deep-drive-autonomous-vehicle-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/c4a04d43252eacdfaa8acbfafe00210d4e05f92c-1090x754.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Kinetics Dataset: Train and Evaluate Video Classification Models\\ \\ Datasets\\ \\ • \\ \\ Apr 13, 2022](https://voxel51.com/blog/the-kinetics-dataset-train-and-evaluate-video-classification-models) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-277-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – April 21, 2023 Apr 22, 2023 • 3 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023#1878647d1ff1) [Basic operations on your dataset using FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023#ff64c0ee0b3c) [Concatenating generated views such as patches in FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023#95e4aac48464) [Exporting FiftyOne datasets](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023#10bb2c1c4a64) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023#9c067f9ca8b9) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023#1878647d1ff1) [Basic operations on your dataset using FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023#ff64c0ee0b3c) [Concatenating generated views such as patches in FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023#95e4aac48464) [Exporting FiftyOne datasets](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023#10bb2c1c4a64) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023#9c067f9ca8b9) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. ![](https://cdn.sanity.io/images/h6toihm1/production/d6290c63fcc4c4f62e3bc46ae391317155adf6f2-960x540.gif?auto=format&dpr=2&fit=max&q=75&w=960) - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## Basic operations on your dataset using FiftyOne Community Slack member Earl asked, _“I want to initially use the FiftyOne App to quickly look at my YOLOv5 dataset and annotations. I'd like to analyze my data and do some very basic tasks like remove duplicates, show distributions, etc. It isn't clear what I can do with the App without writing Python scripts.”_ Hey Earl! You’ll need to first load your dataset into FiftyOne either through the Python SDK or the command-line interface. But loading a [YOLOv5 dataset](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#yolov5dataset) into the App takes only a few lines of code: ```python 1import fiftyone as fo 2name = "my-dataset" 3dataset_dir = "/path/to/yolov5-dataset" 4 5# The splits to load 6splits = ["train", "val"] 7 8# Load the dataset, using tags to mark the samples in each split 9dataset = fo.Dataset(name) 10for split in splits: 11 dataset.add_dir( 12 dataset_dir=dataset_dir, 13 dataset_type=fo.types.YOLOv5Dataset, 14 split=split, 15 tags=split, 16) 17 18# View summary info about the dataset 19print(dataset) 20 21# Print the first few samples in the dataset 22print(dataset.head()) ``` Once your dataset is in FiftyOne, you can visualize it in the App, view distributions, filter on your labels, and much more. Finding duplicate samples will require [a couple more lines of code](https://docs.voxel51.com/user_guide/brain.html#finding-near-duplicate-images) to compute similarity. For example, let’s use the [find\_duplicates()](https://docs.voxel51.com/api/fiftyone.brain.similarity.html#fiftyone.brain.similarity.DuplicatesMixin.find_duplicates) method to find near-duplicate examples based on the provided parameters. ```python 1# Use the similarity index to identify the 1% of images that are least 2# similar w.r.t. the other images 3 4results.find_duplicates(fraction=0.01) 5print(results.neighbors_map) ``` More information on [finding near duplicate images](https://docs.voxel51.com/user_guide/brain.html#finding-near-duplicate-images) is available in the FiftyOne docs. ## Concatenating generated views such as patches in FiftyOne Community Slack member Joy asked, _“I want to load the dataset object as patches directly. Is there a way to load the patch view directly in FiftyOne?”_ Hi Joy! FiftyOne provides a convenient way to generate patches from your images using the to\_patches() method. This method creates a PatchesView for a given SampleCollection, which allows you to view and manipulate patches in your dataset. To concatenate generated patch views, you can first use the to\_patches() method to generate a patch view of your desired SampleCollection, and then use list concatenation to combine multiple patch views. Here's an example code snippet that demonstrates how to concatenate two patch views in FiftyOne: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# Load a dataset 5dataset = foz.load_zoo_dataset("quickstart") 6 7# Generate a patches view 8patches_view = dataset.to_patches("ground_truth") 9 10# Get the first 10 patches 11patches1 = patches_view[:10] 12 13# Get the last 10 patches 14patches2 = patches_view[-10:] 15 16# Concatenate the patches 17patches = patches1 + patches2 18 19# Check that the length of the concatenated patches matches the sum of the lengths of the individual patch views 20print(len(patches) == len(patches1) + len(patches2)) ``` Here’s the snapshot of the expected output: For more information on [patch views](https://docs.voxel51.com/user_guide/using_views.html#object-patches-views), please visit FiftyOne Docs. ## Exporting FiftyOne datasets Community member, Jack asked: _“I'm using .export() to export my dataset and create a custom coco json. The filenames aren't absolute. How do I export the full path for my filename? It only exports the filename.”_ Hi Jack! FiftyOne provides native support for exporting datasets to disk in a variety of common formats, and it can be easily extended to export datasets in custom formats. The export() method provides additional parameters that you can use to configure the export. For example, you can use the data\_path and labels\_path parameters to independently customize the location of the exported media and labels, including labels-only exports: ```python 1# Export **only** labels in the `ground_truth` field in COCO format 2# with absolute image filepaths in the labels 3 4dataset_or_view.export( 5 dataset_type=fo.types.COCODetectionDataset, 6 labels_path="/path/for/export.json", 7 label_field="ground_truth", 8 abs_paths=True, 9) ``` Or you can use the export\_media parameter to configure whether to copy, move, symlink, or omit the media files from the export: ```python 1# Export the labels in the `ground_truth` field in COCO format, and 2# move (rather than copy) the source media to the output directory 3 4dataset_or_view.export( 5 export_dir="/path/for/export", 6 dataset_type=fo.types.COCODetectionDataset, 7 label_field="ground_truth", 8 export_media="move", 9) ``` For more information on [customizing the exports](https://docs.voxel51.com/user_guide/export_datasets.html#exporting-fiftyone-datasets), you can refer to FiftyOne documents. ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,500+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,800+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,700+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 265+ repositories - 59+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [exporting](https://voxel51.com/blog/tag/exporting) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0ecb0645c4938217bcade4d3d80cf59f7b05329b-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision View Stages Tips and Tricks – Jan 20, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4ac1a727dc192a21563cde51b6e345f620e09376-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-278-lllmstxt|> ## FiftyOne 0.20 Webinar Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Webinar Recap: What’s New in FiftyOne 0.20 for Computer Vision Apr 24, 2023 • 6 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#38190bda90da) [What Is FiftyOne?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#1434e74234d9) [What Are the New Features in FiftyOne 0.20?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#8e4031531018) [Prerequisites](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#bbd2dbdc0203) [Natural Language Search](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#8df68e5c4a64) [Quadrant and Pinecone Integrations](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#43f1abb02cbd) [Enhanced Similarity API](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#d0dc47620247) [Upgrades to Point Cloud Support](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#00fa4829fad3) [Other Notes](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#5182ecf2daa4) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#38190bda90da) [What Is FiftyOne?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#1434e74234d9) [What Are the New Features in FiftyOne 0.20?](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#8e4031531018) [Prerequisites](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#bbd2dbdc0203) [Natural Language Search](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#8df68e5c4a64) [Quadrant and Pinecone Integrations](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#43f1abb02cbd) [Enhanced Similarity API](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#d0dc47620247) [Upgrades to Point Cloud Support](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#00fa4829fad3) [Other Notes](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision#5182ecf2daa4) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We recently released [FiftyOne 0.20](https://voxel51.com/blog/announcing-fiftyone-0-20/), which is packed with exciting new features to help you organize, visualize, search, and explore your computer vision datasets. Voxel51 Co-Founder and CTO [Brian Moore](https://www.linkedin.com/in/brimoor/) walked us through the new features in a live webinar, with plenty of live demos and code examples, so that you can see all the awesomeness in action. You can watch the video playback on [YouTube](https://www.youtube.com/watch?v=DUWfP3tNQSc), take a look at the [slides](https://docs.google.com/presentation/d/10942u2FPN4Y1LtusUHDIEUOKnjufeDjCIwBDsVioWf0/edit?usp=sharing), read the [transcript](https://www.rev.com/transcript-editor/shared/vWtaPq-3lj92-uPxio_aKL0OQyA4ibJAYz4goqjLqNAbRJSc0Eu6c_SG1yTeCiUTE4aeqU06oP95gI8i1n_-lj3ck6U?loadFrom=SharedLink), and read the recap below for the highlights. Enjoy! https://www.youtube.com/watch?v=DUWfP3tNQSc ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes was [Wildlife AI](https://www.wildlife.ai/). We were first introduced to Wildlife AI through the FiftyOne community! They are using FiftyOne to enable their users to easily analyze the camera data and create their own models. We are sending a charitable donation of $200 to Wildlife AI on behalf of the computer vision community. ![](https://cdn.sanity.io/images/h6toihm1/production/e80959f169a6a89e3cef73f9bb1bdd1750499fe7-770x146.png?auto=format&dpr=2&fit=max&q=75&w=770) ## What Is FiftyOne? Brian starts with a quick overview of what FiftyOne is for those who might be new to it: “think of FiftyOne as glue between your datasets and your models, and also between your data-centric tools and your model-centric tools.” FiftyOne helps you curate high quality datasets on the left and feed them to your models on the right. Nestled at the center of your data and models, FiftyOne unlocks dozens of computer vision workflows so you can continuously build high quality data, high performing models, and production-grade AI. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Here’s a snapshot of some of the computer vision workflows made possible by FiftyOne: - Visualize, query, and analyze computer vision datasets - Streamline data annotation workflows - Identify and correct labeling mistakes - Analyze model performance, both visually and programmatically - And dozens more workflows! You can get the full FiftyOne overview in the first five minutes of the presentation [recap video](https://www.youtube.com/watch?v=DUWfP3tNQSc). ## What Are the New Features in FiftyOne 0.20? Here are the new features in FiftyOne 0.20 that Brian demonstrated in the webinar and you can read the highlights in the sections below: - Natural language search: you can now perform arbitrary search-by-text queries natively in the FiftyOne App and Python SDK, leveraging multimodal vector indexes on your datasets under-the-hood - Qdrant and Pinecone integrations: new integrations with Qdrant and Pinecone to power text/image similarity queries - Similarity API: significant upgrades to the FiftyOne Brain’s similarity API, including configurable vector database backends and the ability to modify existing indexes - Point cloud-only datasets: you can now create datasets composed only of point cloud samples and visualize them in the App’s grid view ## Prerequisites Before diving into the features, Brian first installed FiftyOne with pip install fiftyone, loaded a dataset from the Dataset Zoo with 200 images, loaded in a model from the Model Zoo that can generate embeddings for both language and images, and then computed embeddings (specifically a similarity index which is at the heart of the new features being demonstrated). Now that there’s a dataset and an index on it, Brian launches the FiftyOne App to show us the new features in action! You can find the prerequisite steps here in the [Jupyter notebook](https://drive.google.com/file/d/15IEvoXNE0F2lAKWHUYynnKohmI86_K9t/view?usp=sharing), or consider joining us for a Getting Started with FiftyOne Workshop – you can find a variety of [upcoming dates and times listed on our events page](https://voxel51.com/computer-vision-events/). Half lecture and half lab, you’ll walk away with everything you need to get up and running with open source FiftyOne. ## Natural Language Search To use [natural language search](https://docs.voxel51.com/user_guide/app.html#text-similarity) now that the prerequisite steps are done, Brian clicks the magnifying glass icon in the FiftyOne App samples grid bar and first searches for “puppies,” but then tries searching a string that’s a little more complicated: “kites high in the air.” What returns by default are the 25 most closely matching images in the dataset (because 25 is the default setting, but can be configured). \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop This dataset may have annotations, but this search feature is completely unsupervised. It’s not relying on annotations; it’s only based on the embeddings that we added to the dataset in the prerequisite step. You can see the demo of the natural language search feature from [~08:37 - 17:42](https://www.youtube.com/watch?v=DUWfP3tNQSc&t=517s) in the presentation, including using a subset of the COCO dataset (the validation split) with 5,000 samples. Brian even goes under the hood to explain how this all works: vector search! Furthermore, not only can you search by natural language, you can also perform an image to image search. Simply use the image similarity icon in the sample grid and find similar images in the dataset to the one you selected. So far Brian has demonstrated vector searches being done using a built-in vector database that runs in memory. But if you want to scale this workflow to very large datasets, let's say millions of images, and you want to be able to quickly do searches across those millions of data points, then it would be more performant to consider using a dedicated vector database solution, which is described in the next section! ## Quadrant and Pinecone Integrations As of the 0.20 release, FiftyOne integrates with two vector databases: Qdrant and Pinecone. Earlier in the presentation, Brian ran the [compute\_similarity()](https://docs.voxel51.com/api/fiftyone.brain.html#fiftyone.brain.compute_similarity) method. The method now supports a backend parameter that you can either set to use the built-in database, or you can specify one of the two supported backends. Assuming you've configured your Qdrant or Pinecone vector database, then you can generate indices for your FiftyOne datasets where the vectors are stored in that separate database. Then performing similarity searches in the FiftyOne App will query that database rather than the built-in one. Brian dives deep into both integrations and shows multiple compelling computer vision workflows. You can watch it all from [~17:42 - 36:28](https://www.youtube.com/watch?v=DUWfP3tNQSc&t=1062s) in the presentation, and you can learn more about them in the integration docs: [Qdrant](https://docs.voxel51.com/integrations/qdrant.html) and [Pinecone](https://docs.voxel51.com/integrations/pinecone.html). ## Enhanced Similarity API Brian notes that we’ve covered the upgrades to the FiftyOne Brain’s similarity API, because it’s what’s makes all this possible from the presentation so far: - Use default backend or configure a custom one (Qdrant, Pinecone, or add your own) - Initialize an empty index - Add vectors to an existing index - Retrieve vectors from an index - Remove vectors from an index Previously, similarity indexes were static objects that could not be edited once created. In FiftyOne 0.20, the similarity indices are now mutable! Get this quick summary of the similarity API updates from [~36:28 - 37:02](https://www.youtube.com/watch?v=DUWfP3tNQSc&t=2188s) in the presentation, and here in the [Similarity API](https://docs.voxel51.com/user_guide/brain.html#similarity-api) docs. ## Upgrades to Point Cloud Support Previously, working with point cloud datasets in the App was to add point cloud samples as slices of [grouped datasets](https://docs.voxel51.com/user_guide/groups.html) that also contain other media modalities (image, video, etc). However, in FiftyOne 0.20 you can now create datasets that contain [only point cloud samples](https://docs.voxel51.com/user_guide/using_datasets.html#point-cloud-datasets) and work with them natively in the App’s grid and modal views. Brian explains how: “You can run a utility that exists in the tool to generate projection images. So, what it's doing is taking each of the point clouds in the dataset and generating what we're calling an orthographic projection image, like a top-down projection of that point cloud, and then storing those in a new field of the dataset. And the reason that's useful is when you open up the grid view, you have a fast rendering representation of each of those point clouds so you can quickly scan through to find the one of interest. And then if you want to click into it and work with it in three dimensions, you can then open the modal and work with the full 3D scene.” Watch Brian’s demo of the new point cloud features from [~37:02 - 45:18](https://www.youtube.com/watch?v=DUWfP3tNQSc&t=2222s) in the presentation. Also check out [the docs](https://docs.voxel51.com/user_guide/using_datasets.html#point-cloud-datasets) for more information about adding point cloud samples and orthographic projections to your FiftyOne datasets and visualizing them in the App. ## Other Notes After demonstrating the new features, Brian shares a few additional points before concluding the presentation. Open source software like FiftyOne doesn’t happen without an amazing community supporting it. Brian gives a shout out to the community members who contributed to FiftyOne 0.20. Also, we’re always open to new contributions and contributors! Check out the [good first issue](https://github.com/voxel51/fiftyone/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22) label on GitHub. If you are interested in working on one of those, feel free to leave a comment there and we will be happy to help you out in any way. [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne 0.20](https://voxel51.com/blog/tag/fiftyone-0-20) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/308698a5aece1d5b1b95ee1bf52811b24448458c-1200x672.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne 0.18 for Computer Vision\\ \\ Event Recaps\\ \\ • \\ \\ Dec 6, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/17422a76c76f14096dce21e43da51945f498a811-1200x673.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne & FiftyOne Teams\\ \\ Event Recaps\\ \\ • \\ \\ Oct 8, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-fiftyone-teams) [![](https://cdn.sanity.io/images/h6toihm1/production/26fba75b02878a8803a43bf05ea3dd9153e4e491-1199x675.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne 0.19 for Computer Vision\\ \\ Event Recaps\\ \\ • \\ \\ Mar 2, 2023](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-19-for-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-279-lllmstxt|> ## FiftyOne Workshop Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Getting Started with FiftyOne Workshop – April 26 Recap Apr 27, 2023 • 6 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#d119508442a0) [Wait, what’s FiftyOne?](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#7373a232b881) [Workshop summary](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#3ebfa9122f9b) [Lecture: get up-to-speed on the basics](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#84eeda8f5c23) [Lab: fire up FiftyOne and experience it for yourself!](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#0e4bd0e3dd2a) [Q&A recap](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#91ea521a862d) [Join an upcoming event](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#184cc5fd68dc) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#d119508442a0) [Wait, what’s FiftyOne?](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#7373a232b881) [Workshop summary](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#3ebfa9122f9b) [Lecture: get up-to-speed on the basics](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#84eeda8f5c23) [Lab: fire up FiftyOne and experience it for yourself!](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#0e4bd0e3dd2a) [Q&A recap](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#91ea521a862d) [Join an upcoming event](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap#184cc5fd68dc) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop [Jacob Marks](https://www.linkedin.com/in/jacob-marks/), PhD and Machine Learning Engineer at Voxel51, recently presented the _Getting Started with FiftyOne Workshop_, which is part of a series of hands-on, educational events to show you step-by-step how to use the open source FiftyOne computer vision toolset. In this blog post, we summarize the highlights, recap the questions and answers from the event, and share the upcoming schedule of events. We’d love to see you at a future event! ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes was [Wildlife AI](https://www.wildlife.ai/). We were first introduced to Wildlife AI through the FiftyOne community! They are using FiftyOne to enable their users to easily analyze the camera data and create their own models. We are sending a charitable donation of $100 to Wildlife AI on behalf of the computer vision community who participated in this event! ![](https://cdn.sanity.io/images/h6toihm1/production/e80959f169a6a89e3cef73f9bb1bdd1750499fe7-770x146.png?auto=format&dpr=2&fit=max&q=75&w=770) ## Wait, what’s FiftyOne? The _Getting Started with FiftyOne Workshop_ was created to help you get up-to-speed on the basics of using open source FiftyOne in your computer vision workflows. If you’re new to FiftyOne – you may be wondering, what is it? FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. ![](https://cdn.sanity.io/images/h6toihm1/production/136887bcc1e07d86237d0f79c0f2bf731805f6fb-1280x720.gif?auto=format&dpr=2&fit=max&q=75&w=1280) ## Workshop summary Half lecture, half lab, the workshop covered these essentials: Lecture: - FiftyOne Basics (terms, architecture, installation, and general usage) - An overview of useful workflows to explore, understand, and curate your data - How FiftyOne represents and semantically slices unstructured computer vision data Lab: - Loading datasets from the FiftyOne [Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html) - How to easily navigate the FiftyOne App’s features - Programmatically inspecting attributes of a dataset - Adding new sample and custom attributes to a dataset - Generating and evaluating model predictions - How to save insightful views into the data … All with the goal of helping you gain greater visibility into the quality of your computer vision datasets and models. ## Lecture: get up-to-speed on the basics Jacob walked us through some popular ways to use FiftyOne. ### Curate data - Find: filter, match, sort, select - Remove: duplicates - Add: tags, metadata, predictions - Correct: annotation mistakes - Save: interesting “views” ### Understand data - Aggregate statistics: FiftyOne makes it easy to compute summary statistics about your datasets and views, including histograms and all of the traditional aggregations for numerical quantities you would expect: min, max, mean, standard aviation, and more. - Embeddings: Visualizing your dataset in a low-dimensional embedding space is a powerful workflow that can help you uncover hidden patterns and clusters in your data so you can take action to improve the quality of your datasets and models. - Interactive visualization: All these visualizations are interactive. Simply lasso points in an embeddings plot to see just those samples, explore a cell in a confusion matrix to see just those samples, and more. ### Evaluate data FiftyOne has support for tons of one-number metrics: precision, recall, F1 score, intersection over union, and more. There’s support for all of your favorite plots including PR curves and confusion matrices. You can also perform analysis on samples, labels, and entire datasets. ### Tap into the flexibility of FiftyOne Because computer vision is not one-size-fits-all, FiftyOne is designed for flexibility and customizability, across all these categories: - Datasets - Models - Media types - Labels - Plugins - More! ### Key components of FiftyOne Jacob provided a few additional tips to prepare everyone as they geared up for the hands-on lab. First, Jacob described the three main components of FiftyOne: the FiftyOne Library, the FiftyOne App, and the FiftyOne Brain. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Then covered two additional helpful concepts: - A description of working with tabular data (structured data) vs computer vision data (unstructured data), and how FiftyOne can be thought of as the pandas of computer vision - A look under the hood of a schema - including a dataset, samples, fields, metadata, filepath, labels, media type, and more ## Lab: fire up FiftyOne and experience it for yourself! The hands-on lab was the star of the second half, enabling you to put your learnings from the lecture into action. The outcome – everyone fired up FiftyOne, explored datasets and models firsthand, and experienced how to: - Install FiftyOne - Load datasets and models from the FiftyOne [Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html) and FiftyOne Model Zoo - Easily navigate the FiftyOne App’s features - Programmatically inspect attributes of a dataset - Add new samples and custom attributes to a dataset - Evaluate model predictions - Save insightful views into the data - More! ## Q&A recap **Does FiftyOne support segmentation?** Yes! Learn more about FiftyOne’s support for [instance segmentation](https://docs.voxel51.com/user_guide/using_datasets.html#instance-segmentations) and [semantic segmentation masks](https://docs.voxel51.com/user_guide/using_datasets.html#semantic-segmentation) in the [FiftyOne Docs](https://docs.voxel51.com/user_guide/using_datasets.html#labels). **Does FiftyOne work with the Segment Anything Model?** Yes! Look for a blog on this topic in the near future. **Can I import a 3D point cloud dataset into FiftyOne?** Absolutely! Check out this [blog](https://voxel51.com/blog/visualize-3d-point-clouds-and-work-with-openai-point-e/) and this [tutorial](https://docs.voxel51.com/tutorials/pointe.html) to learn more about how to work with 3D point cloud data in FiftyOne. If you want to work with point clouds as part of grouped datasets, see the [FiftyOne User Guide](https://docs.voxel51.com/user_guide/groups.html). **Is FiftyOne similar to YOLO?** YOLO is a set of (You Only Look Once) object detection models, while FiftyOne is a toolset that allows you to manage and curate your computer vision data. In a typical workflow they would be used together with FiftyOne boosting the performance of the YOLO model by improving the quality of the data going in. Check out [this YOLO tutorial](https://docs.voxel51.com/tutorials/yolov8.html) for additional details. **In a dataset, can multiple samples point to the same source media file?** Yes! **Do the uniqueness and similarity features in FiftyOne use localized distance metrics?** Out-of-the-box, FiftyOne selects a default metric, as well as a default model, for FiftyOne Brain methods like [uniqueness](https://docs.voxel51.com/user_guide/brain.html?highlight=uniqueness#brain-image-uniqueness) and [similarity](https://docs.voxel51.com/user_guide/brain.html?highlight=uniqueness#brain-similarity). However, if you’d like, you can also specify these via keyword arguments. For instance, [for a Qdrant similarity backend](https://docs.voxel51.com/integrations/qdrant.html#qdrant-config-parameters), you can use a \`cosine\` similarity metric with \`metric=cosine\`. Different vector search backends support different metrics. Learn more about how to use uniqueness and similarity in the [FiftyOne Docs](https://docs.voxel51.com/user_guide/brain.html). **What types of integrations does FiftyOne support?** FiftyOne supports a variety of popular platforms and tools including COCO, PyTorch, AWS, Google Cloud, Qdrant, and more. For a complete list of documented integrations, check out [the integration Docs](https://docs.voxel51.com/integrations/index.html). **What options does FiftyOne support for model evaluations?** FiftyOne provides a variety of builtin methods for evaluating your model predictions, including regressions, classifications, detections, polygons, instance and semantic segmentations, on both image and video datasets. It also supports custom metrics. For a variety of tips and tricks concerning evaluations, check out [this blog](https://voxel51.com/blog/fiftyone-computer-vision-model-evaluation-tips-and-tricks-feb-03-2023/) and [these Docs](https://docs.voxel51.com/user_guide/evaluation.html?highlight=evaluation). **What backend database does FiftyOne use?** FiftyOne uses MongoDB as a backend. Learn more about [configuring a MongoDB backend in the Docs](https://docs.voxel51.com/user_guide/config.html#configuring-a-mongodb-connection). **Does FiftyOne have utilities for dataset transformations and/or conversions?** Yes! A good place to start is with the documentation concerning [Custom Importers](https://docs.voxel51.com/recipes/custom_importer.html) and [Custom Exporters](https://docs.voxel51.com/recipes/custom_exporter.html). For a deeper dive, check out the [FiftyOne data utils](https://docs.voxel51.com/api/fiftyone.utils.data.html). ## Join an upcoming event We are excited to have a number of other events already lined up and hope to see you there! Upcoming Computer Vision Meetups: - May ’23 Computer Vision Meetup (Americas & EMEA) - May ’23 Computer Vision Meetup (APAC) - June ’23 Computer Vision Meetup (Americas, EMEA) Additional dates & times for the Getting Started with FiftyOne Workshop: - May 31 @ 4 PM BST \[11 AM EDT / 15:00 UTC\] - June 28 @ 10 AM PDT \[1 PM EDT / 17:00 UTC\] [See the full schedule.](https://voxel51.com/computer-vision-events/) [Getting Started with FiftyOne Workshop](https://voxel51.com/blog/tag/getting-started-with-fiftyone-workshop) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/83d7ece0c635d6204dd333b0175cc55ec02c8c4a-1199x675.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Getting Started with FiftyOne Workshop – March 29 Recap\\ \\ Event Recaps\\ \\ • \\ \\ Apr 4, 2023](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-march-29-recap) [![](https://cdn.sanity.io/images/h6toihm1/production/308698a5aece1d5b1b95ee1bf52811b24448458c-1200x672.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Webinar Recap: What’s New in FiftyOne 0.18 for Computer Vision\\ \\ Event Recaps\\ \\ • \\ \\ Dec 6, 2022](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-18-for-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ed751b5c0fbc3d2cb74f0b30e6418d3319564c-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — December 2022\\ \\ Event Recaps\\ \\ • \\ \\ Dec 13, 2022](https://voxel51.com/blog/recapping-the-computer-vision-meetup-december-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-280-lllmstxt|> ## Generate Movement from Text [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) Generate Movement from Text Descriptions with T2M-GPT Apr 28, 2023 • 16 min read Article content In this article [Human motion synthesis](https://voxel51.com/blog/generate-movement-from-text-descriptions-with-t2m-gpt#e5769379f454) [Transformers in motion](https://voxel51.com/blog/generate-movement-from-text-descriptions-with-t2m-gpt#8debf842fdcc) [Conclusion](https://voxel51.com/blog/generate-movement-from-text-descriptions-with-t2m-gpt#dfb1f4e3e89e) In this article [Human motion synthesis](https://voxel51.com/blog/generate-movement-from-text-descriptions-with-t2m-gpt#e5769379f454) [Transformers in motion](https://voxel51.com/blog/generate-movement-from-text-descriptions-with-t2m-gpt#8debf842fdcc) [Conclusion](https://voxel51.com/blog/generate-movement-from-text-descriptions-with-t2m-gpt#dfb1f4e3e89e) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) _Editor’s note: This is a guest post by [Chien Vu](https://www.linkedin.com/in/vumichien/), PhD, Machine Learning Researcher at Detomo, Japan_ \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ![](https://cdn.sanity.io/images/h6toihm1/production/3b028b7d87ab4d03213479774f27ca4a63742208-186x167.gif?auto=format&dpr=2&fit=max&q=75&w=186)![](https://cdn.sanity.io/images/h6toihm1/production/cfb3c35a8817df2a121f2a68d55a954fa62a95e8-186x167.gif?auto=format&dpr=2&fit=max&q=75&w=186)![](https://cdn.sanity.io/images/h6toihm1/production/8a933573e547b5284f6a5e09db3ee4430f5570c7-186x167.gif?auto=format&dpr=2&fit=max&q=75&w=186) Creating movement based on written descriptions has various applications in industries such as gaming, filmmaking, robotics animation, and the future of the metaverse. Conventionally, **motion capture**, also known as mocap, is a technique used to capture the movements of an object or individual in real time and convert them into digital data. It involves placing reflective markers on the object or individual's body, which are then tracked by cameras or other sensors. The captured data can then be used to animate digital characters, create visual effects, or analyze movement for scientific or medical research. Motion capture is widely used in the entertainment industry for films, video games, and virtual reality experiences, as well as in sports science, biomechanics, and robotics. However, motion capture is a relatively expensive process due to the equipment required (high-quality cameras, specialized software, and reflective markers), people to operate the equipment and ensure high quality results, and other factors. Generating movement from written descriptions could drastically reduce the time and money required to produce games and movies, opening up new avenues for creators. In this blog post, we present a novel framework called **T2M-GPT** that utilizes a **Vector Quantised Variational AutoEncoder (VQ-VAE)** and a **Generative Pretrained Transformer (GPT)** to generate human motion capture from textual descriptions. The **T2M-GPT** framework achieves superior performance compared to other state-of-the-art approaches, including recent diffusion-based methods. Notably, T2M-GPT produces high-quality motion and demonstrates comparable consistency between text and generated motion. Continue reading for further details about T2M-GPT, including why we created it, how it works, and how to try it out. ## Human motion synthesis Human motion synthesis, also known as motion generation, is a computer graphics technique used to generate realistic and natural-looking human movements in a virtual environment. It involves developing algorithms and models that can simulate human motion, typically by utilizing motion capture data, physical laws, and artificial intelligence techniques. The goal of human motion synthesis is to create virtual characters that can move and interact with their environment in a manner that appears natural and believable to human observers. There are three different methods for generating it: future motion prediction, monocular motion estimation, and conditioned motion synthesis. ### Future motion prediction Future motion prediction is the process of using available information to predict the future trajectory or movement of an object or entity. This can be accomplished through various means, such as mathematical models, machine learning algorithms, or physical simulations. Predicting future frames from past motion or an initial pose has a history dating back to the 1980s, with statistical models commonly used in earlier studies. Recent research has shown promising results using generative models with neural networks such as [DLow](https://arxiv.org/abs/2003.08386) \[1\] (Diversifying Latent Flows) and [SoMoFormer](https://arxiv.org/abs/2208.14023v1) \[2\] (Social Motion Transformer). DLow uses a single random variable and a set of learnable mapping functions to generate correlated latent codes, which are then decoded into a diverse set of correlated samples. During training, DLow optimizes the latent mappings to promote sample diversity using a flexible diversity-promoting prior. The prior can be customized to generate diverse motions with common features. Meanwhile, SoMoFormer is proposed for multi-person 3D pose forecasting. The unique transformer architecture models human motion input as a joint sequence, allowing attention over joints while predicting the entire future motion sequence for each joint in parallel. SoMoFormer extends to multi-person scenes by using the joints of all individuals in a scene as input queries, and by learning relationships between joints and between people. ### Monocular motion estimation Monocular motion estimation refers to the process of estimating the motion of an object or scene using a single camera. This is done by analyzing the changes in the image captured by the camera over time, and using this information to determine the direction and speed of motion. There are several techniques used for monocular motion estimation, including optical flow (Figure 1), structure from motion (Figure 2), and visual odometry. Optical flow is a technique that estimates the motion of image pixels over time, while structure from motion uses the relative motion of different points in the scene to estimate the camera's motion like [VIBE](https://arxiv.org/pdf/1912.05656.pdf) \[3\] (Video Inference for Body Pose and Shape Estimation). ![](https://cdn.sanity.io/images/h6toihm1/production/96c762682cf03de04521f814a785de5e9c148ea9-512x296.gif?auto=format&dpr=2&fit=max&q=75&w=512) VIBE uses an existing large-scale motion capture dataset (AMASS) along with unpaired, in-the-wild, 2D keypoint annotations. The main innovation of VIBE is an adversarial learning framework that uses AMASS to distinguish between real human motions and those produced by the temporal pose and shape regression networks. By using a temporal network architecture and performing adversarial training at the sequence level, VIBE can generate kinematically plausible motion sequences without the need for in-the-wild ground-truth 3D labels. ![](https://cdn.sanity.io/images/h6toihm1/production/1249e63690364976a9c477fa0efc697ffda7ca6a-1600x400.gif?auto=format&dpr=2&fit=max&q=75&w=1600) In addition to optical flow and structure from motion, another technique used for monocular motion estimation is visual odometry, which estimates the 3D structure of the scene based on the camera's motion. Recently, an end-to-end learning approach called [PoseConvGRU](https://arxiv.org/abs/1906.08095) \[4\] has been developed to directly map input image pairs to ego-motion estimates using a two-module Long-term Recurrent Convolutional Neural Network. The feature-encoding module of this network captures short-term motion features in an image pair, while the memory-propagating module captures long-term motion features in consecutive image pairs. The visual memory is implemented with convolutional gated recurrent units (ConvGRU), which propagate information over time. This method stacks two consecutive RGB images to extract motion information and estimate poses, which are then passed through a stacked ConvGRU module to generate the relative transformation pose for each image pair. ### Conditioned motion synthesis While there is a lot of research on future motion prediction, synthesis from scratch has received less attention. Early work used statistical models like [PCA](https://www.sciencedirect.com/science/article/abs/pii/S0262885605001526) \[5\] and [GPLVMs](https://link.springer.com/chapter/10.1007/978-3-540-75703-0_8) \[6\] to learn cyclic motions; and the next generation of synthesis involved guiding conditioning with short text or language representations like [Text2Action](https://arxiv.org/pdf/1710.05298.pdf) \[7\] and [Language2Pose](https://arxiv.org/abs/1907.01108) \[8\]. Text2Action is a generative adversarial network (GAN) built on a sequence-to-sequence (SEQ2SEQ) architecture that can generate a sequence of human actions based on a sentence describing human behavior. The Text2Action model includes a text encoder recurrent neural network (RNN) and an action decoder RNN, enabling the synthesis of various actions for a robot or virtual agent. Language2Pose tries to map linguistic concepts to motion animations by learning a joint embedding of language and pose. This joint embedding space is learned end-to-end using a curriculum learning approach that emphasizes shorter and easier sequences before longer and harder ones. To generate human motion from a joint embedding for language and motion, these generative deep learning models should be able to learn precise mapping from the language space to the motion space. Recently [MotionCLIP](https://arxiv.org/abs/2203.08063) \[9\] aligns its latent space with that of the Contrastive Language-Image Pre-training ( [CLIP](https://arxiv.org/pdf/2103.00020.pdf)) \[10\] model to imbue rich semantic knowledge into the human motion manifold; and [ACTOR](https://arxiv.org/pdf/2104.05670.pdf) \[11\] and [TEMOS](https://arxiv.org/abs/2204.14109) \[12\] proposed transformer-based VAEs for action-to-motion and text-to-motion. While these models represent important steps on the path to generating motion from text, they are limited in their ability to produce high-quality motion for longer and more complex textual descriptions. Additionally, some of these models utilize a complex three-stage process for text-to-motion generation but sometimes they still fail to generate high-quality motion consistent with the text. To improve the performance of motion generation, it is recommended that we explore methods to elevate the traditional approach for acquiring discrete representations with the advanced capabilities of high-quality large language models such as Generative Pre-trained Transformer (GPT). That’s why Text to Image-Generative Pre-trained Transformer (T2M-GPT) was proposed. ## Transformers in motion The [T2M-GPT](https://arxiv.org/pdf/2301.06052.pdf) \[13\] method is a straightforward yet efficient technique that utilizes a two-stage process to produce high-quality motion sequences consistent with complex textual descriptions. T2M-GPT uses a simple and classic framework based on Vector Quantized Variational Autoencoders (VQ-VAE) to learn the discrete representation for motion generation. Comparing to complex diffusion-based techniques such as [Human Motion Diffusion Model](https://arxiv.org/abs/2209.14916) (MDM) \[14\] or [MotionDiffuse](https://arxiv.org/abs/2208.15001) \[15\], T2M-GPT yields comparable performance. Moreover, the T2M-GPT utilizes models similar to GPT, which include distinct representations, enabling the model to produce motion from text with increased precision and competitive outcomes. The two-stage process is outlined below. ### Stage 1 VQ-VAE is employed to convert motion sequences into discrete code indices. Motion VQ-VAE consists of a standard CNN-based architecture with 1D convolution (Conv1D) with stride 2, residual block (ResBlock), and ReLU activation as shown in Figure 3 below. ![](https://cdn.sanity.io/images/h6toihm1/production/497a3366d62172ce6f2ae34a66410e4c144aa156-780x633.png?auto=format&dpr=2&fit=max&q=75&w=780) VQ-VAE has exhibited notable efficacy in generative tasks pertaining to various modalities, including but not limited to image synthesis, text-to-image generation, speech gesture generation, and music generation. The success of VQ-VAE stems from its aptitude to dissociate the process of learning discrete representation and the prior. Nevertheless, a rudimentary training of VQ-VAE can encounter [codebook collapse](https://machinelearning.wtf/terms/codebook-collapse/), wherein only a meager number of codes are stimulated, leading to suboptimal performance in reconstruction and generation. The VQ-VAE frequently becomes trapped in a suboptimal state due to two underlying reasons. Firstly, the encoder is incentivized by the commitment loss (a measure to encourage the encoder output to stay close to the embedding space and to prevent it from fluctuating too frequently from one code vector to another) to solely adhere to the codebook vectors that it has been designated to, which hinders it from learning to prioritize other codebook vectors that are yet to be assigned. Secondly, solely those codebook vectors that have been allocated encoder output vectors will undergo updating. Therefore, any codebook vector that lacks an assigned encoder output vector will retain its position in the embedding space unchanged. Consequently, a considerable portion of the codebook remains unutilized. To mitigate this issue, [VQ-VAE](https://arxiv.org/abs/1711.00937) is trained using exponential moving average (EMA) \[16\] and Code Reset techniques \[16\]. EMA facilitates the convergence of VQ-VAE by expediting the learning process, without being reliant on a specific optimizer selection. On the other hand, the Code Reset techniques serve to randomly reset one of the encoder outputs from the current batch when the mean usage of a codebook vector falls below a specified threshold, ensuring that all codebook vectors are utilized and are able to provide a gradient for learning purposes (Figure 4). ![](https://cdn.sanity.io/images/h6toihm1/production/3e9934ad2aee4ecb5e91c58367a04805c3de20c9-731x344.png?auto=format&dpr=2&fit=max&q=75&w=731) ### Stage 2 In the second stage, a standard GPT-like model is trained to generate sequences of code indices from pre-trained text embeddings. To improve the model's ability to predict motion length, a specific **End token** is introduced to signify the end of a motion. The implementation of the **End token** addresses concerns that other methods face, such as requiring manual input of the motion length in MDM and MotionDiffuse \[14, 15\] which is not a feasible approach for real-world applications, because we don’t know the actual motion length given the input text description. Therefore, the **End token** is the simple solution that helps the model knows when the motion should end GPT, as a Large Language Model (LLM), has the capacity to randomly generate the next token, however training GPT for motion description is difficult. In the training phase, i − 1 correct index predicts the next index. However, in the inference phase, there is no guarantee that indices serving as conditions are correct. To address this issue, sequence corruption is employed by replacing τ × 100% (τ ∈ U\[0, 1\]) ground-truth code indices with random ones during the training process to mitigate this discrepancy (Figure 5). ![](https://cdn.sanity.io/images/h6toihm1/production/111b7d6a12956e01ea4f22afbb8ca6fb88c92938-518x317.png?auto=format&dpr=2&fit=max&q=75&w=518) ### Training For training motion VQ-VAE and T2M-GPT, two standard datasets exist for text-driven motion generations: - [KIT Motion-Language](https://arxiv.org/abs/1607.03827) (KIT-ML) \[17\] contains 3,911 human motion sequences scaled to 12.5 frames per second and 6,278 textual annotations. The total vocabulary size is 1,623 words, and each motion sequence is described in 1 to 4 sentences. The average length of descriptions is approximately 8 words. - [HumanML3D](https://openaccess.thecvf.com/content/CVPR2022/papers/Guo_Generating_Diverse_and_Natural_3D_Human_Motions_From_Text_CVPR_2022_paper.pdf) \[18\] is currently the largest 3D human motion dataset with textual descriptions. The dataset contains 14,616 human motions scaled to 20 FPS and 44,970 text descriptions. The total vocabulary size is 5,371 words. The average length of descriptions is approximately 12 words. ### Evaluating synthetic motion To evaluate the global representations of motion and text descriptions, five metrics were used: 1. R-Precision metric is utilized to evaluate the ranking of Euclidean distances between a motion sequence and 32 text embeddings, consisting of one ground-truth and 31 randomly chosen mismatched descriptions. This metric serves as an effective measure to assess the degree of compatibility between the motion sequence and the textual description. 2. Frechet Inception Distance (FID) also known as Wasserstein-2 distance is used to calculate the distribution distance between the generated and real motion on the extracted motion features. Lower scores indicate the two groups of motion are more similar, or have more similar statistics, with a perfect score being 0.0 indicating that the two groups of motion are identical. FID is an important metric widely used to evaluate the overall quality of generated motions. 3. Multimodal Distance (MM-Dist) calculates the mean Euclidean distance between the text feature and the corresponding generated motion feature. A lower score indicates that the model can comprehend the text well, while a higher score might indicate a possibility of generating a mismatched motion. 4. Diversity measures the variance of the generated motions across all action categories. To determine this, a random sampling process is performed on a set of motions generated from various action types, resulting in two subsets of equal size. The average Euclidean distances between the two subsets are then calculated. 5. Multimodality (MModality) is different from diversity; multimodality measures the average diversity within each action type. A low value of multimodality indicates that the generated motions might be similar, even with varying text prompts. Conversely, a high multimodality value indicates a variety of motions, despite using the same textual input. Thus, it is advisable to maintain a sufficiently high value. After being assessed using these metrics, VQ-VAE demonstrates a level of performance that is close to real motion, indicating the utilization of high-quality discrete representations. For the generation, T2M-GPT achieves comparable performance on text-motion consistency (R-Precision and MM-Dist) compared to the state-of-the-art method MotionDiffuse \[15\]. Manually corrupting sequences during the training of GPT brings consistent improvement, meanwhile, motion length can be predicted straightforwardly and effectively through an additional End token. Here is the same output from T2M-GPT (Figure 6 and Figure 7): ![](https://cdn.sanity.io/images/h6toihm1/production/326a09f3b2a0b68bc3ca8f7e3334a33f4dac4f90-186x167.gif?auto=format&dpr=2&fit=max&q=75&w=186)![](https://cdn.sanity.io/images/h6toihm1/production/47a0c1449ec592bb39a8f8eca3e8782f0a87472c-186x167.gif?auto=format&dpr=2&fit=max&q=75&w=186) Figure 6 (above). A person jogs in place, slowly at first, then increases speed; they then back up and squat down ![](https://cdn.sanity.io/images/h6toihm1/production/6d71a406cdbbfccb636cc20aa3517a631093fcf5-186x167.gif?auto=format&dpr=2&fit=max&q=75&w=186)![](https://cdn.sanity.io/images/h6toihm1/production/f671f9691e49c5c4a71fe13d78f8ed23ad78f7d8-186x167.gif?auto=format&dpr=2&fit=max&q=75&w=186) Figure 7 (above). Someone appears to be stretching; they turn their upper body in a counter-clockwise direction and then lean their upper body from side to side However, there are still some failure cases like the following example (Figure 8). This could be attributed to either insufficient information in the text prompt or inaccurate description of motion in the training dataset. One of the ways to improve this is preparing more training dataset or using different descriptions with the same motion to enhance the diversity of motion generation. ![](https://cdn.sanity.io/images/h6toihm1/production/8ee9fb6753188445c940f5201fe53620149e5cb9-186x167.gif?auto=format&dpr=2&fit=max&q=75&w=186)![](https://cdn.sanity.io/images/h6toihm1/production/448483e11dd0a950af9cee6ee6a8c69f638607a3-186x167.gif?auto=format&dpr=2&fit=max&q=75&w=186) Figure 8. A person slightly crouches down and walks forward then back, then around slowly Check out the [Generate Human Motion Hugging Face Space](https://huggingface.co/spaces/vumichien/generate_human_motion), where you can quickly test the performance of T2M-GPT. The Space supports both the skeleton and Skinned Multi-Person Linear 3D Model (SMPL) of the human body. The training and inference code are also provided in the [T2M-GPT GitHub repository](https://github.com/Mael-zys/T2M-GPT). ## Conclusion Utilizing a traditional structure, VQ-VAE and GPT can be combined to generate high quality human motion from textual depictions. With a fairly small training dataset (around 50% of the HumanML3D dataset), T2M-GPT has already outperformed other models such as Language2Pose, TM2T,... on several metrics (FID, MM-Dist, Diversity,...). If the size of the dataset is expanded, it could enhance the model's performance further. The model can be extended to other objects by retraining the model on a new dataset with relevant textual descriptions and motion sequences. For instance, if you want to generate motion sequences for a new object like a bicycle, you will need to gather a dataset that includes textual descriptions of different bicycle movements along with the corresponding motion sequences. You can then use this dataset to retrain the T2M-GPT model to generate motion sequences for bicycles. Models like T2M-GPT may enable people to create motion synthesis for arbitrarily complex action sequences, from text descriptions alone. As research in this field continues to progress, we may even be able to generate more complex and interactive videos that respond to user input or adapt to different environments and industries. **References**: \[1\] Ye Yuan and Kris Kitani. DLow: Diversifying Latent Flows for Diverse Human Motion Prediction. [https://arxiv.org/abs/2003.08386](https://arxiv.org/abs/2003.08386) \[2\] Edward Vendrow, Satyajit Kumar, Ehsan Adeli, and Hamid Rezatofighi. SoMoFormer: Multi-Person Pose Forecasting with Transformers. [https://arxiv.org/abs/2208.14023v1](https://arxiv.org/abs/2208.14023v1) \[3\] Muhammed Kocabas, Nikos Athanasiou and Michael J. Black. VIBE: Video Inference for Human Body Pose and Shape Estimation. [https://arxiv.org/pdf/1912.05656.pdf](https://arxiv.org/pdf/1912.05656.pdf) \[4\] Guangyao Zhai, Liang Liu, Linjian Zhang, Yong Liu. PoseConvGRU: A Monocular Approach for Visual Ego-motion Estimation by Learning [https://arxiv.org/abs/1906.08095](https://arxiv.org/abs/1906.08095) \[5\] Dirk Ormoneit, Michael J. Black, Trevor Hastie, and Hedvig Kjellstrom. Representing cyclic human motion using functional analysis. [https://www.sciencedirect.com/science/article/abs/pii/S0262885605001526](https://www.sciencedirect.com/science/article/abs/pii/S0262885605001526) \[6\] Raquel Urtasun, David J. Fleet, and Neil D. Lawrence. Modeling Human Locomotion with Topologically Constrained Latent Variable Models. [https://link.springer.com/chapter/10.1007/978-3-540-75703-0\_8](https://link.springer.com/chapter/10.1007/978-3-540-75703-0_8) \[7\] Hyemin Ahn, Timothy Ha, Yunho Choi, Hwiyeon Yoo, and Songhwai Oh. Text2Action: Generative Adversarial Synthesis from Language to Action. [https://arxiv.org/pdf/1710.05298.pdf](https://arxiv.org/pdf/1710.05298.pdf) \[8\] Chaitanya Ahuja and Louis-Philippe Morency. Language2Pose: Natural Language Grounded Pose Forecasting. [https://arxiv.org/abs/1907.01108](https://arxiv.org/abs/1907.01108) \[9\] Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. MotionCLIP: Exposing Human Motion Generation to CLIP Space. [https://arxiv.org/abs/2203.08063](https://arxiv.org/abs/2203.08063) \[10\] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning Transferable Visual Models from Natural Language Supervision. [https://arxiv.org/pdf/2103.00020.pdf](https://arxiv.org/pdf/2103.00020.pdf) \[11\] Mathis Petrovich, Michael J. Black, and Gul Varol. Action-Conditioned 3D Human Motion Synthesis with Transformer VAE. [https://arxiv.org/pdf/2104.05670.pdf](https://arxiv.org/pdf/2104.05670.pdf) \[12\] Mathis Petrovich, Michael J. Black, and Gul Varol. TEMOS: Generating diverse human motions from textual descriptions. [https://arxiv.org/abs/2204.14109](https://arxiv.org/abs/2204.14109) \[13\] Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu , Xi Shen. T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations. [https://arxiv.org/pdf/2301.06052.pdf](https://arxiv.org/pdf/2301.06052.pdf) \[14\] Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Amit H Bermano, and Daniel Cohen-Or. Human motion diffusion model. [https://arxiv.org/abs/2209.14916](https://arxiv.org/abs/2209.14916) \[15\] Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model. [https://arxiv.org/abs/2208.15001](https://arxiv.org/abs/2208.15001) \[16\] Will Williams, Sam Ringer, Tom Ash, David MacLeod, Jamie Dougherty, and John Hughes. Hierarchical Quantized Autoencoders. [https://arxiv.org/pdf/2002.08111.pdf](https://arxiv.org/pdf/2002.08111.pdf) \[17\] Matthias Plappert, Christian Mandery, and Tamim Asfour. The KIT Motion-Language Dataset. [https://arxiv.org/abs/1607.03827](https://arxiv.org/abs/1607.03827) \[18\] Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating Diverse and Natural 3D Human Motions from Text. [https://openaccess.thecvf.com/content/CVPR2022/papers/Guo\_Generating\_Diverse\_and\_Natural\_3D\_Human\_Motions\_From\_Text\_CVPR\_2022\_paper.pdf](https://openaccess.thecvf.com/content/CVPR2022/papers/Guo_Generating_Diverse_and_Natural_3D_Human_Motions_From_Text_CVPR_2022_paper.pdf) [GPT](https://voxel51.com/blog/tag/gpt) [human motion synthesis](https://voxel51.com/blog/tag/human-motion-synthesis) [mocap](https://voxel51.com/blog/tag/mocap) [motion capture](https://voxel51.com/blog/tag/motion-capture) [motion estimation](https://voxel51.com/blog/tag/motion-estimation) [motion prediction](https://voxel51.com/blog/tag/motion-prediction) [T2M-GPT](https://voxel51.com/blog/tag/t2m-gpt) [VQ-VAE](https://voxel51.com/blog/tag/vq-vae) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/99dc871d875865fc200f14931a3cfb3118ded624-1020x1007.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Tunnel vision in computer vision: can ChatGPT see?\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 16, 2022](https://voxel51.com/blog/tunnel-vision-in-computer-vision-can-chatgpt-see) [![](https://cdn.sanity.io/images/h6toihm1/production/a735267ad7effa9f799f850ab7c8ffa241088710-1024x1024.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Why 2022 was the most exciting year in computer vision history (so far)\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Dec 14, 2022](https://voxel51.com/blog/why-2022-was-the-most-exciting-year-in-computer-vision-history-so-far) [![](https://cdn.sanity.io/images/h6toihm1/production/e6ca14f73ab3cebac9a67e469fc0cb4f87f7b08b-960x640.jpg?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Making of Avatar: The Way of Water\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jan 19, 2023](https://voxel51.com/blog/the-making-of-avatar-the-way-of-water) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-281-lllmstxt|> ## Model Selection with FiftyOne [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tutorials](https://voxel51.com/blog/category/tutorials) The ML Menu for Model Selection: Hugging Face, Weights & Biases, and FiftyOne May 1, 2023 • 12 min read Article content In this article [What’s on the menu?](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#23fb4e094290) [Setup](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#8d85ca8dbbfe) [Preparing the dataset](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#1d7b1d33e69a) [Model preparation](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#157d8df53a2f) [Integrate Weights & Biases with FiftyOne](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#7b03f84869e0) [Train](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#77c9e60c827c) [Evaluate results in FiftyOne](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#125c1601fd6c) [Perform a hyperparameter sweep](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#7347474ac977) [Summary](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#24605e220df0) [Join the FiftyOne community!](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#aa56bb7dc80c) In this article [What’s on the menu?](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#23fb4e094290) [Setup](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#8d85ca8dbbfe) [Preparing the dataset](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#1d7b1d33e69a) [Model preparation](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#157d8df53a2f) [Integrate Weights & Biases with FiftyOne](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#7b03f84869e0) [Train](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#77c9e60c827c) [Evaluate results in FiftyOne](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#125c1601fd6c) [Perform a hyperparameter sweep](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#7347474ac977) [Summary](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#24605e220df0) [Join the FiftyOne community!](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone#aa56bb7dc80c) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _Visualize and evaluate Hugging Face models on your dataset with Weights & Biases and FiftyOne_ The development of a machine learning solution is no longer a problem of simply designing the right model for the job because the performance of your model is also bounded by the quality of your dataset. In this era of huge data, it is necessary to co-develop a high quality dataset alongside a capable model, iteratively improving your model and dataset together. Two challenges that you will encounter are 1) keeping track of how your models perform and evolve over time as your dataset and model training scheme change and 2) managing your datasets and being able to visualize and explore them as they grow ever larger. Luckily, there are plenty of tools at your disposal that can help you co-develop high-quality datasets and models. The leading solutions to these two problems, specifically, are [Weights & Biases](https://wandb.ai/site) (W&B) for model and experiment tracking and [FiftyOne](https://voxel51.com/fiftyone/) for dataset visualization and management. In this post, we will also be adding [Hugging Face](https://huggingface.co/models) into the mix to quickly access a model to finetune. ![](https://cdn.sanity.io/images/h6toihm1/production/878ad4ffe165325bdff6478982b74c1ce9f9aa16-560x155.png?auto=format&dpr=2&fit=max&q=75&w=560) In this post, we will show you how to visualize and evaluate Hugging Face models on your dataset by integrating both Weights & Biases and FiftyOne. Specifically, we’ll cover how to: - Curate a custom dataset with FiftyOne - Pull in a popular model from Hugging Face - Integrate FiftyOne and Weights & Biases in your model training loop - Visualize model predictions in FiftyOne - Sweep over hyperparameters with Weights & Biases - Find the best model with Weights & Biases and FiftyOne Specifically, we'll curate a dataset for food detection from the COCO and Open Images datasets, and track the finetuning of a DETR model. This integration between W&B and FiftyOne means that you can browse W&B to view your high-level model evaluation results and dig into the corresponding model predictions and evaluations on your actual samples in FiftyOne. This is critical to find interesting success and failure modes of your model to let you know how to improve the quality of your dataset and, in turn, the performance of your model. ## What’s on the menu? **Weights & Biases:** [Weights & Biases](https://wandb.ai/site) is a developer-first MLOps platform that integrates into your model training loop to track your experiments and your model results. It also makes it easy to scale your experimentation and run large hyperparameter sweeps to optimize your model weights and manage the lifecycle of your model. **FiftyOne:** [FiftyOne](https://voxel51.com/fiftyone/) is the open source toolkit for building high-quality datasets and computer vision models. It's designed to allow fast and effective analysis during dataset and model co-development enabling you to iterate through experiments rapidly. It has a powerful but easy-to-use App and Python SDK letting you curate your custom datasets, visualize and explore them, and integrate them into other computer vision workflows. (Disclaimer: I work at Voxel51 and built portions of FiftyOne.) **Hugging Face:** [Hugging Face](https://huggingface.co/) is focused on letting the AI community build, train, and deploy state of the art models from throughout the open source machine learning community. Specifically, their model hub and transformers Python library make it easy to discover and extend open source machine learning models across a variety of different tasks. ## Setup If you haven’t already, install [FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html), the [Weights & Biases Python client](https://docs.wandb.ai/quickstart), and [Hugg](https://huggingface.co/docs/transformers/index) [ing Face's transformers package](https://huggingface.co/docs/transformers/index): ```bash 1pip install fiftyone wandb transformers ``` Next, you will need to [sign up for a free Weights & Biases account](https://wandb.ai/site) then find your API key [here](https://wandb.ai/authorize). The following will prompt you for this API key to connect the Python client to your account. ```python 1import wandb 2wandb.login() ``` ## Preparing the dataset In this walkthrough, let's pretend that we're working on the next viral recipe app that lets you take a picture of food and returns to you a recipe for how to cook it. We’ll focus on the food detection aspect. FiftyOne makes it easy to curate a training and evaluation dataset for our task. For dataset curation, we can use FiftyOne to [query for subsets of interest](https://docs.voxel51.com/tutorials/pandas_comparison.html), [find maximally unique data samples](https://docs.voxel51.com/user_guide/brain.html#finding-maximally-unique-images), [find annotation mistakes](https://docs.voxel51.com/user_guide/brain.html#label-mistakes), and more. To do this, we're going to need to train an object detection model to detect food in images. We'll be making use of the COCO-2017 and Open Images v7 datasets, both of which exist in the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html) and include annotations for various different types of food items like carrots, hot dogs, and cakes. To start, we need to put together lists of the classes of food objects we want to download from each dataset. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4coco_food = ['banana','apple','sandwich','orange','broccoli','carrot','hot dog','pizza','donut','cake'] 5openimages_food = ['Apple', 'Artichoke', 'Bagel', 'Baked goods', ..., 'Zucchini'] ``` See the full list of classes [here](https://gist.github.com/ehofesmann/92a5b80c549a9783dcf086acae9304a1). Now we can use the FiftyOne Dataset Zoo to download all samples with at least one instance of the listed classes from the validation splits of both the COCO and Open Images datasets. ```python 1coco_dataset = foz.load_zoo_dataset( 2    "coco-2017", 3    split="validation", 4    classes=coco_food, 5    dataset_name="coco-food", 6) 7 8openimages_dataset = foz.load_zoo_dataset( 9    "open-images-v7", 10    split="validation", 11    classes=openimages_food, 12    dataset_name="oi-food", 13    label_types="detections", 14) ``` Next, we want to merge these two datasets together with the [`merge_samples()`](https://docs.voxel51.com/api/fiftyone.core.dataset.html?highlight=merge_samples#fiftyone.core.dataset.Dataset.merge_samples) method. But first, there are a couple of processing steps to take, including making all of the classes the same case and filtering out all non-food classes. The good news is that preprocessing data is easy with FiftyOne because of its handy builtin methods like [`map_labels(),`](https://docs.voxel51.com/user_guide/using_views.html#transforming-fields) [`filter_labels(),`](https://docs.voxel51.com/user_guide/using_views.html#filtering) [`merge_samples()`](https://docs.voxel51.com/api/fiftyone.core.dataset.html?highlight=merge_samples#fiftyone.core.dataset.Dataset.merge_samples) [,](https://docs.voxel51.com/api/fiftyone.core.dataset.html?highlight=merge_samples#fiftyone.core.dataset.Dataset.merge_samples) and more. [See this gist](https://gist.github.com/ehofesmann/35e181991d413cd3a1c54349610aefe0) for these preprocessing and merging steps that make use of FiftyOne's SDK to write [views](https://docs.voxel51.com/user_guide/using_views.html) that query, filter, and mutate the dataset and metadata. ```python 1dataset = preprocess_and_merge_with_fiftyone( 2    coco_dataset, openimages_dataset, coco_food, openimages_food 3) ``` We also want to set the dataset to be [persistent](https://docs.voxel51.com/user_guide/using_datasets.html#dataset-persistence) so that we can easily load it with [fo.load\_dataset(dataset\_name)](https://docs.voxel51.com/user_guide/using_datasets.html#datasets) in the future. ```python 1dataset.name = "food" 2dataset.persistent = True ``` Before moving on, let's visualize and explore the dataset in the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html). ```python 1session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/372c5b0f37e9a3b8d7fa3ffcdd197d3248c19ca8-1697x1091.png?auto=format&dpr=2&fit=max&q=75&w=1600) The last processing step that needs to be done on the dataset is to generate training and validation splits. We can use FiftyOne to generate random 80/20 splits of the dataset, tagging samples as either `train` or `val`. ```python 1import fiftyone.utils.random as four 2 3four.random_split(dataset, {"train": 0.8, "val": 0.2}) 4train_view = dataset.match_tags("train") 5val_view = dataset.match_tags("val") ``` ## Model preparation The object detection model we'll be using is [Facebook Research's DETR](https://huggingface.co/facebook/detr-resnet-50) from the Hugging Face [transformers](https://github.com/huggingface/transformers) Python package. Some of the code in this section has been extended from [this tutorial](https://github.com/NielsRogge/Transformers-Tutorials/blob/master/DETR/Fine_tuning_DetrForObjectDetection_on_custom_dataset_(balloon).ipynb) on finetuning DETR. For the model preparation and training, we'll be sticking fairly closely to native PyTorch code so you can see the bare bones of how to train on a FiftyOne dataset and track your experiment with Weights & Biases. Much of the image and label preprocessing is able to be handled by the model-specific [processor class](https://huggingface.co/docs/transformers/model_doc/detr#transformers.DetrImageProcessor) made available by Hugging Face. ```python 1from transformers import DetrImageProcessor 2 3processor = DetrImageProcessor.from_pretrained("facebook/detr-resnet-50") ``` We then need to create the PyTorch dataset class and data loaders that return images and labels in the format expected by DETR. [See this gist](https://gist.github.com/ehofesmann/a1e73d02941463b554c7f198ecf4488a) for the implementation of these classes in a way that makes use of the FiftyOne dataset we’ve curated. ```python 1train_dataloader, val_dataloader = create_data_loaders(train_view, val_view, processor) ``` Finally, let's load [the pretrained DETR model](https://huggingface.co/facebook/detr-resnet-50) (with a ResNet-50 backbone) from Hugging Face. Since we're training on a custom dataset, we need to specify the number of labels on which we will be finetuning. We can use a FiftyOne [aggregation](https://docs.voxel51.com/user_guide/using_aggregations.html#distinct-values) to quickly find the number of distinct label classes across our ground truth labels in our dataset. ```python 1from transformers import DetrConfig, DetrForObjectDetection 2import torch 3 4num_labels = len(dataset.distinct("ground_truth.detections.label")) 5 6model = DetrForObjectDetection.from_pretrained( 7    "facebook/detr-resnet-50", 8     revision="no_timm", 9     num_labels=num_labels, 10     ignore_mismatched_sizes=True, 11) ``` Note: If you have sufficient training data, you can also attempt to train this model from scratch [as shown here](https://huggingface.co/docs/transformers/v4.27.2/en/model_doc/detr#transformers.DetrConfig.example). ## Integrate Weights & Biases with FiftyOne Now to bring Weights & Biases into the mix! We'll be using W&B to track hyperparameters of our experiment and monitor the training and validation losses throughout the training process. Later, we'll also show how to train and track multiple models with a W&B sweep. Integrating a W&B run with FiftyOne involves creating a link between the FiftyOne dataset and field containing model predictions that is associated with a given W&B training run. There are numerous ways that this could be done, here we show one approach that adds direct links to and from FiftyOne and W&B. W&B allows you to organize your experiments under _projects_, with each experiment constituting a _run_. We will track model and dataset metadata like training hyperparameters and information about the FiftyOne dataset we're training within the _run_ configuration. ```python 1custom_id = "51" 2# Start a new run to track this script 3wandb.init( 4    project="food-finder", 5    name=f"training_run_{custom_id}", 6    config={ 7        "epochs": 300, 8        "lr": 1e-4, 9        "lr_backbone": 1e-5, 10        "weight_decay": 1e-4, 11        "architecture": "DETR", 12        "fiftyone_dataset": dataset.name, 13        "fiftyone_train_split_tag": "train", 14        "fiftyone_val_split_tag": "val", 15        "field_id": custom_id, 16    } 17) ``` Now to actually create the integration between a W&B run and a FiftyOne dataset. For the W&B to FiftyOne direction, we'll be storing the W&B run and project urls on [the field](https://docs.voxel51.com/user_guide/using_datasets.html#storing-info) of our FiftyOne dataset which will contain the predictions from the resulting model of that run. These links are then clickable in the sidebar of the FiftyOne App when hovering over that field name. For the FiftyOne to W&B direction, we'll store a [link on our W&B run](https://docs.wandb.ai/guides/track/log) to the URL of the specific FiftyOne view into our dataset which contains only the ground truth and prediction label fields for that run. (This can all be customized to link to whatever views you want). ```python 1def integrate_wandb_fo(dataset, pred_field, gt_field): 2 3    # Link W&B project within FiftyOne dataset 4    project_url = wandb.run.get_project_url() 5    wandb_url = "%s/runs/%s" % (project_url, wandb.run.id) 6    if pred_field not in dataset.get_field_schema(): 7        dataset.add_sample_field( 8            pred_field, 9            ftype=fo.EmbeddedDocumentField, 10            embedded_doc_type=fo.Detections, 11        ) 12 13    field = dataset.get_field(pred_field) 14    field.info = { 15        "WandB_run_url": wandb_url, 16        "WandB_project_url": project_url, 17    } 18    field.save() 19 20    # Create a view into our FiftyOne dataset containing the exact predictions of this run and the ground truth labels 21    view = dataset.select_fields([gt_field, pred_field]) 22    view_name = f"wandb_run_{wandb.config.field_id}" 23    if not dataset.has_saved_view(view_name): 24        dataset.save_view(view_name, view) 25 26    # Link FiftyOne view within W&B run 27    wandb.log( 28        { 29            "FiftyOne_URL": wandb.Html( 30                'View Dataset in FiftyOne' % (dataset.name, view_name.replace("_", "-")) 31            ) 32        } 33    ) 34 35pred_field = f"predictions_{wandb.config.field_id}" 36gt_field = "ground_truth" 37integrate_wandb_fo(dataset, pred_field, gt_field) ``` Note that since the FiftyOne App runs in a local Python session on your machine accessible at localhost:5151, you'll need to have an instance of the FiftyOne App running whenever you want to click these links and view your dataset. You can run the following in a terminal window to start up a FiftyOne App process: ```bash 1fiftyone app launch -r --wait -1 ``` (With [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) you can just link directly to your deployed FiftyOne Teams App URL.) Below is a sneak peek of the links this method creates between W&B and FiftyOne later in this post. ![](https://cdn.sanity.io/images/h6toihm1/production/a70c92d151efabb9f708d75179cc605b0011ef2d-872x499.png?auto=format&dpr=2&fit=max&q=75&w=872) ## Train Now we're ready to put everything together and get this model trained. In this example, we're just using native PyTorch code to write our training loop, but this can be extended to use your preferred model training framework by following the same principles. The meat of the training code is in this train() method. Here, we set up an optimizer and iteratively train/validate one epoch at a time on our FiftyOne-backed dataloader. Each iteration, we then track whatever hyperparameters are of interest to us. In this example, we'll track the epoch, learning rate, and train/validation losses. Additionally, we save the model weights every 25 epochs throughout the training process, but this can be configured to your needs. ```python 1from tqdm import tqdm 2import os 3 4def train( 5 model, 6 train_dataloader, 7 val_dataloader, 8 output_folder="output", 9 gradient_clip_val=0.1, 10 save_epoch=25, 11): 12    optimizer = setup_optimizer( 13 model, 14 wandb.config.lr, 15 wandb.config.lr_backbone, 16 wandb.config.weight_decay, 17 ) 18    device = torch.device("cuda:0" if torch.cuda.is_available() else "cpu") 19    model.to(device) 20 21    for epoch in range(wandb.config.epochs): 22        print(epoch) 23        train_loss = train_one_epoch( 24 model, 25 train_dataloader, 26 device, 27 optimizer, 28 gradient_clip_val=gradient_clip_val, 29 ) 30        val_loss = validate_one_epoch(model, val_dataloader, device) 31 32        # Log training metrics in W&B 33        wandb.log({ 34 "epoch": epoch, 35 "lr": optimizer.defaults["lr"], 36 "train_loss": train_loss, 37 "val_loss": val_loss 38 }) 39        if epoch % save_epoch == 0: 40            model.save_pretrained( 41                os.path.join(output_folder,str(epoch)) 42            ) 43 44    model.save_pretrained(os.path.join(output_folder, "final")) ``` [This gist](https://gist.github.com/ehofesmann/15137c935472e59685d05b83f8b4e562) provides some fairly boilerplate code to implement the methods used above. The batch processing performed here is specific to DETR and follows [this finetuning tutorial](https://github.com/NielsRogge/Transformers-Tutorials/blob/master/DETR/Fine_tuning_DetrForObjectDetection_on_custom_dataset_(balloon).ipynb). You can find the finetuning steps for other Hugging Face models [here](https://github.com/NielsRogge/Transformers-Tutorials/tree/master). Time to start training! For reference, 20 epochs with this setup took roughly 4 hours on an NVIDIA TITAN V GPU. ```python 1train(model, train_dataloader, val_dataloader) ``` We can also log the resulting model weights as a [W&B artifact](https://docs.wandb.ai/guides/artifacts). Between this and the FiftyOne dataset, we can ensure that it will be easy to reproduce our results in the future. ```python 1art = wandb.Artifact(f'food-detector-{wandb.config.field_id}', type="model") 2art.add_dir("output/final") 3wandb.log_artifact(art) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/fec7b23838b23cb94088f2ba8618011646ca544c-1203x528.png?auto=format&dpr=2&fit=max&q=75&w=1203) ## Evaluate results in FiftyOne FiftyOne not only makes it easy to curate your training dataset, but also to evaluate the results of your model predictions. The flexibility of FiftyOne's data model means that you can add however many custom fields that you want to your dataset, in this case predictions from all of our runs. To start, we need to run inference on the samples of the validation set, then convert the model outputs from the bounding box format of DETR to the [bounding box format](https://docs.voxel51.com/user_guide/using_datasets.html#object-detection) expected by FiftyOne in the form of fo.Detection objects. This is a fairly straightforward conversion of restructuring the bounding box coordinates and putting the labels and confidences in the right spots. At the end, all of the model's detections are then stored in a new field on our FiftyOne dataset. [See this gist](https://gist.github.com/ehofesmann/4cc89eef9a043a3720489738d6c63881) for an implementation of this DETR to FiftyOne conversion function, `add_detections()`. ```python 1add_detections(model, processor, val_view, pred_field) ``` With the inference results added to our dataset, we can now make use of FiftyOne's model evaluation capabilities. For common label types, there are standard evaluation practices available in the FiftyOne SDK. For example, [`fo.evaluate_detections()`](https://docs.voxel51.com/user_guide/evaluation.html#detections) will perform either COCO-style or Open Images-style object detection evaluation comparing your ground truth detections with your model predicted detections. You can use this to compute the same mAP as with pycocotools, but the primary benefit is that this will also keep the individual label-level results around. For example, we’ll know if a prediction was a false positive or a false negative. This becomes invaluable to dig into the dataset and query for specific edge cases of interest, like figuring out “which vegetable is my model worst at detecting”, or “how often does my model miss detecting pizza slices”. ```python 1eval_key = f"eval_{wandb.config.field_id}" 2results = fo.evaluate_detections( 3    val_view, 4    pred_field, 5    gt_field=gt_field, 6    eval_key=eval_key, 7    compute_mAP=True, 8) 9 10print(results.mAP()) 11# 13.99 ``` Using the evaluation results, we can write a query with [FiftyOne's view expression](https://docs.voxel51.com/user_guide/using_views.html#filtering) [language](https://docs.voxel51.com/user_guide/using_views.html#filtering) to find all false positive predictions with a high confidence. These will generally be interesting examples since this is where the model was confident in its prediction, but got it wrong. This often lets us find hard samples, annotation mistakes, or discrepancies between the training and validation splits. ```python 1from fiftyone import ViewField as F 2 3high_conf_fp_view = val_view.filter_labels( 4    pred_field, (F(eval_key)=="fp") and (F(confidence) > 0.7) 5) 6 7session.view = high_conf_fp_view ``` From these results, we can see that even though the mAP is low, the model still performs fairly well and the low score is more due to the evaluation protocol and ground truth labels. Specifically, there are numerous cases where the model produces technically correct predictions which happen to not match the ground truth annotations. ![](https://cdn.sanity.io/images/h6toihm1/production/a92a7366fc91c4f31c0883f77d0add2ff4abadb8-1841x963.png?auto=format&dpr=2&fit=max&q=75&w=1600) In this example, you can see that there are annotation mistakes in the oranges that are missing from the ground truth, as well as a “Baked goods” annotation which may be correct, but may also be missing one of the “Snack”/”Fast food” ground truth labels for these food items. In practice, we may want to use this knowledge to take a pass over our label ontology here to ensure that all of the potential classes for an object are represented in the ground truth annotations. > This is precisely why you should not select the best model based off of aggregate scores like mAP alone, you must visualize and explore the predictions of your model on your dataset itself to build an intuition and trust in the performance of your dataset, and its shortcomings. ## Perform a hyperparameter sweep Combining everything we've learned so far, we can now consolidate it into a [Weights & Biases sweep](https://docs.wandb.ai/guides/sweeps). When training a model, there are often so many hyperparameters and options to choose, that you need to iterate over them to find combinations that result in the best model for your task. W&B makes it really easy to define your training function and configure a sweep over the hyperparameters of interest to you. The following code snippet uses methods defined in [this gist](https://gist.github.com/ehofesmann/d3deae192825c9b2dd9e1d686ca9291c) and follows the exact same workflow outlined in the previous sections. ```python 1def sweep_main(): 2    dataset = fo.load_dataset("food") 3    train_view = dataset.match_tags("train") 4    val_view = dataset.match_tags("val") 5 6    processor = DetrImageProcessor.from_pretrained( 7        "facebook/detr-resnet-50" 8    ) 9    train_dataloader, val_dataloader = create_dataloaders( 10        train_view, val_view, processor 11    ) 12 13    pred_field = f"predictions_{custom_id}" 14    gt_field = "ground_truth" 15 16    num_labels = len( 17        dataset.distinct(f"{gt_field}.detections.label") 18    ) 19    model = DetrForObjectDetection.from_pretrained( 20        "facebook/detr-resnet-50", 21         revision="no_timm", 22         num_labels=num_labels, 23         ignore_mismatched_sizes=True, 24    ) 25 26    custom_id = get_run_id(dataset) 27    init_wandb_run(dataset, custom_id) 28 29    integrate_wandb_fo(dataset, pred_field, gt_field) 30 31    output_folder = f"sweep_results/{custom_id}" 32    train( 33        model, 34        train_dataloader, 35        val_dataloader, 36        output_folder=output_folder, 37    ) 38    log_artifact(output_folder, custom_id) 39 40    add_detections(model, processor, val_view, pred_field) 41 42    eval_key = f"eval_{custom_id}" 43    results = fo.evaluate_detections( 44      val_view, 45      pred_field, 46      eval_key=eval_key, 47      compute_mAP=True, 48    ) 49 50    save_view(dataset, pred_field, gt_field, eval_key, custom_id) ``` Now let’s set up and run our sweep configuration to cover five different combinations of learning rate, backbone learning rate, and weight decay parameters to find which result in the best performance. ```python 1sweep_configuration = { 2    'method': 'random', 3    'metric': {'goal': 'minimize', 'name': 'val_loss'}, 4    'parameters': 5    { 6        'lr': {'max': 1e-3, 'min': 1e-5}, 7        'lr_backbone': {'values': [1e-4, 1e-5, 1e-6]}, 8        'weight_decay': {'values': [1e-3, 1e-4, 1e-5]}, 9     } 10} 11 12num_runs = 5 13sweep_id = wandb.sweep( 14    sweep=sweep_configuration, 15    project='fiftyone-sweep', 16) 17 18wandb.agent(sweep_id, function=sweep_main, count=num_runs) ``` ### Explore results Once our sweep is complete, we can view the results in W&B and FiftyOne. ![](https://cdn.sanity.io/images/h6toihm1/production/2e48837e2fa9ea579826cf2d7181abaf427d21dc-1815x963.png?auto=format&dpr=2&fit=max&q=75&w=1600) From the W&B dashboard, we can see that there is one run where the validation loss dropped significantly more and faster than the other runs. We can click on that run to see the metrics in more detail. ![](https://cdn.sanity.io/images/h6toihm1/production/e5bc47b5671632458a8df63e1a8099c35e12689a-1831x910.png?auto=format&dpr=2&fit=max&q=75&w=1600) From that run's page, we can also click the "View Dataset in FiftyOne" link to open a new browser window to see exactly those model predictions and their evaluation results in FiftyOne. (Assuming the FiftyOne App is running as described in the "Link Weights & Biases with FiftyOne" section above.) From the FiftyOne App, we can now dig in and explore the results of this model. For example, we can perform the same high-confidence false-positive query from before but directly through buttons in the App rather than code, giving you even more flexibility to work how you’d like to. ![](https://cdn.sanity.io/images/h6toihm1/production/a70c92d151efabb9f708d75179cc605b0011ef2d-872x499.png?auto=format&dpr=2&fit=max&q=75&w=872) We can then click on the link in the field's info tooltip to go back to the W&B run or project to explore our other experiments in W&B. ## Summary Weights & Biases is a leading solution for experiment tracking and model management for good reason. Its feature rich toolset allows you to easily integrate it into your model training pipelines and start building dashboards of your experimental results with ease. Integrating Weights & Biases with FiftyOne adds easy dataset management and model evaluation to the equation to round out your MLOps stack and let you start co-developing better datasets and models, faster! ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,500+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,800+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,700+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 274+ repositories - 59+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Hugging Face](https://voxel51.com/blog/tag/hugging-face) [hyperparameter sweep](https://voxel51.com/blog/tag/hyperparameter-sweep) [model evaluation](https://voxel51.com/blog/tag/model-evaluation) [model selection](https://voxel51.com/blog/tag/model-selection) [model training](https://voxel51.com/blog/tag/model-training) [object detection](https://voxel51.com/blog/tag/object-detection) [Weights & Biases](https://voxel51.com/blog/tag/weights-biases) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/7a4d8c6623f69b1f138e3548bb11bd805e5aa323-924x638.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ How to Train Your Dragon (Detector)\\ \\ Tutorials\\ \\ • \\ \\ Feb 4, 2022](https://voxel51.com/blog/how-to-train-your-dragon-detector) [![](https://cdn.sanity.io/images/h6toihm1/production/98e839c6e81bb9c4fa3ad96bf0d5d1b77ee11f6c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 1)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1) [![](https://cdn.sanity.io/images/h6toihm1/production/de34ad70ef11fe3d7daf5b3c8ff6214e76c85522-2560x1440.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 2)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-282-lllmstxt|> ## Computer Vision Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup – April 27, 2023 Apr 28, 2023 • 5 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#87c6b56ad521) [Leveraging Attention for Improved Accuracy and Robustness](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#d0f3ac622ca3) [Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#0319739a7f16) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#263026dca69e) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#f50b28511bf3) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#ac36eac269a2) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#87c6b56ad521) [Leveraging Attention for Improved Accuracy and Robustness](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#d0f3ac622ca3) [Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#0319739a7f16) [Computer Vision Meetup Locations](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#263026dca69e) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#f50b28511bf3) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023#ac36eac269a2) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We just wrapped up the April 27, 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/), and if you missed it or want to revisit it – here's a recap! In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes again this month was Wildlife AI! We were first introduced to Wildlife AI through the FiftyOne community. They are using FiftyOne to enable their users to easily analyze the camera data and create their own models. We are sending this month’s charitable donation of $200 to Wildlife AI on behalf of the computer vision community. ![](https://cdn.sanity.io/images/h6toihm1/production/e80959f169a6a89e3cef73f9bb1bdd1750499fe7-770x146.png?auto=format&dpr=2&fit=max&q=75&w=770) Missed the Meetup? No problem. Here are playbacks and talk abstracts from the event. ## Leveraging Attention for Improved Accuracy and Robustness https://www.youtube.com/watch?v=QYXIHIvesYo In recent years, the naturally interpretable attention mechanism has become one of the most common building blocks of neural networks, allowing us to produce explanations intuitively and easily. However, the applications of such explanations beyond the scope of accountability and interpretability remain limited. In this talk, Hila presents her latest research on leveraging attention to significantly improve the accuracy and robustness of state-of-the-art large neural networks with limited resources. This is achieved by directly manipulating the attention maps based on intuitive objectives and can be applied to a variety of tasks ranging from object classification to image generation. [Hila Chefer](https://www.linkedin.com/in/hila-chefer/) is a PhD student and lecturer at Tel-Aviv University, and an intern at Google research. Her research focuses on constructing faithful explainable AI algorithms, and leveraging explanations to promote model accuracy and robustness. Q&A from the talk included: - Is the attention matrix specific to one self-attention head? Or is it aggregated in some manner across all heads? - In ViT, is the attention score (q,k,v) for a patch initialized from some unsupervised dataset? - Could you give us some intuition on how to create a differentiable loss function that can help to differentiate between foregrounds and backgrounds using vision transformers? - Why do stopwords have “strong” maps since they might not be affecting the output? - If we change the question to "A crown on the lion", will the error be the same? - What are some real-world scenarios where the attention mechanism could be used? You can jump straight to the Q&A [here](https://youtu.be/QYXIHIvesYo?t=1102) and [here.](https://youtu.be/QYXIHIvesYo?t=2737) ## Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO https://www.youtube.com/watch?v=rIz841UeudQ In this workshop, you will learn how to use less data for performant AI models. You will see this in action through real-world computer vision implementations, such as object detection and [anomaly detection](https://github.com/openvinotoolkit/anomalib) use cases, optimization processes, and deployment at the edge. And, you will learn how the open source [OpenVINO](https://github.com/openvinotoolkit/openvino) toolkit can help reduce the gap between theoretical models and real-world implementations. - Train & optimize for the edge with less data - Improve the performance of your model regardless of hardware - Learn how OpenVINO can accelerate AI models [Zhuo Wu](https://www.linkedin.com/in/wuzhuo/) is an AI software evangelist at Intel focusing on the OpenVINO toolkit. Her work ranges from deep learning technologies to 5G wireless communication technologies. She has delivered end2end machine learning and deep learning based solutions to business customers in different industries. Q&A from the talk included: - Are the detections 360 degree or front face? - Which detectors does Anomalib use? - Is OpenVINO similar to OpenCV? - Does OpenVINO support training and inference on CUDA? - Is it possible to fine-tune the model to detect different levels of anomalies? - Are we able to deploy Anomalib to tablet/mobile devices? - For anomaly detection, must the training and testing of the images have to use the same camera angle and lighting conditions? - Does Anomalib work with synthetic data generation? You can jump straight to the Q&A [here](https://youtu.be/rIz841UeudQ?t=2114). ## **Computer Vision Meetup Locations** Computer Vision Meetup membership has grown to over [3,800 members](https://www.meetup.com/pro/computer-vision-meetups/) in just under a year! The goal of the Meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. Join one of the 13 Meetup locations closest to your timezone. - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) ## What’s Next? We have exciting speakers already signed up over the next few months! Become a member of the [Computer Vision Meetup closest to you](https://www.meetup.com/pro/computer-vision-meetups/), then register for the Zoom. Up next on May 11 at 10 AM Pacific we have the US and EU-timezone-friendly Computer Vision Meetup happening with talks including: - **The Role of Symmetry in Human and Computer Vision**– Sven Dickinson _(University of Toronto & Samsung)_ - **Machine Learning for Fast, Motion-Robust MRI**– Nalini Singh _(MIT)_ Register for the Zoom [here](https://voxel51.com/computer-vision-events/may-2023-computer-vision-meetup/?utm_source=blog). You can find a complete schedule of upcoming Meetups on [the Voxel51 Events page](https://voxel51.com/computer-vision-events/). ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. — _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [attention](https://voxel51.com/blog/tag/attention) [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [Edge AI](https://voxel51.com/blog/tag/edge-ai) [OpenVINO](https://voxel51.com/blog/tag/openvino) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/e56436d38d294978c25356ba4c5482b53a29afb3-960x540.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup – February 2023\\ \\ Event Recaps\\ \\ • \\ \\ Feb 14, 2023](https://voxel51.com/blog/computer-vision-meetup-feb-2023-recap) [![](https://cdn.sanity.io/images/h6toihm1/production/73db90a7323e7eb6feb0dad3839e11c6d4ab525b-960x540.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — May 11, 2023\\ \\ Event Recaps\\ \\ • \\ \\ May 12, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/2260b433f7da8171c8f20164bf89e715d04cb7a6-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — June 8, 2023\\ \\ Event Recaps\\ \\ • \\ \\ Jun 9, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-283-lllmstxt|> ## Visualizing Amazon ARMBench Dataset [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets) Visualizing Defects in Amazon’s ARMBench Dataset Using Embeddings and OpenAI’s CLIP Model May 4, 2023 • 12 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#02f86f809b53) [What are “pick and place” robots?](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#438a85bbb023) [About the dataset](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#5b8f940d8489) [Defect detection](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#ba01de1e2dd7) [Object identification](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#86e7ec15af54) [Object segmentation](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#cb8ecca1b7d1) [Dataset quick facts](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#4c436f62bc94) [Step 1: Download the dataset](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#1e4d8ac9253d) [Step 2: Install FiftyOne](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#90b328f53f4c) [Step 3: Import the dataset](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#3754b93d7422) [Step 4: Launch the FiftyOne App to visualize the dataset](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#e9074e6057c2) [Tags](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#363da90ba6eb) [Histograms](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#cbc030a062dd) [Sample details](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#ba1194f3c620) [Embeddings and the FiftyOne Brain](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#f56c59ee6caa) [Start working with the dataset](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#67ca80ade5ca) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#02f86f809b53) [What are “pick and place” robots?](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#438a85bbb023) [About the dataset](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#5b8f940d8489) [Defect detection](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#ba01de1e2dd7) [Object identification](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#86e7ec15af54) [Object segmentation](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#cb8ecca1b7d1) [Dataset quick facts](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#4c436f62bc94) [Step 1: Download the dataset](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#1e4d8ac9253d) [Step 2: Install FiftyOne](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#90b328f53f4c) [Step 3: Import the dataset](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#3754b93d7422) [Step 4: Launch the FiftyOne App to visualize the dataset](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#e9074e6057c2) [Tags](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#363da90ba6eb) [Histograms](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#cbc030a062dd) [Sample details](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#ba1194f3c620) [Embeddings and the FiftyOne Brain](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#f56c59ee6caa) [Start working with the dataset](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip#67ca80ade5ca) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the latest installment of our ongoing blog series where we explore computer vision related datasets. In this post we’ll use the open source FiftyOne toolset to visualize Amazon’s [recently released](https://www.amazon.science/blog/amazon-releases-largest-dataset-for-training-pick-and-place-robots) dataset for training “pick and place” robots, plus we’ll create embeddings with the OpenAI CLIP model to explore defects. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source computer vision toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. ![](https://cdn.sanity.io/images/h6toihm1/production/136887bcc1e07d86237d0f79c0f2bf731805f6fb-1280x720.gif?auto=format&dpr=2&fit=max&q=75&w=1280) ## What are “pick and place” robots? These days, pick and place robots are becoming more and more common in manufacturing and logistics environments. Leveraging robots that don’t require downtime (besides regularly scheduled maintenance) can boost overall production efficiency and free up humans to work on safer, less repetitive tasks. In some cases, humans and robots work side-by-side, leveraging the strengths of each. Pick and place robots come in a variety of forms, with the most common type likely being the 5- or 6-axis articulated arm. There are also more specialized robots that can pick groups of items and place them in specific positions, robots designed to work at very high speeds, and the aforementioned collaborative robots or “cobots” that work in tandem with humans. What are the practical applications of pick and place robots? Although manufacturing is the most obvious use case, you can also find these types of robots performing packaging, sorting, or inspection tasks that need to be done with speed and accuracy. ![](https://cdn.sanity.io/images/h6toihm1/production/eecad5615085d1110540ad840ddabe12e5e646b1-1999x1125.jpg?auto=format&dpr=2&fit=max&q=75&w=1600) To learn more about the intersection between robots, manufacturing, and computer vision, check out: [How Computer Vision Is Changing Manufacturing in 2023](https://voxel51.com/blog/how-computer-vision-is-changing-manufacturing-in-2023/). ## About the dataset Earlier this month, Amazon publicly released the largest computer vision dataset ever captured in an industrial product-sorting setting. In contrast to previous datasets for robotic manipulation that might be limited either in the number of object types or in scene heterogeneity and realism, this new dataset, called [ARMBench](http://armbench.s3-website-us-east-1.amazonaws.com/) (Amazon Robotic Manipulation Benchmark), features more than 235,000 pick and place activities on 190,000 objects taking place in the context of an operating Amazon warehouse. This massive dataset can be used to train pick and place robots that are better able to generalize to new products and contexts. To learn more about the motivations behind the dataset, check out the paper, [ARMBench: An object-centric benchmark dataset for robotic manipulation](https://www.amazon.science/publications/armbench-an-object-centric-benchmark-dataset-for-robotic-manipulation) over on the Amazon Science blog. The basic scenario for ARMBench is one in which a robotic arm must retrieve a single item from a bin full of items and transfer it to a tray on a conveyor belt. According to the authors, “The variety of objects and their configurations and interactions in the context of the robotic system made for a uniquely challenging task.” The ARMBench dataset is broken down as follows: ![](https://cdn.sanity.io/images/h6toihm1/production/e156acc96caaf5357f8fde73000cd36f7265b2e0-1377x619.png?auto=format&dpr=2&fit=max&q=75&w=1377) ## Defect detection - _Image defect detection (66 GB):_ This dataset comprises 13,303 images of objects with defects taken through multiple view-points (Transfer-images). For image defect detection, multi-pick and package-defects are the two defect classes. 100,000 images of objects with no defects or are available in the dataset. _Multi-pick_ is used to describe activities where multiple objects were picked and transferred from the source container to the destination container. _Package-defect_ is used to describe activities where the object packaging opened and/or the object separated into multiple parts. Two subclasses, open and deconstruction, are defined for package-defect. - _Video defect detection (255 GB):_ This dataset comprises 4,075 videos of objects with defects. Multi-picks are not as observable in videos and are excluded from this dataset. At the same time, open and deconstruction defects are observable in videos and are annotated. 100,000 videos of activities that did not result in a defect are available in the dataset. ![](https://cdn.sanity.io/images/h6toihm1/production/f688e34644d19de3ab1fe0cbef5fc0daf60638ab-1089x654.png?auto=format&dpr=2&fit=max&q=75&w=1089) ## Object identification With Object identification the task is to identify an image segment as one of the objects within a database. In the pre-pick stage, identifying an object segment within the tote allows accessing any stored models or attributes of the object from past experience which can be used for manipulation planning purposes. In the post-pick stage, the ID has access to the segment of the object being manipulated both within the tote as well as when it is attached to the robotic arm. - _Picks:_ 235,000 pick activities with images of the picked object in tote and in robotic arm. - _Reference-images:_ Up to 6 images (1.jpg-6.jpg) corresponding to different product-ids. ![](https://cdn.sanity.io/images/h6toihm1/production/2a033b6a49a9efaa306f89d5535758becb679830-1999x1164.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Object segmentation With this dataset, instance segmentation is used to identify and define distinct objects that are stored in containers. The outcomes of instance segmentation can be used to provide information to subsequent robotic processes, such as the identification of objects and generation of grasping strategies. - _Mix-Object-Tote (14 GB):_ This subset consists of close-up images of mixed objects that are stored in either yellow or blue totes. Mix-Object-Tote comprises a total of 44,253 images of size 2448 by 2048 pixels and 467,225 annotations, with an average of 10.5 instances per tote. - _Zoomed-Out-Tote-Transfer-Set (1.5 GB):_ This subset includes mixed objects placed in a yellow tote that were captured with sensors positioned further away from the tote, under different lighting conditions. The dataset contains 5,837 images of size 2046 by 2046 pixels and 43,401 annotations, with an average of 7.5 objects per tote. - _Same-Object-Transfer-Set (3 GB):_ This subset consists of multiple same objects placed in close proximity within various storage units. The Same-Object-Transfer-Set comprises 3,323 images of size 2048 by 1500 pixels and 12,664 annotations, with an average of 3.8 objects per scene. ## Dataset quick facts - **Download:** Request a [download](http://armbench.s3-website-us-east-1.amazonaws.com/data.html) link from Amazon - **License:** Creative Commons - **Paper:** [ARMBench: An object-centric benchmark dataset for robotic manipulation](https://www.amazon.science/publications/armbench-an-object-centric-benchmark-dataset-for-robotic-manipulation) - **Authors:** Chaitanya Mitash, Fan Wang, Shiyang Lu, Vikedo Terhuja, Tyler Garaas, Felipe Polido, Manikantan Nambi Up next, let’s download the dataset, install FiftyOne, and import the dataset into the App so we can visualize it! https://www.youtube.com/watch?v=5FTd4aYWr-k ## Step 1: Download the dataset In order to load the ARMBench dataset into FiftyOne, you’ll need to [request a download link](http://armbench.s3-website-us-east-1.amazonaws.com/data.html) from Amazon. For the purposes of this blog, we’ll be focusing on the Image Defect Detection subset of data that is part of the larger ARMBench dataset. ## Step 2: Install FiftyOne ```python 1pip install fiftyone ``` https://www.youtube.com/watch?v=7mmH-ql\_-zg If you don’t already have FiftyOne installed on your laptop, it takes less than a minute! Learn more about how to [get up and running with FiftyOne](https://voxel51.com/docs/fiftyone/getting_started/install.html) in the Docs. ## Step 3: Import the dataset Now that you have the dataset downloaded and FiftyOne installed, let’s import the dataset, make it compatible with FiftyOne and launch the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html). ```python 1from os import path 2import glob 3import json 4 5import numpy as np 6import imagesize 7 8import fiftyone as fo 9 10# set to download path 11image_defect_root = '/PATH TO ARMBENCH DATASET GOES HERE' 12 13# maximum number of groups to load (set to None for entire dataset) 14max_groups = 100 15 16data_root = path.join(image_defect_root,'data') 17train_csv = path.join(image_defect_root,'train.csv') 18test_csv = path.join(image_defect_root,'test.csv') 19 20 21def readlines(f): 22 with open(f,'r') as fh: 23 lines = [line.strip() for line in fh] 24 return lines 25 26 27def load_json(f): 28 with open(f,'r') as fh: 29 j = json.load(fh) 30 return j ``` To get started, we import FiftyOne and set up some paths and utility functions. The images and annotations live in subfolders under `data_root`. The basic building block of a FiftyOne dataset is a [sample](https://docs.voxel51.com/user_guide/basics.html#samples) – in this case, an image along with its annotations and metadata. This dataset has additional structure, as each `data` subfolder contains multiple images (typically, four) taken of a single object. This structure lends itself naturally to using a [grouped dataset](https://docs.voxel51.com/user_guide/groups.html) in FiftyOne. The `max_groups` parameter limits the number of groups (or subfolders) of data that are imported, as this is a large dataset. ```python 1def parse_data_dir(data_dir): 2 """Parse one data directory (group) 3 4 Args: 5 data_dir: full path to a single data folder, eg /data/ 6 7 Returns: 8 id, 9 """ 10 11 id = path.basename(data_dir) 12 jpg_pat = path.join(data_dir,'*.jpg') 13 ims = sorted(glob.glob(jpg_pat)) 14 jsons = [path.splitext(x)[0]+'.json' for x in ims] 15 jsons = [load_json(x) for x in jsons] 16 17 imsbase = [path.basename(x) for x in ims] 18 imskey = [path.splitext(x)[0] for x in imsbase] 19 json_files = [path.join(data_dir,x+'.json') for x in imskey] 20 jsons = [load_json(x) for x in json_files] 21 22 for im, json in zip(ims,jsons): 23 imbase = path.basename(im) 24 imkey = path.splitext(imbase)[0] 25 assert json['id']==imbase or json['id']==imkey 26 assert imkey.startswith(id + '_') 27 slice = imkey[len(id)+1:] 28 29 imw,imh = imagesize.get(im) 30 new_info = { 31 'filepath': im, 32 'imw': imw, 33 'imh': imh, 34 'slice': slice, 35 } 36 json.update(new_info) 37 38 return id, jsons 39 40 41def parse_all_data_dirs(): 42 """Parse all data folders, up to max_groups 43 44 Returns: 45 list of (id,jsons) 46 """ 47 48 data_dirs = sorted(glob.glob(path.join(data_root,'*'))) 49 data_dirs = data_dirs[:max_groups] 50 data_dir_infos = [parse_data_dir(x) for x in data_dirs] 51 return data_dir_infos ``` `parse_data_dir` parses a single subfolder of `data`. Each image has a corresponding json file that contains a segmentation polygon of the object, as well as labels indicating whether the transfer was without defect, or if a defect occurred, the type and subtype of defect observed. We augment the dictionaries read from the annotation files with image metadata and a `slice` identifier, which just specifies the camera view (1-4) for that image. The image height and width will be important because the annotation polygons in this dataset are stored as absolute pixel values, whereas FiftyOne stores vertices [normalized to image dimensions](https://docs.voxel51.com/user_guide/using_datasets.html#polylines-and-polygons). We will perform the conversion when we create our samples. Let’s create our FiftyOne dataset! ```python 1train_set = set(readlines(train_csv)) 2test_set = set(readlines(test_csv)) 3data_dir_infos = parse_all_data_dirs() 4 5dataset = fo.Dataset('ARMBench-Image-Defect-Detection') 6dataset.persistent = True 7 8samples_all = [] 9 10for id, grp_info in data_dir_infos: 11 12 group = fo.Group() 13 14 for info in grp_info: 15 if id in train_set: 16 tags = ['train'] 17 elif id in test_set: 18 tags = ['test'] 19 else: 20 tags = [] 21 22 if info['label']: 23 tags.append(info['label']) 24 if info['sublabel']: 25 tags.append(info['sublabel']) 26 27 sample = fo.Sample(filepath=info['filepath'], 28 tags=tags, 29 group=group.element(info['slice'])) 30 31 imw = info['imw'] 32 imh = info['imh'] 33 poly_pts = info['polygon'] 34 if poly_pts: 35 poly_pts = np.array(poly_pts,dtype=np.float64) 36 poly_pts[:,0] /= imw 37 poly_pts[:,1] /= imh 38 polyline = fo.Polyline(points=[poly_pts.tolist()],filled=True) 39 detections = fo.Polylines(polylines=[polyline]).to_detections(frame_size=(imw,imh)) 40 sample['object'] = detections 41 42 samples_all.append(sample) 43 44dataset.add_samples(samples_all) ``` The basic recipe for loading a dataset into FiftyOne is simple: create a [Dataset](https://docs.voxel51.com/user_guide/using_datasets.html#datasets), then create and add [Samples](https://docs.voxel51.com/user_guide/using_datasets.html#samples)! Here we have groups as well, and we follow this [basic recipe](https://docs.voxel51.com/user_guide/groups.html#adding-samples) for adding samples to a grouped dataset. We’ll use FiftyOne [tags](https://docs.voxel51.com/user_guide/using_datasets.html#tags) to store our defect annotations, as well as membership in train and test splits. This makes it a snap to visualize and filter by these elements in the App. We have a couple options for our object polygons. We could represent these in FiftyOne as either [polylines](https://docs.voxel51.com/user_guide/using_datasets.html#polylines-and-polygons) or [instance segmentations](https://docs.voxel51.com/user_guide/using_datasets.html#instance-segmentations). The bounding box will come in handy later, so we use instance segmentations in the end. But to get there, we first load into a Polyline and then do a conversion, as Polylines easily accept the list-of-vertices format of these annotations. FiftyOne has support for a huge variety of [label types](https://docs.voxel51.com/user_guide/using_datasets.html#labels) and [dataset formats](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#supported-import-formats), making it a snap to work with all types of data in the way that works best for you. ## Step 4: Launch the FiftyOne App to visualize the dataset ```python 1session = fo.launch_app(dataset) ``` With our dataset created, let’s launch the FiftyOne App in a browser. You should see the following initial view of the `ARMBench-Image-Defect-Detection` dataset by default in the App: ![](https://cdn.sanity.io/images/h6toihm1/production/10e3b3c23c901dfc7059cd6e6742ccb11085c427-1920x1120.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Tags We’ve loaded our annotations as FiftyOne [tags](https://docs.voxel51.com/user_guide/app.html#tags-and-tagging). In the sidebar, click on _Tags > sample tags_ to have these tags populate the samples in the App. By selecting certain tags, you can restrict your view to show (for instance) only those tags of interest. Here, we are focusing on open book jackets. ![](https://cdn.sanity.io/images/h6toihm1/production/a9c2ec647ee4a0aa813551da999aef02262e598c-400x400.jpg?auto=format&dpr=2&fit=max&q=75&w=400) ## Histograms We can explore the distribution of tags or annotations using FiftyOne’s [histogram panel](https://docs.voxel51.com/user_guide/app.html#histograms-panel). In this view, we’ve excluded nominal samples and are focusing on defects. The histogram shows us a breakdown of the distribution of the various defect annotations. Note that the defect tags are not mutually exclusive. As [the paper](https://arxiv.org/abs/2303.16382) describes, the defects in this dataset fall into two overall categories: _package\_defect_ and _multi\_pick_. Some of the tags then further detail the type of defect within these categories. Open book jackets are relatively common, while crushed boxes (fortunately for customers!) are less so. ![](https://cdn.sanity.io/images/h6toihm1/production/5fe3f3679937e78e0047c7c0fc1800be16572e5e-1923x1121.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Sample details Click on any of the samples to get a larger view with additional [view tools](https://docs.voxel51.com/user_guide/app.html#using-the-image-visualizer) and details available. In this image, we’ve used the crop tool to zoom tightly around the object of interest. The image carousel at the top of the interface shows the other views or slices for this sample. Any of these images may be selected to become the primary displayed image. Note that while the first three slices show the same defect tags, the fourth view is listed as nominal. The partial box defect is only visible from certain angles, making multi-view grouped datasets like this one crucial in this type of application. ![](https://cdn.sanity.io/images/h6toihm1/production/b613d108a3db187c4d92ab3b890e7db61407d90c-1920x1200.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Embeddings and the FiftyOne Brain By browsing our data in the FiftyOne App, we can get a pretty good sense of what our data looks like, including how the various defects manifest and how camera views differ from each other. This is just the beginning. Let’s dig into our data deeper with the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html#fiftyone-brain) and its [embeddings functionality](https://docs.voxel51.com/tutorials/image_embeddings.html). If you are new to embeddings, check out [this post](https://towardsdatascience.com/neural-network-embeddings-explained-4d028e6f0526) over on the Towards Data Science blog. The ARMBench dataset comes with tasks or challenges associated with each data subset. In our case, the task at hand is [defect detection](http://armbench.s3-website-us-east-1.amazonaws.com/defects.html). This is a difficult problem, as defects are often rare and unpredictable in nature, challenging supervised learning approaches. The difficulty is magnified here given the huge variety of objects and packaging present. Can analyzing embeddings in FiftyOne help shed some light on the challenge of finding and identifying pick and place defects? To start things off, we’ll import the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html?highlight=fiftyone%20brain). We’ll focus on the fourth slice of our data, which from our exploration in FiftyOne seems qualitatively less noisy in background and lighting heterogeneity. Our analysis will [compute embeddings on detection patches](https://docs.voxel51.com/user_guide/brain.html#object-embeddings-example), so we’ll also filter out the relatively small number of samples that do not have object detections. (This includes the _multi-pick_ defects, as these do not have accompanying segmentations.) For simplicity, we clone a new, smaller dataset restricted to this slice of samples. ```python 1import fiftyone.brain as fob 2from fiftyone import ViewField as F 3dataset = dataset.select_group_slices('4') 4 .filter_labels('object',F()) 5 .clone(name='ArmBench-Image-Defect-Slice4',persistent=True) ``` To assist in visualizing our embeddings, we’ll assign a label field to our object detections based on our tags. In addition to the _nominal_ label, there are two basic types of defects, _multi-pick_ and _package-defect_. _Package-defects_ are broken down further for books, boxes, and bags. To simplify the visualizations, we’ll collapse the book-related and bag-related defects into single categories. ```python 1labels = { 2 'book': ['book_jacket','open_book_jacket','open_book'], 3 'open_box': ['open_box'], 4 'partial_box':['partial_box'], 5 'crush_box':['crush_box'], 6 'bag': ['empty_bag','torn_bag'], 7 'multi_pick': ['multi_pick'], 8 'nominal': ['nominal'], 9} 10 11# set default defect label; this is overwritten for most samples 12dataset.set_values('object.detections.label',[['other_defect']]*len(dataset)) 13 14for ty,tags in labels.items(): 15 view = dataset.match_tags(tags) 16 view.set_values('object.detections.label',[[ty]]*len(view)) ``` It’s time to compute our embeddings! We’ll use the [OpenAI CLIP model](https://openai.com/research/clip) from the [FiftyOne Model Zoo](https://docs.voxel51.com/user_guide/model_zoo/index.html) to compute embeddings on our [object patches](https://docs.voxel51.com/user_guide/brain.html#object-embeddings-example). Meanwhile, the FiftyOne Brain [compute\_visualization](https://docs.voxel51.com/user_guide/brain.html#visualizing-embeddings) method will use the [UMAP](https://umap-learn.readthedocs.io/en/latest/) method to perform structure-preserving dimensionality reduction to enable a 2D visualization. ```python 1fob.compute_visualization(dataset, 2 patches_field='object', 3 embeddings='clip_embeddings', 4 brain_key='object_clip', 5 model='clip-vit-base32-torch') ``` In this method call: - `patches_field` specifies that we will compute embeddings on the bounding box patches (or crops) for the detections stored in the `object` field. - `embeddings` names a field to store our computed embedding vectors. For the CLIP model, these are 512-dimensional vectors. - `brain_key` gives an identifier to this visualization, so we can refer to it later and in the FiftyOne App. - `model` specifies a model from the [FiftyOne Model Zoo](https://docs.voxel51.com/user_guide/model_zoo/index.html). The pixels from our detection crops are passed through this model to generate our embedding vectors. In the FiftyOne App, we’ll enter a [Patches view](https://docs.voxel51.com/user_guide/app.html#viewing-object-patches) to focus on our detected objects and their corresponding embeddings. Using the FiftyOne [Embeddings panel](https://docs.voxel51.com/user_guide/app.html#embeddings-panel), it’s easy to visualize these patch embeddings, colored by label type: ![](https://cdn.sanity.io/images/h6toihm1/production/e483eed17186855a224a5c9cf2b6498645c92fac-1920x1200.png?auto=format&dpr=2&fit=max&q=75&w=1600) As it turns out, our book defects are clustered quite visibly in the lower-left-hand corner! Using the [lasso tool](https://docs.voxel51.com/user_guide/app.html#embeddings-panel), we can select this group of samples. In the samples grid, it is clear that the majority of these samples are indeed defects, and in particular, book defects: ![](https://cdn.sanity.io/images/h6toihm1/production/26db7ac0a4aa117ac9cd42257aabd9c69af6fb08-1920x1200.png?auto=format&dpr=2&fit=max&q=75&w=1600) Setting rigor aside for a moment, this lasso-ed selection of 1830 objects captures 1560 book-related defects out of a total of 1922 in the entire dataset, for a recall of 1560/1922 ~ 81%. The remaining 270 selected samples represent false positives in a pool of 31257 negatives, for a false positive rate of 270/31257 ~ .8%. (Recall and false positive rate are the preferred metrics described in [the paper](https://arxiv.org/abs/2303.16382)). Of course, this is only a rough exploration, but it is suggestive of the rich structure and information available in these embeddings. What if we just focus on samples with defects? We’ll re-use our computed embeddings, but re-compute the visualization against the smaller subset of data: ```python 1view_defects = dataset.match_tags('nominal',bool=False) 2 3fob.compute_visualization(view_defects, 4 patches_field='object', 5 embeddings='clip_embeddings', 6 brain_key='dets_clip') 7 8session = fo.launch_app(view_defects) ``` Again, some clear structure is evident in our embeddings plot. Bag-related defects, for instance, are clustered quite tightly, as selected and shown here. ![](https://cdn.sanity.io/images/h6toihm1/production/8c87fcc7ac73628712a142068c3e4e873c3d5f23-1920x1200.png?auto=format&dpr=2&fit=max&q=75&w=1600) Pretty cool stuff! It’s not the end of the story, but this analysis has definitely generated some insights for us and kick-started our effort on detecting defects in this novel dataset. Let’s take a step back and return to our first embeddings plot, which includes all samples from slice 4. While we have detailed labels for defects, the large mass of _nominal_ samples lacks annotations. You may have noticed however that the embeddings plot already gives some clues about structure in this sea of data. Selecting a cluster near the top of the mass, for instance, reveals a distinct clustering of plastic bottles! ![](https://cdn.sanity.io/images/h6toihm1/production/a0cc4f52ae330196bc97a16d07d7d895e444b0b9-1920x1200.png?auto=format&dpr=2&fit=max&q=75&w=1600) The other ‘lobes’ of the visualization are semantically meaningful as well, representing distinct clusters of cardboard boxes and objects wrapped in plastic. As a final example of the tools available in the FiftyOne Brain, let’s utilize the natural language capabilities of our CLIP embeddings to [search by the text prompt](https://docs.voxel51.com/user_guide/brain.html#text-similarity) “medicine bottle”, returning the top 100 matches. The returned results are quite consistent, and overwhelmingly located in the cluster of bottles we found earlier. Depending on our analysis, we could leverage this [zero-shot labeling](https://arxiv.org/pdf/2103.00020.pdf) capability to automatically add annotations to our dataset to give us more to work with in the large sea of _nominal_ samples. ![](https://cdn.sanity.io/images/h6toihm1/production/8c9980c3e5700224c5907404b6ee0af6f6e60cd1-1920x1200.png?auto=format&dpr=2&fit=max&q=75&w=1600) You can see more in this short video. https://www.youtube.com/watch?v=6n5OqLfXbe8 ## Start working with the dataset That’s a wrap for now! We hope you enjoyed this quick exploration of defect detection and the new [ARMBench](http://armbench.s3-website-us-east-1.amazonaws.com/) dataset. We’ll be adding this massive dataset to the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html#fiftyone-dataset-zoo) in the near future, so you’ll be able to explore it on your own in just a couple lines of Python! [Amazon dataset](https://voxel51.com/blog/tag/amazon-dataset) [ARMBench](https://voxel51.com/blog/tag/armbench) [CLIP](https://voxel51.com/blog/tag/clip) [dataset visualization](https://voxel51.com/blog/tag/dataset-visualization) [defect detection](https://voxel51.com/blog/tag/defect-detection) [FiftyOne App](https://voxel51.com/blog/tag/fiftyone-app) [object segmentation](https://voxel51.com/blog/tag/object-segmentation) [OpenAI](https://voxel51.com/blog/tag/openai) [pick and place robots](https://voxel51.com/blog/tag/pick-and-place-robots) [robotics](https://voxel51.com/blog/tag/robotics) [robots](https://voxel51.com/blog/tag/robots) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/c332c478d66b51893447f19eb71d84a940b94a09-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Finding Images with Words\\ \\ Computer Vision, Vector Search\\ \\ • \\ \\ Jan 11, 2023](https://voxel51.com/blog/finding-images-with-words) [![](https://cdn.sanity.io/images/h6toihm1/production/2867ac2853fae5362ca6bd2d208358dc94556344-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ A Google Search Experience for Computer Vision Data\\ \\ Tutorials, Vector Search\\ \\ • \\ \\ Mar 22, 2023](https://voxel51.com/blog/a-google-search-experience-for-computer-vision-data) [![](https://cdn.sanity.io/images/h6toihm1/production/73db90a7323e7eb6feb0dad3839e11c6d4ab525b-960x540.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — May 11, 2023\\ \\ Event Recaps\\ \\ • \\ \\ May 12, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-284-lllmstxt|> ## YOLO-NAS Object Detection [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Tutorials](https://voxel51.com/blog/category/tutorials) State-of-the-Art Object Detection with YOLO-NAS & FiftyOne May 4, 2023 • 5 min read Article content In this article [Setup](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone#67259c5ac636) [Generating YOLO-NAS Predictions](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone#bc8af607ed29) [Evaluation](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone#ca03e7200958) [Conclusion](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone#e9ea66fdfd27) [Join the FiftyOne community!](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone#e5b17aa18de7) In this article [Setup](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone#67259c5ac636) [Generating YOLO-NAS Predictions](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone#bc8af607ed29) [Evaluation](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone#ca03e7200958) [Conclusion](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone#e9ea66fdfd27) [Join the FiftyOne community!](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone#e5b17aa18de7) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ![](https://cdn.sanity.io/images/h6toihm1/production/f86435297a6a64cebcc63d8c49b2c2639dc92a11-2826x1496.png?auto=format&dpr=2&fit=max&q=75&w=1600) Yesterday, [Deci AI](https://deci.ai/) released a new state of the art object detection model named [YOLO-NAS](https://github.com/Deci-AI/super-gradients/blob/master/YOLONAS.md), which achieves higher mean average precision than prior models running with the same latency. In this blog post, we’ll show you how to generate predictions with YOLO-NAS and load them into FiftyOne! If you are new to [FiftyOne](https://docs.voxel51.com/), it is an open source computer vision toolset for curating better data and building better models. ## Setup If you haven’t done so already, you will need to [install FiftyOne](https://docs.voxel51.com/getting_started/install.html): ```bash 1pip install fiftyone ``` You will also need to install Deci AI’s [SuperGradients](https://github.com/Deci-AI/super-gradients/tree/master) package: ```bash 1pip install super-gradients ``` The next step is importing FiftyOne and loading a dataset from the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html). We will use a random subset of the validation split from the MS COCO dataset. We will also make the dataset persistent: ```python 1​​import fiftyone as fo 2import fiftyone.zoo as foz 3dataset = foz.load_zoo_dataset( 4 "coco-2017", 5 split = "validation", 6 max_samples=1000 7) 8dataset.name = "YOLO-NAS-demo" 9dataset.persistent = True ``` We will also use the `compute_metadata()` method to store image width and height, so that we can use these to convert between absolute and relative coordinates for bounding boxes: ```python 1dataset.compute_metadata() ``` Now we load in the YOLO-NAS model. We’ll use the large architecture (hence the “l”) and will download pretrained weights for a version of the model trained on COCO data. For a list of available weights, see [here](https://github.com/Deci-AI/super-gradients/blob/master/YOLONAS.md#a-next-generation-object-detection-foundational-model-generated-by-decis-neural-architecture-search-technology). ```python 1from super_gradients.training import models 2model = models.get("yolo_nas_l", pretrained_weights="coco") ``` ## Generating YOLO-NAS Predictions We can generate predictions for a single sample by passing the filepath for the sample’s image to our model’s \`predict()\` method, along with an optional confidence threshold: ```python 1sample = dataset.first() 2prediction = model.predict(sample.filepath, conf = 0.25) ``` We can then visualize the object bounding boxes drawn onto the image with the `show()` method for the SuperGradient `DetectionResult`: ```python 1prediction.show() ``` ![](https://cdn.sanity.io/images/h6toihm1/production/b0e7512746a6e89a3bf9e630b65d3110ce4aea1c-2260x1504.png?auto=format&dpr=2&fit=max&q=75&w=1600) To efficiently generate predictions for all of the images in our dataset, we batch these operations by passing the model the names of all the filepaths for all of our images: ```python 1fps, widths, heights = dataset.values( 2 ["filepath", "metadata.width", "metadata.height"] 3 ) 4 5 ## batch predictions 6 preds = model.predict(fps, conf = confidence)._images_prediction_lst ``` To load these predictions into FiftyOne, we need to first convert these into `Detection` label objects. This will require accessing the internals of the prediction objects, extracting the confidence, labels, and bounding boxes. We can access this information via the `_images_prediction_lst` attribute of the prediction objects. Let’s see what this looks like for a single image. ```python 1pred = model.predict(sample.filepath, conf = 0.9) 2print(next(pred._images_prediction_lst)) ``` ImageDetectionPrediction(image=array(\[\[\[170, 136, 73\],\ \ \[173, 142, 77\],\ \ \[175, 144, 79\],\ \ ...,\ \ \[ 69, 76, 42\],\ \ \[ 68, 76, 39\],\ \ \[ 70, 71, 37\]\],\ \ \[\[172, 141, 77\],\ \ \[176, 145, 80\],\ \ \[177, 146, 81\],\ \ ...,\ \ \[ 69, 77, 40\],\ \ \[ 72, 80, 43\],\ \ \[ 71, 75, 40\]\],\ \ \[\[175, 144, 79\],\ \ \[177, 146, 81\],\ \ \[178, 147, 80\],\ \ ...,\ \ \[ 70, 78, 39\],\ \ \[ 69, 77, 40\],\ \ \[ 71, 75, 40\]\],\ \ ...,\ \ \[\[188, 189, 157\],\ \ \[183, 183, 149\],\ \ \[193, 187, 153\],\ \ ...,\ \ \[186, 157, 153\],\ \ \[186, 157, 153\],\ \ \[187, 156, 154\]\],\ \ \[\[186, 183, 152\],\ \ \[187, 184, 153\],\ \ \[186, 183, 152\],\ \ ...,\ \ \[198, 134, 134\],\ \ \[195, 120, 124\],\ \ \[186, 88, 101\]\],\ \ \[\[186, 183, 150\],\ \ \[187, 184, 151\],\ \ \[186, 183, 152\],\ \ ...,\ \ \[129, 60, 63\],\ \ \[126, 57, 60\],\ \ \[107, 41, 45\]\]\], dtype=uint8), prediction=DetectionPrediction(bboxes\_xyxy=array(\[\[ 5.661982, 166.75662 , 154.55098 , 261.8113 \],\ \ \[292.11624 , 217.75893 , 352.51135 , 318.72244 \]\], dtype=float32), confidence=array(\[0.96724653, 0.9323513 \], dtype=float32), labels=array(\[62., 56.\], dtype=float32)), class\_names=\['person', 'bicycle', 'car', 'motorcycle', 'airplane', 'bus', 'train', 'truck', 'boat', 'traffic light', 'fire hydrant', 'stop sign', 'parking meter', 'bench', 'bird', 'cat', 'dog', 'horse', 'sheep', 'cow', 'elephant', 'bear', 'zebra', 'giraffe', 'backpack', 'umbrella', 'handbag', 'tie', 'suitcase', 'frisbee', 'skis', 'snowboard', 'sports ball', 'kite', 'baseball bat', 'baseball glove', 'skateboard', 'surfboard', 'tennis racket', 'bottle', 'wine glass', 'cup', 'fork', 'knife', 'spoon', 'bowl', 'banana', 'apple', 'sandwich', 'orange', 'broccoli', 'carrot', 'hot dog', 'pizza', 'donut', 'cake', 'chair', 'couch', 'potted plant', 'bed', 'dining table', 'toilet', 'tv', 'laptop', 'mouse', 'remote', 'keyboard', 'cell phone', 'microwave', 'oven', 'toaster', 'sink', 'refrigerator', 'book', 'clock', 'vase', 'scissors', 'teddy bear', 'hair drier', 'toothbrush'\]) Because the `_images_prediction_lst` attribute is a generator, we used `next()` to see what it generates, which is an `ImageDetectionPrediction` object. In this result, we can see bounding boxes stored in `xyxy` format, an array of confidence scores, and an array of integers representing the indices of the label classes in the `class_names` array for those detections. The class names are precisely the COCO class names. Here is how we can convert bounding boxes from YOLO-NAS output coordinates: ```python 1def convert_bboxes(bboxes, w, h): 2 tmp = np.copy(bboxes[:, 1]) 3 bboxes[:, 1] = h - bboxes[:, 3] 4 bboxes[:, 3] = h - tmp 5 bboxes[:, 0]/= w 6 bboxes[:, 2]/= w 7 bboxes[:, 1]/= h 8 bboxes[:, 3]/= h 9 bboxes[:, 2] -= bboxes[:, 0] 10 bboxes[:, 3] -= bboxes[:, 1] 11 bboxes[:, 1] = 1 - (bboxes[:, 1] + bboxes[:, 3]) 12 return bboxes ``` Applying this bounding box conversion, we can generate FiftyOne Detection objects for each object, and create a `Detections` object containing a list of detected objects for a given image: ```python 1def generate_detections(p, width, height): 2 class_names = p.class_names 3 dp = p.prediction 4 bboxes, confs, labels = np.array(dp.bboxes_xyxy), dp.confidence, dp.labels.astype(int) 5 if 0 in bboxes.shape: 6 return fo.Detections(detections = []) 7 8 bboxes = convert_bboxes(bboxes, width, height) 9 labels = [class_names[l] for l in labels] 10 11 detections = [\ 12 fo.Detection(\ 13 label = l,\ 14 confidence = c,\ 15 bounding_box = b\ 16 )\ 17 for (l, c, b) in zip(labels, confs, bboxes)\ 18 ] 19 return fo.Detections(detections=detections) ``` Putting it all together, we can efficiently add YOLO-NAS detection predictions to our dataset: ```python 1def add_YOLO_NAS_predictions(dataset, confidence = 0.9): 2 ## aggregation to minimize expensive operations 3 fps, widths, heights = dataset.values( 4 ["filepath", "metadata.width", "metadata.height"] 5 ) 6 7 ## batch predictions 8 preds = model.predict(fps, conf = confidence)._images_prediction_lst 9 10 ## add all predictions to dataset at once 11 dets = [\ 12 generate_detections(pred, w, h)\ 13 for pred, w, h in zip(preds, widths, heights)\ 14 ] 15 dataset.set_values("YOLO-NAS", dets) ``` Applying this to our dataset and launching a session of the FiftyOne App, we can visualize the results: ```python 1add_YOLO_NAS_predictions(dataset, confidence = 0.7) 2session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/364cda98838833db8adfe10afb95eb3f3336f11d-3434x1814.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Evaluation With the data loaded into FiftyOne, we can evaluate the quality of the object detection predictions against the "ground truth" with FiftyOne’s `evaluate_detections()` method. We will store the results with an evaluation key so we can view the resulting evaluation patches in the FiftyOne App: ```python 1res = dataset.evaluate_detections( 2 "YOLO-NAS", 3 eval_key="eval_yolonas" 4) 5session.view = dataset.to_evaluation_patches("eval_yolonas") ``` ![](https://cdn.sanity.io/images/h6toihm1/production/2a211cd4eae4d2673942c9d1a76be0c6b5349116-3434x1814.png?auto=format&dpr=2&fit=max&q=75&w=1600) We can click into one of these evaluation patches and, hovering over the detection, we can see attributes like intersection over union (IoU) score, and whether the prediction was a true positive, false positive, or false negative. ![](https://cdn.sanity.io/images/h6toihm1/production/bf61a1a8d5a18f5146296ffcd511300b8a18a4f2-3434x1814.png?auto=format&dpr=2&fit=max&q=75&w=1600) If we wanted to see images with the most false positives, we could sort the dataset accordingly: ```python 1### get images with the most FP's first 2fp_view = dataset.sort_by("eval_yolonas_fp", reverse = True) ``` We can also dig into the performance data more quantitatively by printing metrics for the most commonly occurring classes in the dataset: ```python 1counts = dataset.count_values("ground_truth.detections.label") 2classes_top10 = sorted(counts, key=counts.get, reverse=True)[:10] 3# Print a report for the top-10 classes 4res.print_report(classes=classes_top10) ``` precision recall f1-score support person 0.98 0.52 0.68 2259 chair 0.88 0.33 0.48 401 car 0.93 0.42 0.58 348 book 1.00 0.03 0.06 212 bottle 0.88 0.28 0.43 210 cup 0.97 0.45 0.61 185 dining table 0.85 0.18 0.30 153 bowl 0.86 0.34 0.49 128 bird 1.00 0.39 0.56 103 backpack 1.00 0.07 0.13 98 micro avg 0.95 0.42 0.58 4097 macro avg 0.93 0.30 0.43 4097 weighted avg 0.95 0.42 0.57 4097 ## Conclusion If you want to take this further, you may be interested in: - Applying YOLO-NAS to [video data](https://docs.voxel51.com/user_guide/dataset_creation/index.html#loading-videos), and loading these predictions into FiftyOne - Fine-tuning YOLO-NAS on your own data, and using FiftyOne’s [Evaluation API](https://docs.voxel51.com/user_guide/evaluation.html) to compare the base and fine-tuned models ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,500+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,900+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 3,900+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 266+ repositories - 58+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [Dataset Zoo](https://voxel51.com/blog/tag/dataset-zoo) [Deci AI](https://voxel51.com/blog/tag/deci-ai) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [object detection](https://voxel51.com/blog/tag/object-detection) [SuperGradient](https://voxel51.com/blog/tag/supergradient) [YOLO](https://voxel51.com/blog/tag/yolo) [YOLO-NAS](https://voxel51.com/blog/tag/yolo-nas) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/98e839c6e81bb9c4fa3ad96bf0d5d1b77ee11f6c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 1)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1) [![](https://cdn.sanity.io/images/h6toihm1/production/803cb935ddffbb5b29b6d3c73104b3a1221ddbfc-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 3)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3) [![](https://cdn.sanity.io/images/h6toihm1/production/713e4352b25d3b4ee12eab92246ceff22f808471-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Spending My First Week With FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ Aug 21, 2023](https://voxel51.com/blog/spending-my-first-week-with-fiftyone) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-285-lllmstxt|> ## FiftyOne Community Update [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) FiftyOne Computer Vision Community Update – May ‘23 May 5, 2023 • 5 min read Article content In this article [Community Spotlights](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#587d51203296) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#e1ebe82d5f7d) [Community Contributions](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#e7cf7c549508) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#e6c16275f486) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#aa9554577fc7) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#2204c2d4fbe4) [Upcoming Computer Vision Events](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#84c1ee63c082) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#38fdd87e9d16) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#9ca9dd871dd8) In this article [Community Spotlights](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#587d51203296) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#e1ebe82d5f7d) [Community Contributions](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#e7cf7c549508) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#e6c16275f486) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#aa9554577fc7) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#2204c2d4fbe4) [Upcoming Computer Vision Events](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#84c1ee63c082) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#38fdd87e9d16) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023#9ca9dd871dd8) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the monthly blog series where we bring you up to speed on recent happenings in the FiftyOne community and celebrate noteworthy milestones. 🙌 🚀 ## Community Spotlights We love hearing how FiftyOne helps you solve challenges and reach new heights! Curious what sorts of use cases are possible with [FiftyOne](https://voxel51.com/fiftyone/)? Here are just a few highlights from what community members have to say. ### **Allstate India - Insurance, IT Services, and IT Consulting** ![](https://cdn.sanity.io/images/h6toihm1/production/420c838228cf1f5eab472634e28d30a82d661a42-628x470.png?auto=format&dpr=2&fit=max&q=75&w=628) [Allstate India](https://www.allstateindia.com/) provides software development, testing, business process management, technology support, analytics and other IT-enabled services to Allstate and its subsidiaries. > _"At Allstate, my team works on auto vehicle damage inspection. Verifying the damage to a vehicle can take an insurance claim agent hours to verify, but using computer vision and FiftyOne, we can segment the parts of vehicles first, then detect the damages, and finally match the damage to repair costs and generate reports for the adjusters."_ > > Pavan Nanjundappa – Data Science Manager ### **Secury360 - Security** ![](https://cdn.sanity.io/images/h6toihm1/production/f3b3e32a40fb6084e5817bb198578e0cef94d6bd-1024x605.png?auto=format&dpr=2&fit=max&q=75&w=1024) The [Secury360](https://en.secury-360.com/) box transforms CCTV setups into proactive perimeter detection solutions that use AI to eliminate false alarms and guarantee only human detection, with 99.998% accuracy. The Secury360 box features edge AI to gradually learn the terrain through deep learning to recognize behavior and discern if someone has bad intentions. Because only human detections get through the filter, operators have more time to respond to and prevent real threats. **How FiftyOne is used:** Secury360 has to manage very large image and video datasets that are constantly being fed by devices. FiftyOne helps Secury360 constantly improve their surveillance model by enabling them to compare similar data, and in turn, deliver a more distributed dataset to train their surveillance model on. ### **See More Stories** [See more stories](https://voxel51.com/success-stories/) from people and organizations building remarkable machine learning and AI using FiftyOne and FiftyOne Teams. ### **Share Your Story!** Is your organization using FiftyOne to solve interesting computer vision problems? [Share your success story](https://voxel51.com/fiftyone-computer-vision-success-story-submission/) and claim a box of community rewards as a thank you! ![](https://cdn.sanity.io/images/h6toihm1/production/e3ed352e497d0e500ff4a1484b8422b3c9bef5cb-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) ## Product Releases In April, we released 0.20.1 to augment the recent [FiftyOne 0.20](https://voxel51.com/blog/announcing-fiftyone-0-20/) release. FiftyOne 0.20.1 contains 90+ enhancements and fixes. You can dive into the details of what’s included in the [0.20.1 release notes.](https://github.com/voxel51/fiftyone/releases/tag/v0.20.1) ## Community Contributions A quick shoutout to the following community members who made their first contributions to the FiftyOne project with the v0.20.1 release. - @karsil - [#2114](https://github.com/voxel51/fiftyone/pull/2114): YOLOv5DatasetExporter: Add flag to use export\_dir for path value - @dlangenk - [#2122](https://github.com/voxel51/fiftyone/pull/2122): Add option to import COCO annotation id - @HoopsMcann - [#2128](https://github.com/voxel51/fiftyone/pull/2128): Add interpolation option for image resizing - @oddeirikigland - [#2145](https://github.com/voxel51/fiftyone/pull/2145): Create annotation run crashes if some but not all samples are labeled - @andife - [#2177](https://github.com/voxel51/fiftyone/pull/2177): Update cvat.rst - @lauralindy - [#2198:](https://github.com/voxel51/fiftyone/pull/2198) Add info to dataset ## FiftyOne on GitHub GitHub is home to the open source FiftyOne project. Here’s the latest snapshot of what’s happening in the [FiftyOne GitHub repo](https://github.com/voxel51/fiftyone): - Total stars: 2,900+ - Total contributors: 60 - Total used by: 283 repositories - Total forks: 343 - Total issues closed so far: 784 ## FiftyOne Community Slack The FiftyOne Community [Slack channel](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) is where you can join more than 1550 machine learning engineers and data scientists using FiftyOne to improve the quality of their computer vision data and build better models. Last month alone we had 55 first time community members. Ask questions, answer questions, or simply follow along with the discussion! To make it easy to catch the highlights, every Friday we recap interesting questions and answers from Slack in [Tips & Tricks blog series](https://voxel51.com/blog/category/tips-tricks/). Recent posts include: - [FiftyOne Computer Vision Tips and Tricks – April 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023/) - [FiftyOne Computer Vision Tips and Tricks – April 7, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023/) - [FiftyOne Computer Vision Embeddings Tips and Tricks – Mar 31, 2023](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023/) - [FiftyOne Computer Vision Tips and Tricks – Mar 24, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-24-2023/) ## Computer Vision Meetups \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Voxel51 sponsors 13 virtual [Computer Vision Meetups](https://www.meetup.com/pro/computer-vision-meetups/) around the world. (To join, visit the Meetup [link](https://www.meetup.com/pro/computer-vision-meetups/) and scroll down to find the location friendliest to your time zone.) The Computer Vision Meetups are geared towards data scientists, machine learning engineers, and open source enthusiasts who want to expand their knowledge of computer vision and complementary technologies. We put an emphasis on open source software, and speakers who are computer vision practitioners or academics doing research in the field. This month’s Meetups include: ### **May ’23 Computer Vision Meetup (Americas and EMEA)** - May 11, 2023 – 10AM PT / 5PM UTC - The Role of Symmetry in Human and Computer Vision – _[Sven Dickinson](https://www.linkedin.com/in/sven-dickinson-1091b73/) (University of Toronto & Samsung)_ - Machine Learning for Fast, Motion-Robust MRI – _[Nalini Singh](https://www.linkedin.com/in/nalinimsingh/) (MIT)_ - [Register for the Zoom](https://voxel51.com/computer-vision-events/may-2023-computer-vision-meetup/?utm_source=blog) ### **May ’23 Computer Vision Meetup (APAC)** - May 25, 2023 – 10AM IST / 04:30 UTC - Wildlife Watcher: A Smart Wildlife Camera - _[Victor Anton](https://www.linkedin.com/in/victor-anton-116b45188/) (Wildlife.ai)_ - Applying Computer Vision to Real Estate at Opendoor - _[Shashwat Srivastava](https://www.linkedin.com/in/shashsrivastava/) (Opendoor)_ - [Register for the Zoom](https://voxel51.com/computer-vision-events/may-2023-computer-vision-meetup-apac/?utm_source=blog) ### **Recapping the April 27 Meetup** If you missed the last Meetup, make sure to check out [the recap blog](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023/) and watch the playbacks! - [Leveraging Attention for Improved Accuracy and Robustness](https://youtu.be/QYXIHIvesYo) _\- Hila Chefer (Tel-Aviv University)_ - [Breaking the Bottleneck of AI Deployment at the Edge with OpenVINO](https://youtu.be/rIz841UeudQ) \- _Zhuo Wu (Intel)_ ## Upcoming Computer Vision Events In addition to meetups, we invite you to join us for one or more of these upcoming [events](https://voxel51.com/computer-vision-events/): - May 31 - [Getting Started with FiftyOne Workshop (Americas & EMEA)](https://voxel51.com/computer-vision-events/) - June 8 - [June Computer Vision Meetup](https://voxel51.com/computer-vision-events/) - June 18-22 - [CVPR in Vancouver, Canada](https://voxel51.com/computer-vision-events/) - June 28 - [Getting Started with FiftyOne Workshop (Americas)](https://voxel51.com/computer-vision-events/) ## New Docs, Blogs, Videos, and Tutorials We want everyone to be successful with FiftyOne, and one of the ways we try to do that is by publishing resources that you might find helpful and handy. Here’s a list of some of the new [documentation](https://docs.voxel51.com/), [blogs](https://voxel51.com/blog/), [videos](https://www.youtube.com/@voxel51/videos), [tutorials](https://docs.voxel51.com/tutorials/index.html), [integrations](https://docs.voxel51.com/integrations/index.html), and [cheat sheets](https://docs.voxel51.com/cheat_sheets/index.html) that you may want to check out. ### **Blogs** - [State-of-the-Art Object Detection with YOLO-NAS & FiftyOne](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone/) - [Visualizing Defects in Amazon’s ARMBench Dataset Using Embeddings and OpenAI’s CLIP Model](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip/) - [The ML Menu for Model Selection: Hugging Face, Weights & Biases, and FiftyOne](https://voxel51.com/blog/ml-menu-for-model-selection-hugging-face-weights-and-biases-fiftyone/) - [Generate Movement from Text Descriptions with T2M-GPT](https://voxel51.com/blog/generate-movement-from-text-descriptions-with-t2m-gpt/) - [Getting Started with FiftyOne Workshop – April 26 Recap](https://voxel51.com/blog/getting-started-with-fiftyone-workshop-april-26-recap/) - [Webinar Recap: What’s New in FiftyOne 0.20 for Computer Vision](https://voxel51.com/blog/webinar-recap-whats-new-in-fiftyone-0-20-for-computer-vision/) - [Exploring Google Research’s Kaggle Image Matching Challenge 2023 Dataset](https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset/) - [Recapping the Computer Vision Meetup — April 13, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-13-2023/) - [Towards Controllable Diffusion Models with GLIGEN](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen/) ### **Videos** - [What’s New in FiftyOne 0.20 for Your Computer Vision Workflows](https://www.youtube.com/watch?v=DUWfP3tNQSc) - [FiftyOne Dataset Zoo: UCF101 YouTube-Based Action Recognition Dataset](https://www.youtube.com/watch?v=gDeByWVIpSE) - [FiftyOne Dataset Zoo: Families in the Wild](https://www.youtube.com/watch?v=95mFsiULngw) - [FiftyOne Dataset Zoo: Berkeley Deep Drive Autonomous Vehicle Dataset](https://www.youtube.com/watch?v=SpM6EjiKbZk) ## Voxel51’s Commitment to Open Source and Community Open source, transparency, and giving back to the computer vision community is what we are all about! Whether it’s developing the open source [FiftyOne computer vision toolset](https://github.com/voxel51/fiftyone) to help engineers and data scientists build high-quality datasets and models, sponsoring [Meetups](https://www.meetup.com/pro/computer-vision-meetups/) to help members boost their computer vision knowledge, or [giving to charitable causes](https://voxel51.com/charitable-giving/) on behalf of the community, Voxel51 is committed to bringing transparency and clarity to the world’s data. [Community Update](https://voxel51.com/blog/tag/community-update) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [open source](https://voxel51.com/blog/tag/open-source) [OSS community](https://voxel51.com/blog/tag/oss-community) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Loading related posts... [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-286-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – May 12, 2023 May 12, 2023 • 3 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#8bbed9f08a11) [Counting classes using the FiftyOne API](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#c399c563b46b) [Adding predictions to videos](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#c7bd24de8250) [Deleting samples with uniqueness less than some value](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#4c6643c54625) [Adding a VOC label to a sample](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#391bd435917a) [Getting started with model predictions in FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#b6d3d95dc3cc) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#111663545e92) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#8bbed9f08a11) [Counting classes using the FiftyOne API](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#c399c563b46b) [Adding predictions to videos](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#c7bd24de8250) [Deleting samples with uniqueness less than some value](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#4c6643c54625) [Adding a VOC label to a sample](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#391bd435917a) [Getting started with model predictions in FiftyOne](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#b6d3d95dc3cc) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023#111663545e92) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Settings QualityAuto SpeedNormal Picture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## Counting classes using the FiftyOne API Community Slack member Nahid asked, _“Is there a way to show the count of different classes in a YOLOv5 dataset with the FiftyOne API programmatically?”_ Although this can be easily done in the UI, you get a similar result with the API with something like: ```python 1import fiftyone as fo 2 3name = "my_dataset" 4dataset_dir = "/home/ec2-user/my-dataset-top-dir" 5# Create the dataset 6dataset = fo.Dataset.from_dir( 7 dataset_dir=dataset_dir, 8 dataset_type=fo.types.YOLOv5Dataset, 9 name=fo.core.dataset.make_unique_dataset_name(name), 10 split = 'test' 11) 12 13print(dataset) 14counts = dataset.count_values("ground_truth.detections.label") 15print(counts) ``` For more information about the `count_values` function, check out [the Docs.](https://docs.voxel51.com/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.count_values) ## Adding predictions to videos Community Slack member Stan asked, _“I want to use FiftyOne to add predictions to all the videos in my video dataset. Is there a way around extracting the frames first or having to rematch the frames to the video in the video dataset?”_ There are two potential patterns you could build off of. The first one is applicable if you are using a custom model. ```python 1from collections import defaultdict 2 3import fiftyone as fo 4import fiftyone.zoo as foz 5 6dataset = foz.load_zoo_dataset("quickstart-video") 7 8# Sample per-frame images on disk to feed to your model 9frames = dataset.to_frames(sample_frames=True) 10 11# Load your model here 12model = ... 13 14# Option 1: save one prediction at a time on `frames` 15# Note that when you add fields to `frames`, they will appear on the frames of `dataset` 16for frame in frames: 17 frame["predictions"] = model.predict(frame.filepath) 18 frame.save() 19 20# Option 2: save batches of predictions directly on `dataset` 21values = frames.values(["sample_id", "frame_number", "filepath"]) 22predictions = defaultdict(dict) 23for sample_id, frame_number, filepath in zip(*values): 24 predictions[sample_id][frame_number] = model.predict(filepath) 25 26dataset.set_values("frames.predictions", predictions, key_field="id") ``` The next one will work if you are using a model from the [FiftyOne Model Zoo](https://docs.voxel51.com/user_guide/model_zoo/index.html). ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("quickstart-video") 5model = foz.load_zoo_model("...") 6 7# If `model` is an image model, this automatically infers that you 8# want to run inference on the frames of the videos 9dataset.apply_model(model, label_field="predictions") ``` Learn more about [working with video datasets](https://docs.voxel51.com/getting_started/troubleshooting.html#videos-do-not-load-in-the-app), [model predictions](https://docs.voxel51.com/user_guide/dataset_creation/index.html#model-predictions), and the [FiftyOne Model Zoo](https://docs.voxel51.com/user_guide/model_zoo/index.html) in the Docs. ## Deleting samples with uniqueness less than some value Community Slack member ZKW asked, _“Is there a method that would allow me to delete samples with uniqueness less than 0.2 in a dataset?”_ Here is one way to get the job done: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5dataset = foz.load_zoo_dataset("quickstart") 6 7# Option 1 8good_view = dataset.match(F("uniqueness") > 0.2) 9good_view.keep() 10 11# Option 2 12bad_view = dataset.match(F("uniqueness") <= 0.2) 13dataset.delete_samples(bad_view) ``` For more information about FiftyOne’s capabilities for computing uniqueness, check out [the Docs](https://docs.voxel51.com/user_guide/brain.html?highlight=uniqueness#brain-image-uniqueness) and the [exploring image uniqueness tutorial](https://docs.voxel51.com/tutorials/uniqueness.html). ## Adding a VOC label to a sample Community Slack member ht asked, _“I have an existing dataset in the following structure:_ ├── labels ├── left └── right _With `labels` containing .xml file in VOC format (boundingbox) and `left` and `right` containing images from the left and right view. I would like to create a `GroupDataset`. How can I add the VOC label to each sample?”_ Check out the `fiftyone.utils.voc.VOCAnnotation` class representing a VOC annotations file. ```python 1voc_annotation = VOCAnnotation(path="path/to/label") 2detections = voc_annotation.to_detections() 3sample.set_field("detection", detections) ``` For more information about utilities for working with datasets in VOC format, check out [the Docs.](https://docs.voxel51.com/api/fiftyone.utils.voc.html?highlight=voc#module-fiftyone.utils.voc) ## Getting started with model predictions in FiftyOne Community Slack member NB asked, _“I am new to FiftyOne and exploring how to use it in my model training pipeline which uses TensorFlow Lite. How can I get started?”_ Model predictions stored in other formats can always be loaded iteratively through a simple Python loop. The example below shows how to add object detection predictions to a dataset, but plenty of other label types are also supported. ```python 1import fiftyone as fo 2 3# Ex: your custom predictions format 4predictions = { 5 "/path/to/images/000001.jpg": [\ 6 {"bbox": ..., "label": ..., "score": ...},\ 7 ...\ 8 ], 9 ... 10} 11 12# Add predictions to your samples 13for sample in dataset: 14 filepath = sample.filepath 15 16 # Convert predictions to FiftyOne format 17 detections = [] 18 for obj in predictions[filepath]: 19 label = obj["label"] 20 confidence = obj["score"] 21 22 # Bounding box coordinates should be relative values 23 # in [0, 1] in the following format: 24 # [top-left-x, top-left-y, width, height] 25 bounding_box = obj["bbox"] 26 27 detections.append( 28 fo.Detection( 29 label=label, 30 bounding_box=bounding_box, 31 confidence=confidence, 32 ) 33 ) 34 35 # Store detections in a field name of your choice 36 sample["predictions"] = fo.Detections(detections=detections) 37 38 sample.save() ``` More information about working with model predictions, check out the [Adding classifier predictions to a dataset](https://docs.voxel51.com/recipes/adding_classifications.html#Adding-Classifier-Predictions-to-a-Dataset) and [Model predictions](https://docs.voxel51.com/user_guide/dataset_creation/index.html#model-predictions) sections of the Docs. ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,600+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 2,950+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 4,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 290+ repositories - 58+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [model predictions](https://voxel51.com/blog/tag/model-predictions) [uniqueness](https://voxel51.com/blog/tag/uniqueness) [video datasets](https://voxel51.com/blog/tag/video-datasets) [VOC annotation](https://voxel51.com/blog/tag/voc-annotation) [VOC label](https://voxel51.com/blog/tag/voc-label) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/ecb6afb20436d0f0e68fbb25bcfc7443657d7b91-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Feb 10, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Feb 10, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-10-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/3f54d0a45faa06a04b5d0244dd7c092603150cf0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Mar 10, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 11, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-10-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/a3a918e30b0553723b9392ea90763379f98480a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – April 7, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Apr 7, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-287-lllmstxt|> ## Computer Vision Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup — May 11, 2023 May 12, 2023 • 5 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#603f208c0b06) [Lightning Talk: Visualizing Defects in Amazon’s ARMBench Dataset Using Embeddings and OpenAI’s CLIP Model](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#d8562719e9ef) [The Role of Symmetry in Human and Computer Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#97d555dfe306) [Machine Learning for Fast, Motion-Robust MRI](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#6b03bbb2b6e9) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#1103a2e9244e) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#6e842c459cd3) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#4daf821a56c6) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#603f208c0b06) [Lightning Talk: Visualizing Defects in Amazon’s ARMBench Dataset Using Embeddings and OpenAI’s CLIP Model](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#d8562719e9ef) [The Role of Symmetry in Human and Computer Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#97d555dfe306) [Machine Learning for Fast, Motion-Robust MRI](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#6b03bbb2b6e9) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#1103a2e9244e) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#6e842c459cd3) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-11-2023#4daf821a56c6) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We just wrapped up the May 11, 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/), and if you missed it or want to revisit it, here’s a recap! In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes this month was BRAC! We are sending this month’s charitable donation of $200 to [BRAC](https://bracusa.org/), an organization powering people to rise above poverty, on behalf of the computer vision community. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Missed the Meetup? No problem. Here are playbacks and talk abstracts from the event. ## Lightning Talk: Visualizing Defects in Amazon’s ARMBench Dataset Using Embeddings and OpenAI’s CLIP Model https://www.youtube.com/watch?v=6n5OqLfXbe8 In this lightning talk, machine learning engineer [Allen Lee](https://www.linkedin.com/in/leap-allen/) from Voxel51 gave us a quick tour of Amazon’s recently released ARMBench dataset for training “pick and place” robots. You can learn more about how to create embeddings on the dataset using the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html) to derive interesting insights in the companion [blog](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip/) and [notebook on GitHub](https://github.com/voxel51/fiftyone-examples/blob/master/examples/armbench_defect_detection.ipynb). ## The Role of Symmetry in Human and Computer Vision https://www.youtube.com/watch?v=ohN5YoGfses Symmetry is one of the most ubiquitous regularities in our natural world. For almost 100 years, human vision researchers have studied how the human vision system has evolved to exploit this powerful regularity as a basis for grouping image features. While computer vision is a much younger discipline, the trajectory is similar, with symmetry playing a major role in both perceptual grouping and object representation. After briefly reviewing some of the milestones in symmetry-based perceptual grouping and object representation/recognition in both human and computer vision, I will review our efforts that draw on computer vision to understand the role that symmetry plays in human scene perception. Conversely, I will also look at how these results in human scene perception can strengthen the performance of modern deep learning computer vision systems for scene perception. [Sven Dickinson](https://www.linkedin.com/in/sven-dickinson-1091b73/) is Professor of Computer Science at the University of Toronto, and is also Vice President and Head of the new Samsung Toronto AI Research Center. [Learn more](https://www.cs.toronto.edu/~sven/) about his research and publications. Q&A from the talk included: - How do you think these learnings can inform future deep learning foundational models like SAM, Symmetry-Net, etc? - MAT is very susceptible to noise in the edges (which are likely present); how do you handle that to get a "reasonable" medial axis? - How can you tell whether it’s our ability to use symmetry vs our ability to extrapolate lines to the implied junction point that explains why removing the middle is so important? - Is there a way to remove different types of symmetry and does it have an effect on human categorization? - Perhaps the same experiment would show a different result if it used Resnet instead? - Separation score had a reverse trend for VGG16 vs human, any insights on why this is the case? - Have similar tests been made on a Transformer architecture? Some symmetry is “enforced” by convolution (translational), so could the transformer architecture be less rigid for symmetry and if so, would this have less of an effect? - Is there a role of something like "amount of expected symmetry"? i.e., Might the visual system weigh the importance of symmetries more in scenes that are expected to have them? - What software was used to create the visualizations? ## Machine Learning for Fast, Motion-Robust MRI https://www.youtube.com/watch?v=mugTIp7MfSQ Magnetic resonance imaging (MRI) is a powerful imaging modality that enables detailed visualization of tissue content and structure. However, MRI suffers from long acquisition times and significant susceptibility to motion artifacts. This talk will explore deep learning approaches that incorporate imaging physics to produce high-quality MR images from highly accelerated and/or motion-corrupted data. [Nalini Singh](https://www.linkedin.com/in/nalinimsingh/) is a Ph.D. student at the Harvard-MIT Program in Health Sciences and Technology, working in the Medical Vision Group at the MIT Computer Science and Artificial Intelligence Laboratory. Her primary research interests are in medical image reconstruction and analysis, signal processing, and inverse problems. Q&A from the talk included: - With adjacent pixels being of the same portion, how is occlusion by other vessel/body parts handled? - Can we implement this method in real time? For the example by attaching eeg to the head? - Would it be preferable to incorporate ordinary cameras (presumably using mirrors) to recover the motion directly, rather than hoping that MRI-only reconstruction is accurate (rather than merely plausible)? - When imposing consistency, would it be possible to account for the possible motion, not enforce that they match exactly, but that the change can be explained by a motion? ## Join the Computer Vision Meetup! \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Computer Vision Meetup membership has grown to over [4,000 members](https://www.meetup.com/pro/computer-vision-meetups/) in just under a year! The goal of the Meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. Join one of the 13 Meetup locations closest to your timezone. - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) ## What’s Next? We have exciting speakers already signed up over the next few months! Become a member of the [Computer Vision Meetup closest to you](https://www.meetup.com/pro/computer-vision-meetups/), then register for the Zoom. Up next on May 25 at 10 AM IST we have the APAC-timezone-friendly Computer Vision Meetup happening with talks including: - **YOLO-NAS - SOTA Object Detection Generated by NAS**– Ofri Masad _(Deci.ai)_ - **Wildlife Watcher: A Smart Wildlife Camera**– Victor Anton _(Wildlife.ai)_ - **Applying Computer Vision to Real Estate at Opendoor**– Shashwat Srivastava _(Opendoor)_ You can find a complete schedule of upcoming Meetups on [the Voxel51 Events page](https://voxel51.com/computer-vision-events/). ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. — _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [Amazon dataset](https://voxel51.com/blog/tag/amazon-dataset) [ARMBench](https://voxel51.com/blog/tag/armbench) [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/991b89d515d9f7fe4eea26c14396eb51116b4a8b-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup – April 27, 2023\\ \\ Event Recaps\\ \\ • \\ \\ Apr 28, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/67b2a5cfcef141ff0b39958c5528e2c8648c3d4c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visualizing Defects in Amazon’s ARMBench Dataset Using Embeddings and OpenAI’s CLIP Model\\ \\ Datasets\\ \\ • \\ \\ May 4, 2023](https://voxel51.com/blog/visualize-amazon-armbench-dataset-using-embeddings-and-clip) [![](https://cdn.sanity.io/images/h6toihm1/production/2260b433f7da8171c8f20164bf89e715d04cb7a6-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — June 8, 2023\\ \\ Event Recaps\\ \\ • \\ \\ Jun 9, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-288-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – May 19, 2023 May 19, 2023 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#078556953bea) [Deleting specific labels in a dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#064aa116a889) [Improving the performance of extracting and downloading of object patches](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#6125e3452f6c) [Creating dataset views after filtering](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#9ac71f839f86) [How to quickly load large thumbnails into the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#f3156044d03a) [Adding custom sidebar groups](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#0d1399393fc3) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#9928437f21f6) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#6d90b8bcbc19) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#078556953bea) [Deleting specific labels in a dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#064aa116a889) [Improving the performance of extracting and downloading of object patches](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#6125e3452f6c) [Creating dataset views after filtering](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#9ac71f839f86) [How to quickly load large thumbnails into the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#f3156044d03a) [Adding custom sidebar groups](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#0d1399393fc3) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#9928437f21f6) [What’s next?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023#6d90b8bcbc19) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## Deleting specific labels in a dataset Community Slack member ZKW asked, _“Is there a method for deleting labels with specific classes in a field? For example, in a prediction I have classes `['person', 'car', 'bus', 'bike']`, but I want to delete all the labels with the class ' `bus`' and ' `bike`_' _and only leave `['person', 'car']`. How can I delete them in the dataset, rather than creating a new view that omits them?”_ If you want to omit certain classes from view, you can create a view using \`filter\_labels()\`. For instance, we can use the negation operator ‘~\` in conjunction with \`is\_in()\` to filter our view to labels whose classes are not in a certain list: ```python 1from fiftyone import ViewField as F 2classes_to_omit = ['bus','bike'] 3view = dataset.filter_labels( 4 'ground_truth', 5 ~F('label').is_in(classes_to_omit) 6 ) ``` To delete classes from the dataset entirely, you can use the \`delete\_labels()\` method. One way to use this method is by passing in a list of label IDs that you want to delete. You can obtain the IDs of the labels you want to delete using \`filter\_labels()\` as above, but without the negation operator, and then using the \`values()\` aggregation method (with \`unwind = True\` to flatten the list) to extract the IDs. ```python 1from fiftyone import ViewField as F 2 3classes_to_delete = ['bus','bike'] 4 5labels_to_delete = dataset.filter_labels( 6 'ground_truth', 7 F('label').is_in(classes_to_delete) 8 ) 9 10label_ids_to_delete = labels_to_delete.values( 11 "ground_truth.detections.id", unwind=True 12 ) 13 14dataset.delete_labels(ids = label_ids_to_delete) ``` For more information about the [attributes of labels](https://docs.voxel51.com/user_guide/annotation.html#label-attributes) stored in dataset samples, check out the FiftyOne Docs. ## Improving the performance of extracting and downloading of object patches Community Slack member mserrari asked, _“I have images that are stored on a remote server that are about 60 MB. As I understand it, in order to generate a patch view and visualize it, I have to first download it into my browser. The patches are very small, less than 100KB. Is there a way to generate and fetch pre-generated patches from large images similar to the way thumbnails are generated, in order to reduce the time and overhead required to view them?_ One possible solution is to use FiftyOne to generate the patches view on the same remote server where the dataset is located. ```python 1dataset.to_patches(my_field).export("/path/", dataset_type=fo.types.ImageClassificationDirectoryTree, label_field=my_field) ``` Learn more about [object patches](https://docs.voxel51.com/user_guide/using_views.html#object-patches) in the FiftyOne Docs. ## Creating dataset views after filtering Community Slack member Joy asked, _“How can I create a view of my dataset after a filtering operation?”_ One solution is to create a [saved view](https://docs.voxel51.com/user_guide/using_views.html#saving-views) from the [filtered view](https://docs.voxel51.com/user_guide/using_views.html#filtering), and then [load that saved view](https://docs.voxel51.com/user_guide/using_views.html#saving-views) in code. For example loading a saved view you have created: ```python 1import fiftyone as fo 2 3dataset = fo.load_dataset("quickstart") 4 5# Retrieve a saved view 6cats_view = dataset.load_saved_view("cats-view") 7print(cats_view) ``` ## How to quickly load large thumbnails into the FiftyOne App Community Slack member Alexey asked, _“I'm creating a dataset with very large image files using the FiftyOne app running on a Ubuntu host. On a remote MacOS host I have observed slow image loading in the FiftyOne App. It looks like the frontend downloads the entire image before any of the dataset manipulations take place. Is there a way to download only the thumbnail images that populate the grid in order to shorten the load time in the App?”_ Yes! You can use [multiple media fields](https://docs.voxel51.com/user_guide/app.html#app-multiple-media-fields) to have lower-res thumbnails displayed in the Apps thumbnail grid. For example, here we create thumbnail images for use in the App’s grid view and store their paths in a `thumbnail_path` field: ```python 1import fiftyone as fo 2import fiftyone.utils.image as foui 3import fiftyone.zoo as foz 4 5dataset = foz.load_zoo_dataset("quickstart") 6 7# Generate some thumbnail images 8foui.transform_images( 9 dataset, 10 size=(-1, 32), 11 output_field="thumbnail_path", 12 output_dir="/tmp/thumbnails", 13) ``` More specifically in the snippet above, `size=(-1, 32)` is using `fiftyone.utils.image.transform_image` to specify an optional (width, height) for the image. One dimension can be -1, in which case the aspect ratio is preserved. For more information on [FiftyOne App configuration changes](https://docs.voxel51.com/api/fiftyone.core.odm.dataset.html#fiftyone.core.odm.dataset.DatasetAppConfig) you can make, check out the FiftyOne Docs. ## Adding custom sidebar groups Community Slack member Agfian asked, _“I want to make a custom sidebar group in the FiftyOne App to better organize my fields. Is it possible to add the sidebar group using code at the same time as creating my dataset?”_ Yes! For example in the snippet we add `brightness` to the `Meta Features` sidebar group: ```python 1sidebar_groups = fo.DatasetAppConfig.default_sidebar_groups(dataset) 2 3features_sg = fo.SidebarGroupDocument(name="Meta Features") 4features_sg.paths = ["brightness"] 5 6# Add a new group 7sidebar_groups.append(features_sg) 8 9dataset.app_config.sidebar_groups = sidebar_groups 10dataset.save() ``` ![](https://cdn.sanity.io/images/h6toihm1/production/8431241d2f21c6f79da8bf6a0196a271e8030f8f-302x512.png?auto=format&dpr=2&fit=max&q=75&w=302) For more information about working with [Sidebar Groups](https://docs.voxel51.com/user_guide/app.html#sidebar-groups), check out the FiftyOne Docs. ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,600+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 3,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 4,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 266+ repositories - 58+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## What’s next? - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [dataset views](https://voxel51.com/blog/tag/dataset-views) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne App](https://voxel51.com/blog/tag/fiftyone-app) [object patches](https://voxel51.com/blog/tag/object-patches) [sidebar](https://voxel51.com/blog/tag/sidebar) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/205569e4c6b9ed68023e0ed430acbb40a981026f-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Tips and Tricks for Customizing your Computer Vision Workflows – Mar 03, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 4, 2023](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-customizing-your-computer-vision-workflows-mar-03-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/3347737caaf6592de04f0f98a3dc822f7543fa58-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Feb 24, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Feb 25, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-feb-24-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/5eca7f455938fa6e0b4755f60529b4123ae91046-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – May 26, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ May 26, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-289-lllmstxt|> ## CVPR 2023 Survival Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) CVPR 2023 Survival Guide May 25, 2023 • 10 min read Article content In this article [10 papers you won’t want to miss](https://voxel51.com/blog/cvpr-2023-survival-guide#a2badd7f906d) [Beyond Appearance: a Semantic Controllable Self-Supervised Learning Framework for Human-Centric Visual Tasks](https://voxel51.com/blog/cvpr-2023-survival-guide#467900befd29) [DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation](https://voxel51.com/blog/cvpr-2023-survival-guide#1765fa297719) [F2-NeRF: Fast Neural Radiance Field Training with Free Camera Trajectories](https://voxel51.com/blog/cvpr-2023-survival-guide#05453f0d3143) [GLIGEN: Open-Set Grounded Text-to-Image Generation](https://voxel51.com/blog/cvpr-2023-survival-guide#bb29e25d8c69) [ImageBind: One Embedding Space To Bind Them All](https://voxel51.com/blog/cvpr-2023-survival-guide#75105ce603aa) [Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation](https://voxel51.com/blog/cvpr-2023-survival-guide#8dad1c2c1dd9) [MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures](https://voxel51.com/blog/cvpr-2023-survival-guide#88417e9bc901) [Planning-oriented Autonomous Driving](https://voxel51.com/blog/cvpr-2023-survival-guide#35d8e6cefc53) [SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation](https://voxel51.com/blog/cvpr-2023-survival-guide#c0b09e93d83d) [VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation](https://voxel51.com/blog/cvpr-2023-survival-guide#e0e53b1dcfac) [Visit Voxel51 at CVPR!](https://voxel51.com/blog/cvpr-2023-survival-guide#3c1b754ffc09) [Join the FiftyOne community!](https://voxel51.com/blog/cvpr-2023-survival-guide#8c8a7492eaf1) In this article [10 papers you won’t want to miss](https://voxel51.com/blog/cvpr-2023-survival-guide#a2badd7f906d) [Beyond Appearance: a Semantic Controllable Self-Supervised Learning Framework for Human-Centric Visual Tasks](https://voxel51.com/blog/cvpr-2023-survival-guide#467900befd29) [DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation](https://voxel51.com/blog/cvpr-2023-survival-guide#1765fa297719) [F2-NeRF: Fast Neural Radiance Field Training with Free Camera Trajectories](https://voxel51.com/blog/cvpr-2023-survival-guide#05453f0d3143) [GLIGEN: Open-Set Grounded Text-to-Image Generation](https://voxel51.com/blog/cvpr-2023-survival-guide#bb29e25d8c69) [ImageBind: One Embedding Space To Bind Them All](https://voxel51.com/blog/cvpr-2023-survival-guide#75105ce603aa) [Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation](https://voxel51.com/blog/cvpr-2023-survival-guide#8dad1c2c1dd9) [MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures](https://voxel51.com/blog/cvpr-2023-survival-guide#88417e9bc901) [Planning-oriented Autonomous Driving](https://voxel51.com/blog/cvpr-2023-survival-guide#35d8e6cefc53) [SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation](https://voxel51.com/blog/cvpr-2023-survival-guide#c0b09e93d83d) [VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation](https://voxel51.com/blog/cvpr-2023-survival-guide#e0e53b1dcfac) [Visit Voxel51 at CVPR!](https://voxel51.com/blog/cvpr-2023-survival-guide#3c1b754ffc09) [Join the FiftyOne community!](https://voxel51.com/blog/cvpr-2023-survival-guide#8c8a7492eaf1) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## 10 papers you won’t want to miss \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop The annual IEEE/CVF [Conference on Computer Vision and Pattern Recognition](https://cvpr2023.thecvf.com/) (CVPR) is just around the corner. Every year, thousands of computer vision researchers and engineers from across the globe come together to take part in this monumental event. The prestigious conference, which can trace its origin back to 1983, represents the pinnacle of progress in computer vision. With CVPR playing host to some of the field’s most pioneering projects and painstakingly crafted papers, it's no wonder that the conference has the fourth highest h5-index of any conference or publication, trailing only _Nature_, _Science_, and _The New England Journal of Medicine._ This year, the conference will take place in Vancouver, Canada, from June 18th - June 22nd. With [2359 accepted papers](https://cvpr2023.thecvf.com/Conferences/2023/AcceptedPapers), [100 workshops](https://cvpr2023.thecvf.com/Conferences/2023/workshop-list), [33 tutorials](https://cvpr2023.thecvf.com/Conferences/2023/tutorial-list), and Flash Sessions happening in the Expo (including two by Voxel51), CVPR will have something for everyone. To help you make the most of a jam-packed June week, we’ve compiled a list of the top 10 papers you just can’t miss, with links and summaries. We selected these papers based on GitHub project star counts, perceived impact to the field, and personal interest. Here are our top picks in alphabetical order: 01. [Beyond Appearance: a Semantic Controllable Self-Supervised Learning Framework for Human-Centric Visual Tasks](https://voxel51.com/blog/cvpr-2023-survival-guide#beyond-appearance) 02. [DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation](https://voxel51.com/blog/cvpr-2023-survival-guide#dream-booth) 03. [F2-NeRF: Fast Neural Radiance Field Training with Free Camera Trajectories](https://voxel51.com/blog/cvpr-2023-survival-guide#f2-nerf) 04. [GLIGEN: Open-Set Grounded Text-to-Image Generation](https://voxel51.com/blog/cvpr-2023-survival-guide#gligen) 05. [ImageBind: One Embedding Space To Bind Them All](https://voxel51.com/blog/cvpr-2023-survival-guide#imagebind) 06. [Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation](https://voxel51.com/blog/cvpr-2023-survival-guide#mask-dino) 07. [MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures](https://voxel51.com/blog/cvpr-2023-survival-guide#mobile-nerf) 08. [Planning-oriented Autonomous Driving](https://voxel51.com/blog/cvpr-2023-survival-guide#uniad) 09. [SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation](https://voxel51.com/blog/cvpr-2023-survival-guide#sad-talker) 10. [VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation](https://voxel51.com/blog/cvpr-2023-survival-guide#video-fusion) ## Beyond Appearance: a Semantic Controllable Self-Supervised Learning Framework for Human-Centric Visual Tasks - Links: ( [Arxiv](https://arxiv.org/abs/2303.17602) \| [Code](https://github.com/tinyvision/SOLIDER)) - Authors: Weihua Chen, Xianzhe Xu, Jian Jia, Hao luo, Yaohua Wang, Fan Wang, Rong Jin, Xiuyu Sun _Beyond Appearance_ introduces a new approach to human-centric visual tasks like pose estimation and pedestrian tracking which aims to learn a generalized human representation from a plethora of unlabeled human images. Unlike traditional self-supervised learning methods, this method, dubbed SOLIDER (Semantic cOntrollable seLf-supervIseD lEaRning framework), incorporates prior knowledge from human images to create pseudo-semantic labels as a way to integrate more semantic information into the learned representation. The real game-changer, however, is SOLIDER's conditional network with a semantic controller. Different downstream tasks demand varying ratios of semantic and appearance information. For instance, [human parsing](https://paperswithcode.com/task/human-parsing#:~:text=%E2%80%A2%202%20datasets-,Human%20parsing%20is%20the%20task%20of%20segmenting%20a%20human%20image,%2C%20torso%2C%20arms%20and%20legs.) requires high semantic information, while person re-identification instead benefits more from appearance information. SOLIDER’s semantic controller allows users to tailor these ratios to meet specific task requirements. ## DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation - Links: ( [Arxiv](https://arxiv.org/abs/2208.12242) \| [Project Page](https://dreambooth.github.io/) \| [Dataset](https://github.com/google/dreambooth)) - Authors: Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, Kfir Aberman DreamBooth’s fine-tuned text-to-image diffusion models bridge the gap between human imagination and AI interpretation. The innovation lies in the authors’ approach to subject-driven generation: the user guides the AI's creative process. Gone are the days of AI merely replicating what it has learned. With DreamBooth, it's about molding AI's understanding in real time. DreamBooth empowers users to have more control over the AI's image generation process, effectively making the AI a collaborator rather than just a tool. This fine-tuning of diffusion models enables a more nuanced and responsive generation of images from text input. With DreamBooth, your words guide the brushstrokes of the model’s metaphorical paint brush. ## F2-NeRF: Fast Neural Radiance Field Training with Free Camera Trajectories - Links: ( [Arxiv](https://arxiv.org/abs/2303.15951v1) \| [Code](https://github.com/Totoro97/f2-nerf) \| [Project Page](https://totoro97.github.io/projects/f2-nerf/)) - Authors: Peng Wang, Yuan Liu, Zhaoxi Chen, Lingjie Liu, Ziwei Liu, Taku Komura, Christian Theobalt, Wenping Wang In _F2-NeRF_, the paper’s authors present a novel approach to training Neural Radiance Fields (NeRFs). The approach is based on a new type of NeRF called a _Fast-Free-NeRFs (F2-NeRF),_ which make use of fast Fourier feature-based volume rendering to significantly reduce the computational complexity of novel view synthesis - hence the ‘Fast’. The second F, ‘Free’, derives from a new method which the authors dub _perspective warping_, that can be applied to create 3D scenes from arbitrary _free_ camera trajectories, including images captured along non-linear and non-uniform paths. This method brings an unprecedented level of flexibility to NeRFs. F2-NeRF isn't just a paper to read; it's a glimpse into the future of 3D scene understanding and creation. ## GLIGEN: Open-Set Grounded Text-to-Image Generation - Links: ( [Arxiv](https://arxiv.org/abs/2301.07093) \| [Code](https://github.com/gligen/GLIGEN) \| [Project Page](https://gligen.github.io/)) - Authors: Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, Yong Jae Lee _GLIGEN_ sets a new standard for text-to-image generation with **G** rounded **L** anguage- **I** mage **GEN** eration. The method excels at handling open-set conditions - a context where the model is required to generate images of objects or scenes it hasn't been explicitly trained on. This represents a dramatic improvement over traditional models that struggle when faced with concepts outside of their training set. The model achieves this versatility by _grounding_ the text-to-image generation process in the semantics of the input text, as opposed to relying solely on learnt correlations. This enables GLIGEN to better handle novel scenarios and generate images with a level of detail and accuracy previously unseen. The paper's exploration of open-set grounded text-to-image generation paves the way for applications in visual storytelling, content creation, and AI-powered design tools. For a thorough presentation of GLIGEN by the project’s lead author, [Yuheng Li](https://yuheng-li.github.io/), see the blog post [Towards Controllable Diffusion Models with GLIGEN](https://voxel51.com/blog/towards-controllable-diffusion-models-with-gligen/). ## ImageBind: One Embedding Space To Bind Them All - Links: ( [Paper](https://dl.fbaipublicfiles.com/imagebind/imagebind_final.pdf) \| [Code](https://github.com/facebookresearch/ImageBind)) - Authors: Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, Ishan Misra With _ImageBind_, FAIR and Meta AI researcherspresent a framework that learns a common embedding space for six modalities: images, text, audio, depth, thermal, and inertial measurement unit (IMU). This approach signifies a major departure from previous models, which typically required individual mappings for each domain pair, such as [text-to-image](https://openai.com/research/clip) and [audio-to-image](https://arxiv.org/abs/1705.08168). ImageBind streamlines image-to-image translation, and gives rise to a variety of emergent applications, including cross-modal retrieval, and state-of-the-art zero shot recognition tasks. It’s no surprise that the project’s GitHub repo [accumulated more than 5,000 stars within its first week](https://star-history.com/#facebookresearch/ImageBind&Date). ## Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation - Links: ( [Arxiv](https://arxiv.org/abs/2206.02777) \| [Code](https://github.com/IDEA-Research/MaskDINO)) - Authors: Feng Li, Hao Zhang, Huaizhe xu, Shilong Liu, Lei Zhang, Lionel M. Ni, Heung-Yeung Shum In analogy with unified CNN-based models like Mask-R-CNN, _Mask DINO_ presents a unified, transformer-based framework for object detection and segmentation, bringing together the performance of transformers with the simplicity and efficiency of single-stage pipelines. The model does so by extending DINO (DETR with Improved Denoising Anchor Boxes) with a mask prediction branch, which generates binary prediction masks from dot products between embeddings of the query and the pixel maps. This makes the model capable of instance, semantic, and panoptic segmentation. At the time of the paper’s Arxiv submission, Mask DINO achieved state-of-the-art performance on the COCO datasets for all three of these segmentation tasks. ## MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures - Links: ( [Arxiv](https://arxiv.org/abs/2208.00277) \| [Code](https://github.com/google-research/jax3d/tree/main/jax3d/projects/mobilenerf)) - Authors: Zhiqin Chen, Thomas Funkhouser, Peter Hedman, Andrea Tagliasacchi Traditional volumetric rendering of neural radiance fields (NeRFs), is compute-intensive. With MobileNeRF, researchers have brought novel view synthesis to mobile devices. This breakthrough is achieved through the introduction of a new NeRF representation based on textured polygons, which shifts some of the computational load from over-utilized [shader](https://en.wikipedia.org/wiki/Shader) cores to the [rasterizer](https://en.wikipedia.org/wiki/Rasterisation), which is typically underutilized in NeRF computations. By optimizing NeRFs for mobile architectures, MobileNeRF opens up a wide array of potential applications, from on-the-go 3D model visualization to real-time augmented reality experiences, even on devices with limited computational capabilities. This paper is an absolute must-read for anyone interested in 3D rendering or the future of mobile computing. ## Planning-oriented Autonomous Driving - Links: ( [Arxiv](https://arxiv.org/abs/2212.10156) \| [Code](https://github.com/OpenDriveLab/UniAD) \| [Project Page](https://opendrivelab.github.io/UniAD/)) - Authors: Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, Lewei Lu, Xiaosong Jia, Qiang Liu, Jifeng Dai, Yu Qiao, Hongyang Li In _Planning-oriented Autonomous Driving_, present a compelling argument for the value of strategic planning in autonomous driving systems. Whereas autonomous driving has traditionally been treated as a sequence of modular tasks, each performed by a different model, this can lead to the accumulation of errors from subtasks, as well as from poor coordination. Instead, the authors of the paper propose a “unified” paradigm, where a single model is trained to optimize the _planning_ of the driving itself. This 'planning-oriented' approach offers autonomous vehicles the ability to make decisions akin to human drivers, accounting for variables such as traffic, pedestrian behavior, and road conditions. To justify the new paradigm, the paper’s authors introduce a full-stack framework for unified autonomous driving (UniAD) which achieves state-of-the-art performance across the board on the notoriously challenging [nuScenes](https://www.nuscenes.org/) dataset. ## SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation - Links: ( [Arxiv](https://arxiv.org/abs/2211.12194v2) \| [Code](https://github.com/OpenTalker/SadTalker) \| [Project Page](https://sadtalker.github.io/)) - Authors: Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang, Xi Shen, Yu Guo, Ying Shan, Fei Wang You’ll likely feel more excitement, fear, and amazement than sadness when you see how _SadTalker_ can bring a portrait to life. SadTalker marries audio and images to generate often-uncanny talking head videos with intricate facial expressions and head poses. The model works with photographs or cartoons, and English or Chinese audio - it even works for singing! The key to SadTalker’s success is its use of 3D Morphable Model (3DMM). While models based on 2D motion can result in unnatural head movement, distorted expressions, and identity modification, and models based on full 3D information can result in stiff or incoherent videos, SadTalker threads the needle by restricting its 3D modeling to motion. The model generates coefficients for head, pose, and expression, and uses these to render the final video. In so doing, SadTalker bypasses the problems of both fully-2D and fully-3D approaches, and surpasses prior methods in video and motion quality. ## VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation - Links: ( [Arxiv](https://arxiv.org/abs/2303.08320v3) \| [Code](https://github.com/modelscope/modelscope)) - Authors: Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, Tieniu Tan VideoFusion tackles the challenging problem of extending diffusion-based models from image generation to video generation. The paper’s novel involves decomposing the complex, high-dimensional diffusion process into two types of noise: a _uniform_ noise that is shared across all frames, and a time-varying _residual_ noise. Breaking the problem down in this way - essentially _content_ and _motion_ \- accounts for the content redundancy and temporal correlation inherent in video data. One major benefit that VideoFusion’s decomposition brings is enhanced efficiency. Because the base noise is shared by all frames, prediction of this noise can be reduced to an image-generation diffusion task. This simplifies the learning of time-dependent features of video data, and reduces latency by 57.5% compared to video-based diffusion methods. ## Visit Voxel51 at CVPR! Want to discuss the state of computer vision, talk about data-centric AI, or learn how the open source computer vision toolkit [FiftyOne](https://github.com/voxel51/fiftyone) can help you overcome data quality issues and build higher quality models? Come by our booth #1618 at CVPR. Not only will we be available to discuss all things computer vision with you, we’d also simply love to meet fellow members of the CV community, and swag you up with some of our latest and greatest threads. Want to make your dataset easily accessible to the fastest growing community in machine learning and computer vision? Add your dataset to the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html) so anyone can load it with a single line of code. Reach out to me on [Linkedin](https://www.linkedin.com/in/jacob-marks/)! I'd love to discuss how we can work together to help bring your research to a wider audience :) ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,600+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 3,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 4,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 266+ repositories - 58+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [CVPR](https://voxel51.com/blog/tag/cvpr) [DreamBooth](https://voxel51.com/blog/tag/dreambooth) [F2-NeRF](https://voxel51.com/blog/tag/f2-nerf) [GLIGEN](https://voxel51.com/blog/tag/gligen) [ImageBind](https://voxel51.com/blog/tag/imagebind) [Mask DINO](https://voxel51.com/blog/tag/mask-dino) [MobileNeRF](https://voxel51.com/blog/tag/mobilenerf) [SadTalker](https://voxel51.com/blog/tag/sadtalker) [SOLDIER](https://voxel51.com/blog/tag/soldier) [VideoFusion](https://voxel51.com/blog/tag/videofusion) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0aa3f8dad8ae1464d05d81ac4a301bd92aea55e3-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2024 Survival Guide: Five Vision-Language Papers You Don’t Want to Miss\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Apr 15, 2024](https://voxel51.com/blog/cvpr-2024-survival-guide-five-vision-language-papers-you-dont-want-to-miss) [![](https://cdn.sanity.io/images/h6toihm1/production/964c6084b2c194bb2816f2888b6d2e2d9a2ba447-1200x675.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2024 Datasets and Benchmarks – Part 1: Datasets\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Apr 23, 2024](https://voxel51.com/blog/cvpr-2024-datasets-and-benchmarks-part-1-datasets) [![](https://cdn.sanity.io/images/h6toihm1/production/770b8cfdbd7944916b1195dc11e5b173dbab8e97-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 and the State of Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 18, 2023](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-290-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – May 26, 2023 May 26, 2023 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#dd37ffcc96a9) [Computing embeddings on a PyTorch model](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#ed8383a13ff6) [Exporting with absolute paths vs filenames](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#da4fec4b2f89) [Simplest way to get data into the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#f79c17d12c34) [Downloading annotated videos stored locally or remotely](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#e113ec4c0826) [Deleting a batch of samples using a DatasetView](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#aa24692d4628) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#eaf09b45a51e) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#dd37ffcc96a9) [Computing embeddings on a PyTorch model](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#ed8383a13ff6) [Exporting with absolute paths vs filenames](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#da4fec4b2f89) [Simplest way to get data into the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#f79c17d12c34) [Downloading annotated videos stored locally or remotely](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#e113ec4c0826) [Deleting a batch of samples using a DatasetView](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#aa24692d4628) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023#eaf09b45a51e) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## Computing embeddings on a PyTorch model Community Slack member Pedro asked, _“How can I compute embeddings using my PyTorch model? I have trained using ResNet-18 and want to also use it for the embeddings computation. Is this possible?”_ FiftyOne provides powerful [embeddings visualization](https://docs.voxel51.com/user_guide/brain.html#visualizing-embeddings) capabilities via the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html#). With it you can generate low-dimensional representations of the samples and objects in your datasets. ![](https://cdn.sanity.io/images/h6toihm1/production/1f8711e98f0973fbded8235b95ac1719d11e6b46-1999x1191.png?auto=format&dpr=2&fit=max&q=75&w=1600) Check out the [Using Image Embeddings](https://docs.voxel51.com/tutorials/image_embeddings.html) tutorial in the FiftyOne Docs for step-by-step instructions on how to: - Load a dataset into FiftyOne - Use [`compute_visualization()`](https://voxel51.com/docs/fiftyone/api/fiftyone.brain.html#fiftyone.brain.compute_visualization) to generate 2D representations of images - Provide custom embeddings to [`compute_visualization()`](https://voxel51.com/docs/fiftyone/api/fiftyone.brain.html#fiftyone.brain.compute_visualization) - Visualize embeddings via [interactive plots](https://voxel51.com/docs/fiftyone/user_guide/plots.html) - Identify anomalous/incorrect image labels - Find examples of scenarios of interest - Pre-annotate unlabeled data for training ## Exporting with absolute paths vs filenames Community Slack member ZKW asked, _“Is there a method in FiftyOne to add absolute paths in annotation.xml rather than providing a file name?_ Yes! You can pass the optional `abs_paths=True` option to export absolute paths rather than filenames. For example: ```python 1dataset.export( 2 ... 3 dataset_type=fo.types.CVATImageDataset, 4 abs_paths=True, 5) ``` For more information on the options available when [exporting FiftyOne datasets](https://docs.voxel51.com/user_guide/export_datasets.html?highlight=abs_paths%20true), check out the FiftyOne Docs. ## Simplest way to get data into the FiftyOne App Community Slack member J asked, _“What is the simplest way to get example datasets into the FiftyOne App so I can start experimenting with the tool?”_ The easiest way to get a dataset into the FiftyOne App is to load the `quickstart` dataset which consists of 200 images from the validation split of COCO-2017, with model predictions generated by an out-of-the-box Faster R-CNN model from [torchvision.models](https://pytorch.org/docs/stable/torchvision/models.html). ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("quickstart") 5session = fo.launch_app(dataset) ``` https://www.youtube.com/watch?v=7mmH-ql\_-zg Alternatively, you can also easily load popular datasets you may already be familiar with from the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html). For example ActivityNet, COCO, ImageNet, Kinetics, Open Images, and more. Any of these datasets can be loaded (and downloaded if necessary) using the `load_zoo_dataset()` command. Check out all the [available datasets](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html) in the FiftyOne Docs. ## Downloading annotated videos stored locally or remotely Community Slack member Patrick asked, _“Is it possible to download an annotated version of a video from FiftyOne? I have a dataset I created with many videos and their associated object detection predictions. How can I download this annotated video locally?”_ FiftyOne supports the annotation of datasets and views via the [`draw_labels`](https://docs.voxel51.com/user_guide/draw_labels.html) method in the SDK, and support for this directly in the App is coming very soon! You can use the [`export`](https://docs.voxel51.com/api/fiftyone.core.dataset.html?highlight=dataset.export#fiftyone.core.dataset.Dataset.export) method on your dataset to download your annotated videos. You can perform exports with this method by following the basic patterns detailed below: - Provide `export_dir` and `dataset_type` to export the content to a directory in the default layout for the specified format, as documented on [this page](https://docs.voxel51.com/user_guide/export_datasets.html#exporting-datasets) - Provide `dataset_type` along with `data_path`, `labels_path`, and/or `export_media` to directly specify where to export the source media and/or labels (if applicable) in your desired format; this syntax provides the flexibility to, for example, perform workflows like labels-only exports - Provide a `dataset_exporter` to which to feed samples to perform a fully-customized export If the dataset is local, pass in your local directory. If the dataset is remote, pass in the directory of your remote machine and then download the video. Here’s an example: ```python 1import fiftyone as fo 2 3# The Dataset or DatasetView containing the samples you wish to export 4dataset_or_view = fo.Dataset(...) 5 6# The directory to which to write the exported dataset 7export_dir = "/path/for/export" 8 9# The name of the sample field containing the label that you wish to export 10# Used when exporting labeled datasets (e.g., classification or detection) 11label_field = "ground_truth" # for example 12 13# The type of dataset to export 14# Any subclass of `fiftyone.types.Dataset` is supported 15dataset_type = fo.types.COCODetectionDataset # for example 16 17# Export the dataset 18dataset_or_view.export( 19 export_dir=export_dir, 20 dataset_type=dataset_type, 21 label_field=label_field, 22) ``` ## Deleting a batch of samples using a DatasetView Community Slack member Dan asked, _“How can I delete a group of samples that I have selected with a view?”_ You can easily remove a batch of samples from a `Dataset` by constructing a `DatasetView` that contains the samples, and then deleting them from the dataset as follows: ```python 1# Choose 10 samples at random 2unlucky_samples = dataset.take(10) 3 4dataset.delete_samples(unlucky_samples) ``` For more information on [removing a batch of samples from a dataset](https://docs.voxel51.com/user_guide/using_views.html#removing-a-batch-of-samples-from-a-dataset), check out the FiftyOne Docs. ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,600+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 3,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 4,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 266+ repositories - 58+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [embeddings](https://voxel51.com/blog/tag/embeddings) [exporting](https://voxel51.com/blog/tag/exporting) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne App](https://voxel51.com/blog/tag/fiftyone-app) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [FiftyOne Computer Vision Tips and Tricks – Mar 24, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 25, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-mar-24-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/03107d477b7db4be03031293fa4fe15aaea806f0-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Sept 16, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 17, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-sept-16-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-291-lllmstxt|> ## Computer Vision Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup — May 25, 2023 May 26, 2023 • 5 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#09fb1b80d236) [Applying Computer Vision to Real Estate at Opendoor](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#ef45ded14ac9) [YOLO-NAS - SOTA Object Detection Generated by NAS](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#dcb4cbb58549) [Wildlife Watcher: A Smart Wildlife Camera](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#94af6a5aa565) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#5593c23ebde7) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#f73f5cd4d308) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#3bd9cdd7cd57) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#09fb1b80d236) [Applying Computer Vision to Real Estate at Opendoor](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#ef45ded14ac9) [YOLO-NAS - SOTA Object Detection Generated by NAS](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#dcb4cbb58549) [Wildlife Watcher: A Smart Wildlife Camera](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#94af6a5aa565) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#5593c23ebde7) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#f73f5cd4d308) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023#3bd9cdd7cd57) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We just wrapped up the May 25, 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/), and if you missed it or want to revisit it, here’s a recap! In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes this month was Wildlife AI! We were first introduced to Wildlife AI through the FiftyOne community. They are using FiftyOne to enable their users to easily analyze the camera data and create their own models. We loved their use case so much, we invited them to present, and as it turns out, they said yes and Victor Anton, Founder and CEO of Wildlife.ai, spoke at this very Meetup! We are sending this event’s charitable donation of $200 to Wildlife AI on behalf of the computer vision community. ![](https://cdn.sanity.io/images/h6toihm1/production/e80959f169a6a89e3cef73f9bb1bdd1750499fe7-770x146.png?auto=format&dpr=2&fit=max&q=75&w=770) Missed the Meetup? No problem. Here are playbacks and talk abstracts from the event. ## Applying Computer Vision to Real Estate at Opendoor https://www.youtube.com/watch?v=BSkhgJ4invc There are many applications for computer vision in the real estate sector. Real estate is one of the few industries which still uses tools from the 2000s and can stand to benefit greatly from the advances in machine learning. In this talk, we will deep dive into Enricher - the computer vision pipelines that help [Opendoor](https://www.opendoor.com/) conduct virtual assessments of homes across the country at scale. Also discussed were some high level computer vision projects being worked on at Opendoor. [Shashwat Srivastava](https://www.linkedin.com/in/shashsrivastava/) is a Senior Engineer on the Data Platform team at Opendoor. Opendoor is a real estate company which greatly simplifies the process of buying and selling your home. Over the 5 years, Shashwat worked on Data and Analytics at Opendoor on everything from data infrastructure to ETL pipelines to building backend services for real time serving of data. Before this role, he was a backend engineer at AppDynamics which he joined right after his undergrad at Carnegie Mellon. Learn more about Opendoor’s computer vision work at the [Opendoor Engineering and Data Science](https://medium.com/opendoor-labs) blog. Q&A from the talk included: - Do you use continuous training? - Are there issues with the time delay between getting the video and the estimate on the cost of repairs? - How do you monitor the results of your models? - What was the rationale for doing a bytes conversion for Spark instead of other formats, say linearised or numpy? - Is the home inspection done completely through computer vision or is having a human in loop still required? - What about insulation and plumbing work? How do you inspect that? ## YOLO-NAS - SOTA Object Detection Generated by NAS https://www.youtube.com/watch?v=\_Nc9z3S0X64 [Deci.ai](https://deci.ai/) is thrilled to announce the release of a new object detection model, [YOLO-NAS](https://deci.ai/blog/yolo-nas-object-detection-foundation-model/). This model is a game-changer in the world of object detection, providing superior real-time object detection capabilities and production-ready performance. Eugene will discuss the new model, and the hardware-aware NAS approach. [Eugene Khvedchenia](https://www.linkedin.com/in/cvtalks/) is a Deep Learning Engineer at Deci AI. Q&A from the talk included: - In knowledge distillation, how are the teacher network's weights initialized and maintained? - Is NAS applied only in the backbone? or in the decoder as well? - Why is YOLO-NAS considered to be part of the YOLO family? - YOLOv8 is licensed with limitations for commercial use. YOLO-NAS is Apache 2.0. What are the distribution limitations of YOLO-NAS? - What kind of augmentations were tried in the training process? - Did you consider pre-training with another dataset before training with COCO? ## Wildlife Watcher: A Smart Wildlife Camera https://www.youtube.com/watch?v=exuIqF1U4Qw Wildlife conservation is more critical now than ever, and monitoring biodiversity is key to protecting our planet. [Wildlife.ai](https://wildlife.ai/), an environmental non-profit, is using cutting-edge technology like artificial intelligence, computer vision and community collaboration to help biodiversity conservation. Their [Wildlife Watchers](https://www.wildlife.ai/projects/weta-watcher/) program combines open-source, low-powered cameras with user-friendly software to make conservation accessible to everyone. [Victor Anton](https://www.linkedin.com/in/victor-anton-116b45188/) is Founder and CEO of Wildlife.ai, Victor is a wildlife biologist who works with research institutions, governments and communities around the world to translate data into improved conservation practices. Q&A from the talk included: - Depending on where you place your Wildlife Watcher, there would be more or less animal traffic through it. Given this, how do you deal with this for reproducibility? - Has tinyML been useful for any of these projects? - Can your camera be powered via a solar panel? - Do you record and process sound, as well? ## Join the Computer Vision Meetup! \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Computer Vision Meetup membership has grown to over [4,000 members](https://www.meetup.com/pro/computer-vision-meetups/) in just under a year! The goal of the Meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. Join one of the 13 Meetup locations closest to your timezone. - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) ## What’s Next? We have exciting speakers already signed up over the next few months! Become a member of the [Computer Vision Meetup closest to you](https://www.meetup.com/pro/computer-vision-meetups/), then register for the Zoom. Up next on June 8 at 10 AM Pacific we have the US and EMEA timezone-friendly Computer Vision Meetup happening with talks including: - **Redefining State-of-the-Art with YOLOv5 and YOLOv8** \- Glenn Jocher (Ultralytics) - **Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation** \- Narek Tumanyan & Michal Geyer (Weizmann Institute of Science) - **Re-annotating MS COCO, An Exploration of Pixel Tolerance** \- Jerome Pasquero & Eric Zimmermann (Sama) Register for the Zoom here. You can find a complete schedule of upcoming Meetups on [the Voxel51 Events page](https://voxel51.com/computer-vision-events/). ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. — _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [Deci AI](https://voxel51.com/blog/tag/deci-ai) [object detection](https://voxel51.com/blog/tag/object-detection) [Opendoor](https://voxel51.com/blog/tag/opendoor) [real estate](https://voxel51.com/blog/tag/real-estate) [Wildlife AI](https://voxel51.com/blog/tag/wildlife-ai) [YOLO-NAS](https://voxel51.com/blog/tag/yolo-nas) ![](https://cdn.sanity.io/images/h6toihm1/production/b447c3f47d7e0ddcf4272c2034a8431fc05a809f-300x300.png?auto=format&dpr=2&fit=max&q=75&w=42) Jimmy Guerrero Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/047b21a97f6c858334f9f35ed89fa7655ebf5767-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ State-of-the-Art Object Detection with YOLO-NAS & FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ May 4, 2023](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/689c84b752e280cf575981e56bcbec52695d882f-512x288.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — Oct 12, 2023\\ \\ Computer Vision, Event Recaps\\ \\ • \\ \\ Oct 13, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-oct-12-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/3262b9c1e8ad8f4d6eea690685cb83472fe398de-960x540.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the AI, Machine Learning and Data Science Meetup — Nov 2, 2023\\ \\ Computer Vision, Event Recaps\\ \\ • \\ \\ Nov 3, 2023](https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-nov-2-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-292-lllmstxt|> ## FiftyOne 0.21 Release [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Announcing FiftyOne 0.21 with Operators, Dynamic Groups, and Custom Color Schemes Jun 1, 2023 • 8 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-fiftyone-0-21#9f25966936c3) [tl;dr: What’s new in FiftyOne 0.21?](https://voxel51.com/blog/announcing-fiftyone-0-21#372f00e06605) [See a recap of the live demo and AMA held on June 15](https://voxel51.com/blog/announcing-fiftyone-0-21#3323edefcb4a) [Operators](https://voxel51.com/blog/announcing-fiftyone-0-21#1f6506d51c04) [Dynamic groups](https://voxel51.com/blog/announcing-fiftyone-0-21#bac89d2f877d) [Custom color schemes](https://voxel51.com/blog/announcing-fiftyone-0-21#c8c6a6eca2ad) [Custom field visibility](https://voxel51.com/blog/announcing-fiftyone-0-21#6d3747ebd1ab) [Multiple point cloud overlays](https://voxel51.com/blog/announcing-fiftyone-0-21#7fb7977a2f41) [Community contributions](https://voxel51.com/blog/announcing-fiftyone-0-21#f5ff5f907357) [FiftyOne community updates](https://voxel51.com/blog/announcing-fiftyone-0-21#c53331a5d5f8) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-fiftyone-0-21#9f25966936c3) [tl;dr: What’s new in FiftyOne 0.21?](https://voxel51.com/blog/announcing-fiftyone-0-21#372f00e06605) [See a recap of the live demo and AMA held on June 15](https://voxel51.com/blog/announcing-fiftyone-0-21#3323edefcb4a) [Operators](https://voxel51.com/blog/announcing-fiftyone-0-21#1f6506d51c04) [Dynamic groups](https://voxel51.com/blog/announcing-fiftyone-0-21#bac89d2f877d) [Custom color schemes](https://voxel51.com/blog/announcing-fiftyone-0-21#c8c6a6eca2ad) [Custom field visibility](https://voxel51.com/blog/announcing-fiftyone-0-21#6d3747ebd1ab) [Multiple point cloud overlays](https://voxel51.com/blog/announcing-fiftyone-0-21#7fb7977a2f41) [Community contributions](https://voxel51.com/blog/announcing-fiftyone-0-21#f5ff5f907357) [FiftyOne community updates](https://voxel51.com/blog/announcing-fiftyone-0-21#c53331a5d5f8) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Voxel51 in conjunction with the FiftyOne community is excited to announce the general availability of [FiftyOne 0.21](https://docs.voxel51.com/release-notes.html#fiftyone-0-21-0). This release is packed with upgrades to FiftyOne’s plugin framework and a variety of App features that unlock new ways to explore your datasets and customize your App experience to suit your needs. How? Read on! ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is the open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. ![](https://cdn.sanity.io/images/h6toihm1/production/315bf223808d036ae077ae83e7953f6560ed3e7f-1491x759.gif?auto=format&dpr=2&fit=max&q=75&w=1491) - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://slack.voxel51.com/), we’re always happy to help. Ok, let’s dive into the release. ## tl;dr: What’s new in FiftyOne 0.21? This release includes: - **Operators**: we added a new concept called Operators to FiftyOne’s Plugin framework that you can use to trigger arbitrary custom functionality directly from the App via a powerful custom form language, all written in Python(!) - **Dynamic groups:** you can now dynamically group the samples in your dataset (in the App or Python) based on any field or expression of interest, allowing you to both quickly navigate _between_ groups and explore variations _within_ a particular group - **Custom color schemes:** you now have full control to customize the field and label colors that are used to render content in the App! - **Custom field visibility:** there’s a new settings menu in the App sidebar that you can use to visually configure which sample fields appear in the App - **Multiple point cloud overlays**: you can now overlay multiple point clouds into the same scene in the App Check out the [release notes](https://voxel51.com/docs/fiftyone/release-notes.html#fiftyone-0-21-0) for a full rundown of additional enhancements and bugfixes in FiftyOne 0.21. By the way, FiftyOne Teams 1.3 is also generally available! Check out its [release notes](https://docs.voxel51.com/release-notes.html#fiftyone-teams-1-3-0) to explore the new collaboration-minded features we rolled out to commercial users (remember, Teams is fully compatible with existing FiftyOne workflows). ## See a recap of the live demo and AMA held on June 15 \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Want to see FiftyOne 0.21 in action, along with VoxelGPT, your new AI assistant for computer vision? Check out the [webinar recap video and blog post](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins/). You'll get a taste of all the amazing capabilities you can build using plugins and operators, plus all the capabilities unlocked by VoxelGPT (which is a plugin)! Now, here’s a quick overview of some of the new features we packed into this release. ## Operators FiftyOne 0.21 adds a powerful new feature called Operators to [FiftyOne’s Plugin framework](https://docs.voxel51.com/plugins/index.html) that allows you to trigger arbitrary _custom_ functionality directly from the App via custom forms, all written in Python. For example, wish you could rename a sample field from within the App? Now you can add it yourself! ![](https://cdn.sanity.io/images/h6toihm1/production/78b04bcec65d48a6c72a29a6f105bf22efc4b46a-916x670.gif?auto=format&dpr=2&fit=max&q=75&w=916) And here are the few lines of Python code required to do it, from soup to nuts: ```python 1import fiftyone.operators as foo 2import fiftyone.operators.types as types 3 4class RenameSampleField(foo.Operator): 5 @property 6 def config(self): 7 return foo.OperatorConfig( 8 name="rename_sample_field", 9 label="Rename sample field", 10 dynamic=True, 11 ) 12 13 def resolve_input(self, ctx): 14 inputs = types.Object() 15 fields = ctx.dataset.get_field_schema(flat=True) 16 field_keys = list(fields.keys()) 17 field_selector = types.AutocompleteView() 18 for key in field_keys: 19 field_selector.add_choice(key, label=key) 20 21 inputs.enum( 22 "field_name", 23 field_keys, 24 label="Field to rename", 25 view=field_selector, 26 required=True, 27 ) 28 field_name = ctx.params.get("field_name", None) 29 new_field_name = ctx.params.get("new_field_name", None) 30 if field_name and field_name in field_keys: 31 new_field_prop = inputs.str( 32 "new_field_name", 33 required=True, 34 label="New field name", 35 default=f"{field_name}_copy", 36 ) 37 if new_field_name and new_field_name in field_keys: 38 new_field_prop.invalid = True 39 new_field_prop.error_message = ( 40 f"Field '{new_field_name}' already exists" 41 ) 42 inputs.str( 43 "error", 44 label="Error", 45 view=types.Error( 46 label="Field already exists", 47 description=f"Field '{new_field_name}' already exists", 48 ), 49 ) 50 51 return types.Property( 52 inputs, view=types.View(label="Rename sample field") 53 ) 54 55 def execute(self, ctx): 56 ctx.dataset.rename_sample_field( 57 ctx.params["field_name"], 58 ctx.params["new_field_name"], 59 ) 60 ctx.trigger("reload_dataset") 61 62def register(p) 63 p.register(RenameSampleField) ``` In the above code: - `config` defines the name and label (display name) of the operator, which you can access by clicking the Browse operations button in the App as shown in the GIF above - `resolve_input()` defines the form that accepts user input in the App. FiftyOne provides a rich set of prebuilt types (dropdowns, selectors, etc) that you can use to build sophisticated custom forms - `execute()` defines the operation to perform when the user invokes it - `ctx` is a context object that provides access to the current state of the App, including things like form inputs, the current dataset, the current view, whether any samples are currently selected, etc., that you can use to implement the desired action Operators represent bite-sized pieces of functionality that can be executed either directly by users or, importantly, chained together for internal use to build up more complex sequences of instructions. For example, the `ctx.trigger("reload_dataset")` statement above demonstrates how operators can invoke other operators. In this case a builtin `reload_dataset` operator is triggered that causes the dataset to be reloaded in the App to pull in the schema changes that our custom Operator performed. How do Operators relate to Plugins? FiftyOne Plugins _contain_ zero or more Operators, which are automatically added to the Operator browser in the App whenever you install the plugin. Refer to [the docs](https://docs.voxel51.com/plugins/index.html#how-to-write-plugins) for more information about writing your own Operators and Plugins and [distributing](https://docs.voxel51.com/plugins/index.html#publishing-plugins) them via GitHub. Want to get started with Operators without writing any code? Check out the [fiftyone-plugins](https://github.com/voxel51/fiftyone-plugins) repository on GitHub to explore a growing list of plugins that you can easily download and install [via the CLI](https://docs.voxel51.com/cli/index.html#fiftyone-plugins): ```python 1# Download plugin(s) from a GitHub repository 2fiftyone plugins download https://github.com//[/tree/branch] 3 4# Install any requirements required to use the plugin 5fiftyone plugins requirements -–install 6 7# List the plugins that you’ve downloaded locally 8fiftyone plugins list 9 10# List the available operators that you’ll be able to run 11fiftyone operators list ``` ## Dynamic groups FiftyOne 0.21 adds a new data exploration feature we’re calling _dynamic groups_. With [dynamic groups](https://docs.voxel51.com/user_guide/using_views.html#grouping) you can organize the samples in your dataset by a particular field or expression. For example, you can group a classification dataset like CIFAR-10 by label: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("cifar10", split="test") 5 6# Take 100 samples and group by ground truth label 7view = dataset.take(100, seed=51).group_by("ground_truth.label") 8 9print(view.media_type) # group 10print(len(view)) # 10 ``` which yields a new [view](https://docs.voxel51.com/user_guide/using_views.html#) that contains 10 groups, one per each object class. When you iterate over a dynamic grouped view, you get one example from each group. Like any other view, you can chain additional view stages to further refine the view’s contents: ```python 1# Sort the groups by label 2sorted_view = view.sort_by("ground_truth.label") 3 4for sample in sorted_view: 5 print(sample.ground_truth.label) ``` ```python 1airplane 2automobile 3bird 4cat 5deer 6dog 7frog 8horse 9ship 10truck ``` In addition, you can use [get\_dynamic\_group()](https://docs.voxel51.com/api/fiftyone.core.view.html#fiftyone.core.view.DatasetView.get_dynamic_group) to retrieve a view containing all samples _within_ a specific group: ```python 1group = view.get_dynamic_group("horse") 2print(len(group)) # 11 ``` You can use the new group action in the App’s menu to dynamically group your samples by a field of your choice directly from within the App: ![](https://cdn.sanity.io/images/h6toihm1/production/5088ffe569c040910f5705814e89e4cb5dbf5dc4-1209x690.gif?auto=format&dpr=2&fit=max&q=75&w=1209) In this mode, the App’s grid shows the first sample from each group, and you can click on a sample to view all elements of the group in the modal. By default, the samples in each group are unordered, but you can provide the optional order\_by and reverse arguments to [group\_by()](https://docs.voxel51.com/api/fiftyone.core.collections.html#fiftyone.core.collections.SampleCollection.group_by) to specify an ordering for the samples in each group: ```python 1# Create an image dataset that contains one sample per frame of the 2# `quickstart-video` dataset 3dataset = ( 4 foz.load_zoo_dataset("quickstart-video") 5 .to_frames(sample_frames=True) 6 .clone() 7) 8 9print(len(dataset)) # 1279 10 11# Group by video ID and order each group by frame number 12view = dataset.group_by("sample_id", order_by="frame_number") 13 14print(len(view)) # 10 15print(view.values("frame_number")) 16# [1, 1, 1, ..., 1] 17 18sample_id = dataset.take(1).first().sample_id 19video = view.get_dynamic_group(sample_id) 20 21print(video.values("frame_number")) 22# [1, 2, 3, ..., 120] ``` And as usual, you can achieve the same result directly in the App: ![](https://cdn.sanity.io/images/h6toihm1/production/7955c059575605e4c50034a0d5d171314470ceba-1207x687.gif?auto=format&dpr=2&fit=max&q=75&w=1207) When viewing ordered groups in the App, the modal shows a pagination UI at the bottom that you can use to navigate sequentially or via random access through the elements of the group. Bonus: you can even group by arbitrary expressions! ```python 1from fiftyone import ViewField as F 2 3dataset = foz.load_zoo_dataset("quickstart") 4 5# Group samples by the number of ground truth objects they contain 6expr = F("ground_truth.detections").length() 7view = dataset.group_by(expr) 8 9print(len(view)) # 26 10print(len(dataset.distinct(expr))) # 26 ``` Check out [the docs](https://docs.voxel51.com/user_guide/using_views.html#grouping) for more information about creating dynamic grouped views in Python and the App. ## Custom color schemes In FiftyOne 0.21 you can now fully customize the color scheme used by the App to render content in the grid/modal! The simplest way to get started is to click on the new [color palette icon](https://docs.voxel51.com/user_guide/app.html#color-schemes-in-the-app) above the sample grid. The GIF below demonstrates how to: - Configure a custom color pool from which to draw colors for otherwise unspecified fields/values - Configure the colors assigned to specific fields in color by field mode - Configure the colors used to render specific annotations based on their attributes in color by value mode - Save the customized color scheme as the default for the dataset ![](https://cdn.sanity.io/images/h6toihm1/production/09a4bc867e748e355ac7d2a074ee98f833610412-914x703.gif?auto=format&dpr=2&fit=max&q=75&w=914) Note that any customizations you make only apply to the current dataset. Each time you load a new dataset, the color scheme will revert to that dataset’s default color scheme (if any) or else the global default color scheme. You can press `Save as default` to save your current color scheme as the dataset’s default scheme. As you’d expect from FiftyOne, you can also programmatically define your color scheme [through Python](https://docs.voxel51.com/user_guide/app.html#color-schemes-in-python)! For example, the code below configures a custom color scheme suitable for analyzing evaluation results where: - All true positive objects are **green** - All false negatives objects are **blue** - All false positives objects are **red** ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("quickstart") 5dataset.evaluate_detections( 6 "predictions", gt_field="ground_truth", eval_key="eval" 7) 8 9# Create a custom color scheme 10fo.ColorScheme( 11 fields=[\ 12 {\ 13 "path": "ground_truth",\ 14 "colorByAttribute": "eval",\ 15 "valueColors": [\ 16 {"value": "fn", "color": "#0000ff"}, # false negatives: blue\ 17 {"value": "tp", "color": "#00ff00"}, # true positives: green\ 18 ]\ 19 },\ 20 {\ 21 "path": "predictions",\ 22 "colorByAttribute": "eval",\ 23 "valueColors": [\ 24 {"value": "fp", "color": "#ff0000"}, # false positives: red\ 25 {"value": "tp", "color": "#00ff00"}, # true positives: green\ 26 ]\ 27 }\ 28 ] 29) 30 31# Option 1: launch App with a custom color scheme 32session = fo.launch_app(dataset, color_scheme=color_scheme) 33 34# Option 2: store the color scheme as default for the dataset 35dataset.app_config.color_scheme = color_scheme 36dataset.save() ``` In the above example, you can see TP/FP/FN colors in the App by clicking on the Color palette icon and switching to `color by value` mode. Check out [the docs](https://docs.voxel51.com/user_guide/app.html#color-schemes) for more information about creating and saving custom App color schemes. ## Custom field visibility Do you work with datasets that contain many fields and find your App is getting cluttered? FiftyOne 0.21 is here to help! There’s a new settings menu in the App sidebar that you can use to visually configure which sample fields appear in the App: ![](https://cdn.sanity.io/images/h6toihm1/production/7e02a671baa84e3b9fa6a4955bf93664f0b47565-914x686.gif?auto=format&dpr=2&fit=max&q=75&w=914) In the simplest case, you can use the checkboxes to manually control which fields appear in the sidebar. By default, only top-level fields are available for selection, but if you want fine-grained control you can opt to include nested fields (e.g. [dynamic attributes](https://docs.voxel51.com/user_guide/using_datasets.html#dynamic-attributes) of your label fields) in the selection list as well. Note that, after a field visibility change is applied, a filter icon appears to the left of the settings icon in the sidebar indicating how many fields are currently excluded. You can reset your selection by clicking this icon or reopening the modal and pressing the reset button at the bottom. What if you want to switch between multiple contexts, each containing different sets of fields? You can persist and reload different field selections by [saving views](https://docs.voxel51.com/user_guide/app.html#app-saving-views)! For more advanced use cases, you can use the [Filter rule tab](https://docs.voxel51.com/user_guide/app.html#filter-rules) to define a rule that is _dynamically_ applied to the dataset’s [field metadata](https://docs.voxel51.com/user_guide/using_datasets.html#storing-field-metadata) each time the App loads to determine which fields to include in the sidebar: ![](https://cdn.sanity.io/images/h6toihm1/production/2701e695a7d7d7c4a1072c6313f0e8eeba6446ab-1494x1298.jpg?auto=format&dpr=2&fit=max&q=75&w=1494) ## **Multiple point cloud overlays** Last but not least, FiftyOne 0.21 includes a heavily requested feature by users that work with [grouped datasets](https://docs.voxel51.com/user_guide/groups.html) that contain multiple point cloud slices: you can now overlay multiple point clouds in the App’s 3D Visualizer! ![](https://cdn.sanity.io/images/h6toihm1/production/e1fd1039e37398073314736814885ffaeff91151-1208x689.gif?auto=format&dpr=2&fit=max&q=75&w=1208) ## Community contributions We’re excited to announce that we’ve partnered with the amazing team at [Sama](https://www.sama.com/) to make the [Sama-Coco dataset](https://www.sama.com/sama-coco-dataset/) available for download in the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#sama-coco)! As usual, downloading all or partial subsets of Sama-Coco from the zoo is easy: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# Load 50 random samples from the validation split 5dataset = foz.load_zoo_dataset( 6 "sama-coco", 7 split="validation", 8 max_samples=50, 9 shuffle=True, 10) 11 12# Or download the entire dataset 13dataset = foz.load_zoo_dataset("sama-coco") ``` Sama-Coco is a relabeling of the COCO-2017 dataset, an industry standard benchmark dataset for large-scale object detection and segmentation. Among other improvements ( [read more here](https://www.sama.com/sama-coco-dataset/)), masks in Sama-Coco are tighter and many crowd instances have been decomposed into their components. The project was undertaken as part of Sama’s ongoing effort to [redefine data quality for the modern age](https://www.sama.com/blog/redefining-data-quality/) and to contribute to the wider research and development efforts of the ML community. ![](https://cdn.sanity.io/images/h6toihm1/production/749bfe651f06ed37d0a2c872d413e42b04955cfd-1152x291.jpg?auto=format&dpr=2&fit=max&q=75&w=1152) Here’s what Jerome Pasquero, Principal Product Manager at Sama, had to say about the collaboration: > _“While re-annotating Sama-Coco, we sensed that we were working on something remarkable. However, when we had the opportunity to visualize and explore it in FiftyOne, we truly grasped its potential impact. The ability to easily compare the COCO and Sama-Coco datasets at scale allowed us to see the influence the re-annotations could have on machine learning. We are excited to partner with Voxel51 and make it available for download directly from the FiftyOne App as part of their 0.21 release.”_ Also, shoutout to the following community members who contributed to FiftyOne 0.21: - [Jonathan Badger](https://github.com/jbadger3) contributed [#2706 - Labelstudio instances support](https://github.com/voxel51/fiftyone/pull/2706) - [Akshit Priyesh](https://github.com/akshitpriyesh) contributed [#2776 - Adding support for evaluating keypoints](https://github.com/voxel51/fiftyone/pull/2776) - [Rustem Galiullin](https://github.com/Rusteam) contributed [#2777 - add optional dice score computation for evaluate\_segmentations](https://github.com/voxel51/fiftyone/pull/2777) - [Felix Nobis](https://github.com/fnobis) contributed [#3117 - multiclass classification support for scale.com labeling](https://github.com/voxel51/fiftyone/pull/3117) - [Colin Wong](https://github.com/democat3457) contributed [#3127 - Only match .txt files when reading YOLO labels](https://github.com/voxel51/fiftyone/pull/3127) ## FiftyOne community updates The FiftyOne community continues to grow! - 1,700+ [FiftyOne Slack](https://slack.voxel51.com/) members - 3,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 4,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 300+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [color schemes](https://voxel51.com/blog/tag/color-schemes) [custom plugins](https://voxel51.com/blog/tag/custom-plugins) [dynamic groups](https://voxel51.com/blog/tag/dynamic-groups) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne 0.21](https://voxel51.com/blog/tag/fiftyone-0-21) [operators](https://voxel51.com/blog/tag/operators) [plugins](https://voxel51.com/blog/tag/plugins) [product release](https://voxel51.com/blog/tag/product-release) ![](https://cdn.sanity.io/images/h6toihm1/production/8d61ff90b31d151405f9e21a33c2802509f34651-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Brian Moore Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/b847f6d923a2c98ff957ee27e8b18e41c3fcc0ac-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Announcing FiftyOne Teams 1.4 with Dataset Versioning, Delegated Operations, and Ultralytics Integration\\ \\ Product & News\\ \\ • \\ \\ Sep 20, 2023](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration) [![](https://cdn.sanity.io/images/h6toihm1/production/a4c2bee9ed053c5be2a1c161e5abf758c9a12ff8-1400x923.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.18 with App Performance Improvements, Sidebar Modes, and Custom Attributes\\ \\ Product & News\\ \\ • \\ \\ Nov 15, 2022](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes) [![](https://cdn.sanity.io/images/h6toihm1/production/e94f20fa81716294c7a6caccf7256e9106cb6e89-967x800.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.17 with Grouped Datasets, 3D, Geolocation, and Custom Plugins\\ \\ Product & News\\ \\ • \\ \\ Sep 21, 2022](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-293-lllmstxt|> ## FiftyOne Community Update [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) FiftyOne Computer Vision Community Update – June ‘23 Jun 2, 2023 • 5 min read Article content In this article [Community Spotlight](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#533d0fd9c185) [See More Stories](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#018ab023fd70) [Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#799e4fb011d5) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#13eeea83d4b5) [Community Contributions](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#02e4561eabd6) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#7d1891535181) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#aad721e39846) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#ce4b5ba1a0a7) [Upcoming Computer Vision Events](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#e62636546417) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#c682dbe2d147) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#49d8f7eab89c) In this article [Community Spotlight](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#533d0fd9c185) [See More Stories](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#018ab023fd70) [Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#799e4fb011d5) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#13eeea83d4b5) [Community Contributions](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#02e4561eabd6) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#7d1891535181) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#aad721e39846) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#ce4b5ba1a0a7) [Upcoming Computer Vision Events](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#e62636546417) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#c682dbe2d147) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023#49d8f7eab89c) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the monthly blog series where we bring you up to speed on recent happenings in the FiftyOne community and celebrate noteworthy milestones. 🙌 🚀 ## Community Spotlight We love hearing how FiftyOne helps you solve challenges and reach new heights! Curious what sorts of use cases are possible with [FiftyOne](https://voxel51.com/fiftyone/)? Check out this story from a member of the FiftyOne community. ### METU (Middle East Technical University) - Research & Medical Imaging ![](https://cdn.sanity.io/images/h6toihm1/production/a794207311f6472c445350b07aa67966954982fd-900x400.png?auto=format&dpr=2&fit=max&q=75&w=900) [METU](https://www.metu.edu.tr/) collaborates with the Department of Gastroenterology at Marmara University, Turkey to conduct medical research on inflammatory bowel disease (IBD). One of the main challenges in IBD diagnosis is the high interobserver and intraobserver variability among the assessments of the doctors. With the aid of deep learning, real-time and objective feedback can be provided to doctors during colonoscopy operations. METU uses FiftyOne at various stages of research. First, they make use of FiftyOne's CVAT integration to request and load annotations from doctors. Next, they compare doctors' annotations and their models' predictions. Then, using FiftyOne Brain's embedding space visualization tools (t-SNE specifically), they get a better understanding of their data distribution. In a branch of this research, they produce synthetic images with GANs and again, refer to FiftyOne Brain tools for help identifying the most informative/unique images by eliminating "near duplicates". In the coming months, doctors will start sending videos and METU will start to make use of FiftyOne's video analysis capabilities. ## See More Stories [See more stories](https://voxel51.com/success-stories/) from people and organizations building remarkable machine learning and AI using FiftyOne and FiftyOne Teams. ## Community Rewards ![](https://cdn.sanity.io/images/h6toihm1/production/e3ed352e497d0e500ff4a1484b8422b3c9bef5cb-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) Is your organization using FiftyOne to solve interesting computer vision problems? [Share your success story](https://voxel51.com/fiftyone-computer-vision-success-story-submission/) and claim a box of community rewards as a thank you! ## Product Releases ### FiftyOne 0.21 Is Here! Today, we [announced FiftyOne 0.21](https://voxel51.com/blog/announcing-fiftyone-0-21/). This release is packed with upgrades to FiftyOne’s plugin framework and a variety of App features that unlock new ways to explore your datasets and customize your App experience to suit your needs. You can dive into the details of what’s included in the [0.21 release notes](https://docs.voxel51.com/release-notes.html#fiftyone-0-21-0). Also, stop by our webinar on June 15 at 10 AM Pacific Time. We’ll give you a taste of what you can build using plugins and operators followed by an open Q&A where you can get answers to any questions you might have. [Register here.](https://voxel51.com/computer-vision-events/build-fiftyone-plugins-operators/?utm_source=blog) ### FiftyOne Teams 1.3 Is Also Generally Available! FiftyOne Teams 1.3 is also generally available! Check out its [release notes](https://docs.voxel51.com/release-notes.html#fiftyone-teams-1-3-0) to explore the new collaboration-minded features we rolled out to commercial users (remember, Teams is fully compatible with existing FiftyOne workflows). ## **Community Contributions** A quick shoutout to the following community members who made contributions to the FiftyOne project with the v0.21.0 release. - [Jonathan Badger](https://github.com/jbadger3) contributed [#2706 - Labelstudio instances support](https://github.com/voxel51/fiftyone/pull/2706) - [Akshit Priyesh](https://github.com/akshitpriyesh) contributed [#2776 - Adding support for evaluating keypoints](https://github.com/voxel51/fiftyone/pull/2776) - [Rustem Galiullin](https://github.com/Rusteam) contributed [#2777 - add optional dice score computation for evaluate\_segmentations](https://github.com/voxel51/fiftyone/pull/2777) - [Felix Nobis](https://github.com/fnobis) contributed [#3117 - multiclass classification support for scale.com labeling](https://github.com/voxel51/fiftyone/pull/3117) - [Colin Wong](https://github.com/democat3457) contributed [#3127 - Only match .txt files when reading YOLO labels](https://github.com/voxel51/fiftyone/pull/3127) ## FiftyOne on GitHub GitHub is home to the open source FiftyOne project. Here’s the latest snapshot of what’s happening in the [FiftyOne GitHub repo](https://github.com/voxel51/fiftyone): - Total stars: 3,000+ - Total contributors: 63 - Total used by: 306 repositories - Total forks: 352 ## FiftyOne Community Slack The FiftyOne Community [Slack channel](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) is where you can join more than 1550 machine learning engineers and data scientists using FiftyOne to improve the quality of their computer vision data and build better models. Last month alone we had 55 first time community members. Ask questions, answer questions, or simply follow along with the discussion! To make it easy to catch the highlights, every Friday we recap interesting questions and answers from Slack in [Tips & Tricks blog series](https://voxel51.com/blog/category/tips-tricks/). Recent posts include: - [FiftyOne Computer Vision Tips and Tricks – May 26, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023/) - [FiftyOne Computer Vision Tips and Tricks – May 19, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023/) - [FiftyOne Computer Vision Embeddings Tips and Tricks – May 12, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-12-2023/) - [FiftyOne Computer Vision Tips and Tricks - April 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-21-2023/) ## Computer Vision Meetups \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Voxel51 sponsors 13 virtual [Computer Vision Meetups](https://www.meetup.com/pro/computer-vision-meetups/) around the world. (To join, visit the Meetup [link](https://www.meetup.com/pro/computer-vision-meetups/) and scroll down to find the location friendliest to your time zone.) The Computer Vision Meetups are geared towards data scientists, machine learning engineers, and open source enthusiasts who want to expand their knowledge of computer vision and complementary technologies. We put an emphasis on open source software, and speakers who are computer vision practitioners or academics doing research in the field. This month’s Meetups include: ### June ’23 Computer Vision Meetup (Americas and EMEA) - June 8 – 10AM PDT (1:00PM EDT) - Redefining State-of-the-Art with YOLOv5 and YOLOv8 - [Glenn Jocher](https://www.linkedin.com/in/glenn-jocher/) (Ultralytics) - Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation: [Narek Tumanyan](https://www.linkedin.com/in/narek-tumanyan-5a3334122) & [Michal Geyer](https://www.linkedin.com/in/michal-geyer-aa3a021b1) (Weizmann Institute of Science) - Re-annotating MS COCO, An Exploration of Pixel Tolerance - [Jerome Pasquero](https://www.linkedin.com/in/jeromepasquero/) & [Eric Zimmermann](https://www.linkedin.com/in/eric-zimmermann/) (Sama) - [_Register for the Zoom_](https://voxel51.com/computer-vision-events/june-computer-vision-meetup-2023/?utm_source=blog) ### Recapping the May 24 Meetup If you missed the last Meetup on May 25, make sure to check out [the recap blog](https://voxel51.com/blog/recapping-the-computer-vision-meetup-may-25-2023/) and watch the playbacks! - [Applying Computer Vision to Real Estate at Opendoor](https://youtu.be/BSkhgJ4invc) \- Shashwat Srivastava (Opendoor) - [YOLO-NAS – SOTA Object Detection Generated by NAS](https://youtu.be/_Nc9z3S0X64) \- Eugene Khvedchenia (Deci.ai) - [Wildlife Watcher: A Smart Wildlife Camera](https://youtu.be/exuIqF1U4Qw) \- Victor Anton (Wildlife.ai) ## Upcoming Computer Vision Events In addition to meetups, we invite you to join us for one or more of these upcoming [events](https://voxel51.com/computer-vision-events/): - June 15 - [\[Recap\] How to Build Custom FiftyOne Plugins and Operators](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins/) - June 18-23 - [CVPR 2023 in Vancouver, Canada - Booth 1618](https://cvpr2023.thecvf.com/) - June 28 - [\[Recap\] Getting Started with FiftyOne Workshop (Americas)](https://www.youtube.com/playlist?list=PLuREAXoPgT0SJLKsgFzKxffMApbXp90Gi) - July 13 - Virtual Computer Vision Meetup, Focused on Vector Search ## New Docs, Blogs, Videos, and Tutorials We want everyone to be successful with FiftyOne, and one of the ways we try to do that is by publishing resources that you might find helpful and handy. Here’s a list of some of the new [documentation](https://docs.voxel51.com/), [blogs](https://voxel51.com/blog/), [videos](https://www.youtube.com/@voxel51/videos), [tutorials](https://docs.voxel51.com/tutorials/index.html), [integrations](https://docs.voxel51.com/integrations/index.html), and [cheat sheets](https://docs.voxel51.com/cheat_sheets/index.html) that you may want to check out. ### Blogs - [CVPR 2023 and the State of Computer Vision](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision/) - [CVPR 2023 Survival Guide](https://voxel51.com/blog/cvpr-2023-survival-guide/) ### Videos - [FiftyOne Dataset Zoo: Exploring Google Research's Kaggle Image Matching Challenge 2023 Dataset](https://www.youtube.com/watch?v=iSwBd3xXPUE) - [FiftyOne Dataset Zoo: Visualizing Defects in Amazon's ARMBench Using Embeddings and OpenAI's CLIP](https://www.youtube.com/watch?v=5FTd4aYWr-k) - [What’s New in FiftyOne 0.20 for Your Computer Vision Workflows](https://www.youtube.com/watch?v=DUWfP3tNQSc) ## Voxel51’s Commitment to Open Source and Community Open source, transparency, and giving back to the computer vision community is what we are all about! Whether it’s developing the open source [FiftyOne computer vision toolset](https://github.com/voxel51/fiftyone) to help engineers and data scientists build high-quality datasets and models, sponsoring [Meetups](https://www.meetup.com/pro/computer-vision-meetups/) to help members boost their computer vision knowledge, or [giving to charitable causes](https://voxel51.com/charitable-giving/) on behalf of the community, Voxel51 is committed to bringing transparency and clarity to the world’s data. [community rewards](https://voxel51.com/blog/tag/community-rewards) [Community Update](https://voxel51.com/blog/tag/community-update) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne 0.21](https://voxel51.com/blog/tag/fiftyone-0-21) [OSS community](https://voxel51.com/blog/tag/oss-community) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/338b38d41e6072dd11af86f21f5309337c52f36b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – May ‘23\\ \\ Product & News\\ \\ • \\ \\ May 5, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/e3e0afa63ab1b67d46ae267f97fd9de6139864d2-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – September 2023\\ \\ Product & News\\ \\ • \\ \\ Sep 8, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/2be41b07bd86d7efc5916442ca5b94aa5115234d-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Greatest Hits of 2022: FiftyOne & Voxel51\\ \\ Product & News\\ \\ • \\ \\ Jan 10, 2023](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-294-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – June 2, 2023 Jun 2, 2023 • 3 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#cca857e091b5) [Making fields from duplicates\_view available in exported annotations](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#327f49b92a40) [Deleting label fields](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#1191c99f2340) [Loading datasets on a headless server](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#0e29756c8980) [Efficiently modifying and deleting fields in views](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#aeb460cd38d0) [Accessing samples from Amazon S3](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#fa121c5e9c92) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#e8c3c557da6f) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#cca857e091b5) [Making fields from duplicates\_view available in exported annotations](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#327f49b92a40) [Deleting label fields](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#1191c99f2340) [Loading datasets on a headless server](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#0e29756c8980) [Efficiently modifying and deleting fields in views](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#aeb460cd38d0) [Accessing samples from Amazon S3](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#fa121c5e9c92) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023#e8c3c557da6f) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## Making fields from duplicates\_view available in exported annotations Community Slack member Vinicius asked (and then figured out the solution on his own!): _“Is it possible to make fields such as `dup_type` from the `duplicates_view()` method available in exported annotations?”_ Yes! There are a few ways to accomplish this. Here’s the solution that Vinicius came up with. First, move the information to detection level: ```python 1for sample in dataset.match(F('dup_type').is_in(['duplicate','nearest'])): 2 sample.detections.detections[0]['dup_type'] = sample.dup_type 3 sample.save() ``` Then export the dataset with `include_attributes = True`: ```python 1dataset.export(..., 2dataset_type=fo.types.FiftyOneImageDetectionDataset, 3include_attributes = True) ``` After importing back the dataset, move the attributes to the sample level: ```python 1for dup_type in ['duplicate','nearest']: 2 dataset.filter_labels('detections',F('dup_type')==dup_type).tag_samples(dup_type) ``` Learn more about [duplicate detection](https://docs.voxel51.com/user_guide/brain.html#duplicate-detection) and the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html) in the Docs. ## Deleting label fields Community Slack member Joy asked, _"How can I remove the label field named ‘test?’ in the screenshot below? I've tried `delete_labels` and `remove_dynamic_sample_field`, but neither does the trick."_ ![](https://cdn.sanity.io/images/h6toihm1/production/e98bcb9ab611fb1d371a5035e996b51ea3d80973-272x310.png?auto=format&dpr=2&fit=max&q=75&w=272) You can delete the “test” field from all samples in the dataset using `delete_sample_field(field_name, error_level=0)` The parameters breakdown as such: - **field\_name** – the field name or `embedded.field.name` - **error\_level** (0) – the error level to use. Valid values are: - **0** (-) – raise error if a top-level field cannot be deleted - **1** (-) – log warning if a top-level field cannot be deleted - **2** (-) – ignore top-level fields that cannot be deleted For more information on [delete\_sample\_field](https://docs.voxel51.com/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.delete_sample_field) check out the FiftyOne Docs. UPDATE: This issue was resolved in the [FiftyOne 0.21.0 release](https://docs.voxel51.com/release-notes.html) via the new [built-in operators](https://github.com/voxel51/fiftyone/pull/2679). ## Loading datasets on a headless server Community Slack member Wes asked, _“I’m having some trouble figuring out how to load annotations on a headless server running FiftyOne. Any tips?”_ If your dataset is in a common format, you can use [these instructions](https://docs.voxel51.com/user_guide/dataset_creation/index.html#common-formats) to load your data. If it is not in one of the [natively supported formats](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#supported-import-formats), use the [custom formats instructions.](https://docs.voxel51.com/user_guide/dataset_creation/index.html#custom-formats) To visualize your dataset in FiftyOne using a headless machine, aka remote server, check out the [remote sessions](https://docs.voxel51.com/user_guide/app.html#remote-sessions) documentation. Learn more about [using the FityOne App](https://docs.voxel51.com/user_guide/app.html#using-the-fiftyone-app) and about [how to load datasets](https://docs.voxel51.com/user_guide/dataset_creation/index.html#) into FiftyOne in the Docs. ## Efficiently modifying and deleting fields in views Community Slack member Tobias asked, _"According to [the documentation](https://docs.voxel51.com/user_guide/using_views.html#filtering-sample-contents) when you use `view = dataset.filter_labels(...)` to create a view and then save it, it should not override labels currently not in the view. On the other hand if I am selecting certain fields from a sample like this:_ ```python 1sample["my_custom_field1"] = "custom_value1" 2sample["my_custom_field2"] = "custom_value2" 3sample.save() 4view = dataset.select_fields("my_custom_field1") 5for sample in view: 6 sample["my_custom_field1"] = "new_value" 7 8view.save() ``` _My custom field named `my_custom_field2` has been deleted from all samples in the dataset! I am not sure whether this is working as intended, as I can't use `select_fields` then in this case if I want to update the samples from the view.”_ This is intended behavior, as [specified in the Docs](https://docs.voxel51.com/api/fiftyone.core.view.html#fiftyone.core.view.DatasetView.save). This method does not delete samples or frames from the underlying dataset that this view excludes. If a view has excluded fields or filtered list values, this method will permanently delete this data from the dataset, unless `fields` is used to omit such fields from the save. You can use the `view.save(fields="my_custom_field1")` to save only the field you modified and leave the others untouched. However, there is a more efficient way to do this using `set_values()`. This will let you avoid having to iterate over your samples in memory, and just set the new values that you want to your field directly: ```python 1dataset.set_values("my_custom_field1", ["new_value"]*len(dataset)) ``` Learn more about working with [dataset views](https://docs.voxel51.com/user_guide/using_views.html) in the Docs. ## Accessing samples from Amazon S3 Community Slack member ht asked, _"Using FiftyOne, can I load images directly from S3?"_ The open source version of FiftyOne does not support the direct access of images stored on S3. You’ll want to upgrade to FiftyOne Teams which supports cloud-backed media, multiple users, and role-based security. Learn more about FiftyOne Teams on [the website](https://voxel51.com/fiftyone-teams/) and [Docs.](https://docs.voxel51.com/teams/index.html#) ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,700+ [FiftyOne Slack](https://slack.voxel51.com/) members - 3,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 4,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 300+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/4d37703d72d4b83a85bda19eb1999d5247915fc0-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 02, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 3, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0ecb0645c4938217bcade4d3d80cf59f7b05329b-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision View Stages Tips and Tricks – Jan 20, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-295-lllmstxt|> ## AI Assistant for Computer Vision [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) VoxelGPT: Your AI Assistant for Computer Vision Jun 7, 2023 • 10 min read Article content In this article [What is VoxelGPT](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#4dabadc17f00) [VoxelGPT capabilities](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#087c11e7d360) [Getting up and running](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#1a1593076c53) [Conclusion](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#68cd50cf514c) [Join the FiftyOne community](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#663dd6d23b74) In this article [What is VoxelGPT](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#4dabadc17f00) [VoxelGPT capabilities](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#087c11e7d360) [Getting up and running](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#1a1593076c53) [Conclusion](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#68cd50cf514c) [Join the FiftyOne community](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#663dd6d23b74) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Want to surface interesting insights about your image and video datasets without writing code? Now you can with [VoxelGPT](http://gpt.fiftyone.ai/)! https://www.youtube.com/watch?v=MDeG6lM7hUg VoxelGPT combines the power of large language models (LLMs) with [FiftyOne](https://docs.voxel51.com/)’s flexible computer vision query language, making it easier than ever to semantically slice computer vision datasets and build better machine learning models. This breakthrough gives you unprecedented control over your data via natural language. You can filter your image and video datasets, uncover new insights, and become a better computer vision engineer, all _without writing a single line of code_! _Yes, you read that right: VoxelGPT translates your English queries into Python code that filters your datasets for you!_ Even better, VoxelGPT is completely open source and free to use. The code is located at [https://github.com/voxel51/voxelgpt](https://github.com/voxel51/fiftyone-gpt) and you can try it out live at [gpt.fiftyone.ai](http://try.fiftyone.ai/). Want to use VoxelGPT locally on your own data? You can install the plugin locally (read on for instructions). ## What is VoxelGPT VoxelGPT is an LLM-powered application built with [FiftyOne](https://github.com/voxel51/fiftyone), GPT-3.5, and [L](https://github.com/hwchase17/langchain) [angChain](https://github.com/hwchase17/langchain). VoxelGPT provides a chat-like interface for interacting with your computer vision dataset by translating natural language queries into FiftyOne Python syntax and constructing the resulting [`DatasetView`](https://docs.voxel51.com/user_guide/using_views.html?highlight=datasetview). This is a game-changer because mastering the FiftyOne query language, with all of its flexibility, can have a steep learning curve. With VoxelGPT, you can immediately wield the full power of FiftyOne to semantically slice your data without any prior knowledge of the query language. What’s more, VoxelGPT can also answer specific FiftyOne usage questions and machine learning questions more broadly. You can use VoxelGPT via a simple Python API, or you can use VoxelGPT natively inside the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html) by installing it as a [FiftyOne Plugin](https://docs.voxel51.com/plugins/index.html)! _Try VoxelGPT for yourself at [gpt.fiftyone.ai](http://try.fiftyone.ai/)_ VoxelGPT won’t do your computer vision work for you, but it can certainly increase your speed and efficiency. It is a _pair programmer_, a _translator_, and an _educational_ _tool_ all rolled into one. ## VoxelGPT capabilities VoxelGPT is capable of handling any of the following types of queries: - [Dataset queries](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#dataset-queries) - [FiftyOne docs queries](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#docs-queries) - [General computer vision queries](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision#computer-vision-queries) When you ask VoxelGPT a question, it will interpret your intent, and determine which type of query you are asking. If VoxelGPT is unsure, it will ask you to clarify. ### Dataset queries https://www.youtube.com/watch?v=rSJAAy5uCg0 How does it work? VoxelGPT interprets your query, translates it into FiftyOne query language Python code, and displays the resulting view. It knows how to work with datasets where the samples are _images or videos_, and has full support for _one and two-stage queries_ based on the set of examples on which it was developed. When interpreting your query, VoxelGPT will do the following: - **Recognize field and class names:** VoxelGPT is able to select the appropriate fields and class names based on the natural language query and information about the specific dataset. It uses named entity recognition to identify fields and class names, plus semantic matching for class names for fields with fewer than 1,000 classes. - **Infer relevant computations:** VoxelGPT determines whether [Brain runs](https://docs.voxel51.com/user_guide/brain.html#managing-brain-runs) or [Evaluation runs](https://docs.voxel51.com/user_guide/evaluation.html#managing-evaluations) might be relevant to a natural language query and, if they are, automatically selects the relevant runs. - **Print helpful messages:** If VoxelGPT determines that a computation may need to be run on the dataset, or that the query does not contain all information required for conversion to ViewStages, it will respond with a message indicating this. Here are some examples of dataset queries you can ask VoxelGPT—try them out live at [gpt.fiftyone.ai](http://try.fiftyone.ai/): - Retrieve 10 random samples - Display the most unique images with a false positive prediction - Just the images with at least two people detected with high confidence - Show the 25 images that are most similar to the first image with a dog ### FiftyOne docs queries https://www.youtube.com/watch?v=\_CfBYMMdTHk VoxelGPT is not just a pair programmer; it is also an educational tool. The model has access to the entire [FiftyOne docs](https://docs.voxel51.com/) \- including tutorials, user guide, and the API reference—and can use this information to answer your questions. Here are some examples of documentation queries you can ask VoxelGPT: - How do I load a dataset from the FiftyOne Zoo? - Docs: What does the `match()` stage do? - Can I export my dataset in COCO format? By effortlessly switching between dataset queries and docs queries, you can use VoxelGPT to better understand how the FiftyOne query language works. ### General computer vision queries https://www.youtube.com/watch?v=0g\_PVdCSOPU VoxelGPT can also answer general questions in computer vision, machine learning, and data science. It can help you to understand basic concepts and overcome data quality issues. Here are some examples of computer vision queries you can ask VoxelGPT: - What is the difference between precision and recall? - How can I detect faces in my images? - What are some ways I can reduce redundancy in my dataset? ### What VoxelGPT cannot do While VoxelGPT is powerful, we have purposely limited its scope to provide a focused user experience while also leaving the door open for high value upgrades in the future. VoxelGPT is currently not able to: - **Have general conversations:** VoxelGPT is not a generic chatbot. If your query is deemed out of scope, VoxelGPT will prompt you for a new query. - **Perform computations:** Some computations, such as generating a vector similarity index, may be expensive and time consuming. We don’t (yet) give VoxelGPT the power to do these computations on your behalf, but VoxelGPT will recognize when they may need to be run, and will tell you as much. - **Permanent operations:** In a similar vein, VoxelGPT is _not yet_ able to permanently change your data—for instance, it cannot delete samples from the underlying dataset, copy a dataset, or move the locations of any of your media files. - **Delegate tasks:** At present, VoxelGPT is not equipped with any HuggingGPT-style dispatching capabilities. It cannot delegate computer vision tasks to other ML models. ### Generalizability and feedback VoxelGPT’s current implementation is based on a limited set of examples, so it may not generalize well to all data. The more specific your query, the better the results will be.If you have a more involved task, see if you can split it up into multiple natural language queries and combine VoxelGPT’s results. We’d love to hear from you about how VoxelGPT performs on your use cases. Bugs, feature requests, and surprising findings that VoxelGPT uncovered are all welcome. Please drop us a line! ## Getting up and running ### Live demo If you want to experience VoxelGPT and see for yourself how the model turns natural language into computer vision insights, check out the live demo at [gpt.fiftyone.ai](http://try.fiftyone.ai/), where you can use VoxelGPT natively in the FiftyOne App on a few example datasets. ### Install VoxelGPT locally You can also interact with VoxelGPT programmatically in Python and/or work with your own datasets by installing it locally by following the instructions below. #### Install FiftyOne If you haven't already, install [FiftyOne](https://github.com/voxel51/fiftyone): ```bash 1pip install fiftyone ``` #### Provide an OpenAI API key Next provide an OpenAI API key: ```bash 1export OPENAI_API_KEY=XXXXXXXX ``` If you do not have an OpenAI key, you will need to [create one](https://platform.openai.com/account/api-keys). _A single query typically costs only $0.01 in OpenAI API calls_ #### App-only use Now if you only want to use VoxelGPT in the FiftyOne App, you can install it as a plugin by running: ```bash 1fiftyone plugins download https://github.com/voxel51/voxelgpt 2fiftyone plugins requirements @voxel51/voxelgpt --install ``` #### Python use Alternatively, if you want to programmatically interact with VoxelGPT - or you want to contribute to the project - then clone `voxelgpt` the repository: ```bash 1git clone https://github.com/voxel51/voxelgpt 2cd voxelgpt ``` and install the requirements: ```bash 1pip install -r requirements.txt ``` To make the plugin available within the FiftyOne App as well, you can symlink it into your FiftyOne plugins directory: ```bash 1ln -s "$(pwd)" "$(fiftyone config plugins_dir)/voxelgpt" ``` ### Using VoxelGPT in the App With VoxelGPT installed, you can use it natively within the FiftyOne App with any dataset: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4## load quickstart dataset. 5## -- If you want to load another dataset from the Zoo, 6## -- Ask VoxelGPT: "what datasets are in the FiftyOne Dataset Zoo?" 7 8dataset = foz.load_zoo_dataset("quickstart") 9session = fo.launch_app(dataset) ``` In the FiftyOne App, simply: - Click on the OpenAI icon above the grid, or - Press the `+` icon next to the Samples tab and select VoxelGPT **Pro tip:** use the `now` keyword to incorporate your previous prompts as context for your next query! ### Using VoxelGPT in Python If you opted for a Python/development install, you can also interact with VoxelGPT programmatically in Python. #### Interactive mode You can use VoxelGPT in _interactive mode_ via `ask_voxelgpt_interactive()`. In interactive mode, you can converse with VoxelGPT via `input()` prompts. To get started, just pass in a dataset and an App session: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4from voxelgpt import ask_voxelgpt_interactive 5 6dataset = foz.load_zoo_dataset("quickstart") 7session = fo.launch_app(dataset) 8 9ask_voxelgpt_interactive(dataset, session=session) ``` As usual, you can prompt VoxelGPT with any combination of dataset, documentation, and general computer vision queries. For example, a conversation might look like: You:     what is the difference between precision and recall? VoxelGPT:     Precision and recall are two important metrics used to evaluate the     performance of a machine learning model in classification tasks. Precision     measures the proportion of true positive predictions among all positive     predictions made by the model. Recall, on the other hand, measures the     proportion of true positive predictions among all actual positive instances     in the dataset.     In other words, precision is the ability of the model to correctly identify     positive instances, while recall is the ability of the model to identify     all positive instances. A high precision means that the model is making     very few false positive predictions, while a high recall means that the     model is able to identify most of the positive instances in the dataset.     It is important to balance precision and recall when evaluating a model's     performance. A model with high precision but low recall may be missing     important positive instances, while a model with high recall but low     precision may be making too many false positive predictions.     In the context of the FiftyOne library, precision and recall can be     computed using the \`fiftyone.utils.metrics\` module, which provides     functions for computing various classification metrics. For example, the     \`fiftyone.utils.metrics.binary\_classification\` function can be used to     compute precision and recall for binary classification tasks. You:     does FiftyOne integrate with Label Studio? VoxelGPT:     Yes, FiftyOne integrates with Label Studio, which is an open-source data     labeling tool with a user-friendly UI. The integration allows you to easily     upload your data directly from FiftyOne to Label Studio for labeling. You     can get started with Label Studio through a simple pip install to get a     local server up and running. FiftyOne provides simple setup instructions     that you can use to specify the necessary account credentials and server     endpoint to use. FiftyOne also provides an API to create projects, upload     data, define label schemas, and download annotations using Label Studio,     all programmatically in Python. You:     show me predicted airplanes VoxelGPT:     Identified potential algorithms: evaluation     No evaluation runs found.     Identified potential fields: predictions     Identified potential label classes: {'predictions': \['airplane'\]}     Identified potential view stages: \['match', 'filter\_labels', 'match\_labels', 'exclude\_labels', 'filter\_field'\]     Okay, I'm going to load dataset.filter\_labels("predictions",F("label") == "airplane") You:     now only show me the first 10 samples VoxelGPT:     Identified potential algorithms: evaluation     No evaluation runs found.     Identified potential fields: predictions     Identified potential label classes: {'predictions': \['airplane'\]}     Identified potential view stages: \['match', 'limit', 'limit\_labels', 'skip', 'sort\_by'\]     Okay, I'm going to load dataset.match(F("predictions.detections.label").contains("airplane")).limit(10) You: exit In interactive mode, VoxelGPT automatically loads any views it creates in the App, and you can access them via your [session](https://docs.voxel51.com/user_guide/app.html#sessions) object: ```python 1print(session.view.count("predictions.detections")) ``` #### Single query mode If you just want to run a single query, you can use `ask_voxelgpt()`: ```python 1from voxelgpt import ask_voxelgpt 2ask_voxelgpt("Does FiftyOne integrate with CVAT?") ``` Yes, FiftyOne integrates with CVAT, which is an open-source image and video annotation tool. You can upload your data directly from FiftyOne to CVAT to add or edit labels. FiftyOne provides simple setup instructions that you can use to specify the necessary account credentials and server endpoint to use. CVAT provides three levels of abstraction for annotation workflows: projects, tasks, and jobs. If you pass a dataset along with your query and VoxelGPT interprets your prompt as a request to load a view into your dataset, it will be returned to you: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset("quickstart") 5view = ask_voxelgpt("show me 10 random samples", dataset) ``` From there you can then interact with or further refine the view as you would any other `DatasetView` in FiftyOne. ```python 1num_objects = view.count_values("ground_truth.detections.label") ``` ## Conclusion VoxelGPT brings the power of large language models to your computer vision datasets. It is entirely open source and you can try it out today at [gpt.fiftyone.ai](http://try.fiftyone.ai/)! If you like VoxelGPT, consider showing your support by giving [voxelgpt](https://github.com/voxel51/voxelgpt) a star on GitHub. And while you’re at it, give [fiftyone](https://github.com/voxel51/fiftyone) a star too; it’s all open source! We appreciate your support 🙂 ## Join the FiftyOne community Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,700+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 3,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 4,200+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 300+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [ChatGPT](https://voxel51.com/blog/tag/chatgpt) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [LangChain](https://voxel51.com/blog/tag/langchain) [product release](https://voxel51.com/blog/tag/product-release) [VoxelGPT](https://voxel51.com/blog/tag/voxelgpt) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/770b8cfdbd7944916b1195dc11e5b173dbab8e97-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 and the State of Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 18, 2023](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/3ba0c5048ae0de4e4edc0784fc332521ba19f10f-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ The Spirit of Competition – OpenCV AI Competition 2023\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Aug 22, 2023](https://voxel51.com/blog/opencv-ai-competition-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/5068fe2d5a454e641a9ad3cc910a9b84d1dec0a6-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ NeurIPS 2023 and the State of AI Research\\ \\ Computer Vision, Event Recaps, Product & News\\ \\ • \\ \\ Dec 8, 2023](https://voxel51.com/blog/neurips-2023-and-the-state-of-ai-research) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-296-lllmstxt|> ## Visit Voxel51 at CVPR [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) 5 Reasons to Visit Voxel51 at CVPR Jun 9, 2023 • 3 min read Article content In this article [1 - 🔥 Hot off the press: Meet your new AI assistant for computer vision, VoxelGPT!](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr#9de4643f8cb2) [2 - ⚡ Experience the tremendous power of FiftyOne, the open source computer vision toolset](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr#9fdf2ca7480c) [3 - 👋 Meet fellow enthusiasts in the computer vision community](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr#7f6a6b633208) [4 - 👕 Score some epic FiftyOne swag](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr#42451690cb43) [5 - 🤫 Be there to witness the secret, must-see CVPR surprise](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr#bfdbf6ac127c) In this article [1 - 🔥 Hot off the press: Meet your new AI assistant for computer vision, VoxelGPT!](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr#9de4643f8cb2) [2 - ⚡ Experience the tremendous power of FiftyOne, the open source computer vision toolset](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr#9fdf2ca7480c) [3 - 👋 Meet fellow enthusiasts in the computer vision community](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr#7f6a6b633208) [4 - 👕 Score some epic FiftyOne swag](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr#42451690cb43) [5 - 🤫 Be there to witness the secret, must-see CVPR surprise](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr#bfdbf6ac127c) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) CVPR is just around the corner! We're excited to take part in this amazing event again this year and to join the community of computer vision researchers and engineers from around the world gathering in Vancouver, Canada June 18-22. If you're also attending, we'd love to meet you! Here are the top 5 reasons you'll want to visit Voxel51 at [CVPR 2023](https://cvpr2023.thecvf.com/). ## 1 - 🔥 Hot off the press: Meet your new AI assistant for computer vision, VoxelGPT! Wish you could search your images or videos without writing a line of code? Now you can! The [VoxelGPT](https://voxel51.com/voxelgpt/) plugin combines the power of GPT-3.5 with FiftyOne’s computer vision query language, enabling you to filter, sort, and semantically slice your data with natural language. See it in action at booth 1618! ![](https://cdn.sanity.io/images/h6toihm1/production/e228c74eb5167c87f59620b02626c3e318931471-1920x1080.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## 2 - ⚡ Experience the tremendous power of FiftyOne, the open source computer vision toolset Nothing limits model performance more than a bad dataset. Let us show you how you can be the hero against bad data with [FiftyOne](https://github.com/voxel51/fiftyone) – the open source toolkit that helps you easily improve the quality of your datasets, gain insights about your models, and build better AI. Now with VoxelGPT! Visit the Voxel51 booth #1618 to experience FiftyOne: - See a live demo! Tell us what you’d like to see and we’ll show you. Or we can show you some of the most loved features. Either way, come see FiftyOne and VoxelGPT in action on the big screen! - Want to try it for yourself? Visit the hands-on demo station to let your fingers do the walking as you visualize and explore datasets, gain insights about models, and much more. Simply walk up and give it a try! Or get a 15-min guided tour of FiftyOne in one of these Flash Sessions in the Expo Hall (area #1718): - Wednesday, **June 21 at Noon-12:15pm** \- Eric Hofesmann will present the talk, _Better Data + Better Models: Your Superpowers with FiftyOne!_ - Thursday, **June 22 at 10:30-10:45am** \- Jacob Marks will present the talk, _Researcher Survival Guide: Bye Bye, Bad Data_ In just 15 minutes, you'll walk away with a solid understanding of how to easily improve the quality of your datasets, gain insights about your models, and build better AI with FiftyOne. ## 3 - 👋 Meet fellow enthusiasts in the computer vision community We're as excited about the latest happenings in computer vision as you are! Want to discuss the state of computer vision, talk about data-centric AI, or best practices for overcoming data quality issues and building higher quality models? Come meet the Voxel51 team of ML engineers, developers, co-founders, and open source enthusiasts at booth #1618. Not only will we be available to discuss all things computer vision, we’d also simply love to meet you, and swag you up with some of our latest and greatest threads (while supplies last!). ![](https://cdn.sanity.io/images/h6toihm1/production/779e12dfa3d53389bff790a5eebbff54b67e2317-3797x2849.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## 4 - 👕 Score some epic FiftyOne swag Who doesn't like a good giveaway to commemorate an amazing event? Stop by our booth to see our lineup of goodies, including tshirts, water bottles, hats, and stickers. And snag some swag for yourself! It's easy to walk away with gear you'll love. Simply become a member of the open source FiftyOne community. For example: star FiftyOne on GitHub, follow Voxel51 on LinkedIn, join the Community Slack, join our mailing list, or one of the other convenient and easy ways to join in. The more ways you engage, the more gear you can grab. It's that easy. As much as we love swag, we also simply love welcoming new members to the open source FiftyOne community. Why? We know how valuable FiftyOne truly is for data engineers and scientists and therefore getting it into the hands of even more people is what we're all about. Even if you come for the swag, we're convinced you'll walk away with two things you'll love: sweet gear and first-hand experience of how FiftyOne can help you level up your computer vision workflows. ![](https://cdn.sanity.io/images/h6toihm1/production/816039978ffb5f77fc3409d7d340a4e3ad6bf23b-1080x1000.png?auto=format&dpr=2&fit=max&q=75&w=1080) ## 5 - 🤫 Be there to witness the secret, must-see CVPR surprise We went all in on CVPR this year. First we dug into the CVPR paper data and published some patterns and trends that we uncovered in the blog post [CVPR 2023 and the State of Computer Vision](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision/). Next, to help you make the most of a jam-packed June week, we compiled a list of the top 10 papers you just can’t miss, with links and summaries, in the blog post [CVPR 2023 Survival Guide](https://voxel51.com/blog/cvpr-2023-survival-guide/). … But that’s not all. We have another big, exciting surprise that you won’t want to miss! Visit the Voxel51 booth 1618 to witness this must-see experience, and let us know what you think! [CVPR](https://voxel51.com/blog/tag/cvpr) [CVPR 2023](https://voxel51.com/blog/tag/cvpr-2023) [events](https://voxel51.com/blog/tag/events) [tradeshows](https://voxel51.com/blog/tag/tradeshows) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/e6fca750f87dc5c6c8adc03d909d17706d81295b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How to Get the Most out of CVPR\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jun 20, 2023](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr) [![](https://cdn.sanity.io/images/h6toihm1/production/7f15954f9509d8b465aa7b1c6e3d8e3fe1e4bb71-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ 3 Reasons to Visit Voxel51 at ICCV23!\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Oct 2, 2023](https://voxel51.com/blog/3-reasons-to-visit-voxel51-at-iccv23) [![](https://cdn.sanity.io/images/h6toihm1/production/770b8cfdbd7944916b1195dc11e5b173dbab8e97-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 and the State of Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 18, 2023](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-297-lllmstxt|> ## Computer Vision Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup — June 8, 2023 Jun 9, 2023 • 6 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#5865e0aa0d1f) [Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#33bfe2e94eb9) [Re-Annotating MS COCO, an Exploration of Pixel Tolerance](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#791f38944bb2) [Redefining State-of-the-Art with YOLOv5 and YOLOv8](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#ff1fff59f8ba) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#98364593fee7) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#639ca9493c0e) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#79338a39f570) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#5865e0aa0d1f) [Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#33bfe2e94eb9) [Re-Annotating MS COCO, an Exploration of Pixel Tolerance](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#791f38944bb2) [Redefining State-of-the-Art with YOLOv5 and YOLOv8](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#ff1fff59f8ba) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#98364593fee7) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#639ca9493c0e) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023#79338a39f570) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We just wrapped up the June 8, 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/), and if you missed it or want to revisit it, here’s a recap! In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes this month was [Global Empowerment Mission](https://www.globalempowermentmission.org/) (GEM)! GEM aims to bring the most aid, to the most people, in the least amount of time. We are sending this event’s charitable donation of $200 to GEM on behalf of the computer vision community. Missed the Meetup? No problem. Here are playbacks and talk abstracts from the event. ## Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation https://www.youtube.com/watch?v=0dAUPNWQIEg Large-scale text-to-image generative models have been a revolutionary breakthrough in the evolution of generative AI, allowing us to synthesize diverse images that convey highly complex visual concepts. However, a pivotal challenge in leveraging such models for real-world content creation tasks is providing users with control over the generated content. In this paper, we present a new framework that takes text-to-image synthesis to the realm of image-to-image translation — given a guidance image and a target text prompt, our method harnesses the power of a pre-trained text-to-image diffusion model to generate a new image that complies with the target text, while preserving the semantic layout of the source image. Specifically, we observe and empirically demonstrate that fine-grained control over the generated structure can be achieved by manipulating spatial features and their self-attention inside the model. [Michal Geyer](https://www.linkedin.com/in/michal-geyer-aa3a021b1) and [Narek Tumanya](https://www.linkedin.com/in/narek-tumanyan-5a3334122) are Masters students at the Weizmann Institute of Science in the Computer Vision department. Learn more about the [Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation](https://arxiv.org/abs/2211.12572) paper on arXiv or check out their poster at [CVPR 2023](https://cvpr2023.thecvf.com/). Q&A from the talk included: - Can your approach be extended by including a portion of the image where the text is supposed to make a change/diffuse? ## Re-Annotating MS COCO, an Exploration of Pixel Tolerance https://www.youtube.com/watch?v=LPj9oXRNBZ8 The release of the COCO dataset has served as a foundation for many computer vision tasks including object and people detection. In this session, we’ll introduce the Sama-Coco dataset, a re-annotated version of COCO focused on fine-grained annotations. We’ll also cover interesting insights and learnings during the annotation phase, illustrative examples, and results of some of our experiments on annotation quality as well as how changes in labels affect model performance and prediction style. [Jerome Pasquero](https://www.linkedin.com/in/jeromepasquero/) is Principal Product Manager at Sama. Jerome holds a Ph.D. in electrical engineering and is listed as inventor on more than 120 US patents along with published over 10 peer-reviewed journal and conference articles. [Eric Zimmermann](https://www.linkedin.com/in/eric-zimmermann/) is an Applied Scientist at Sama helping to redefine annotation quality guidelines. He is also responsible for building internal curation tools which aim to improve the process on how clients and annotators interact with their data. Q&A from the talk included: - With respect to Sama-COCO vs MS-COCO, were any tests done to compare corrections in label only versus, mask only, versus both mask and label relative to model performance changes? - I am currently working on airborne object segmentation. The basic idea is to segment an incoming airplane by a still camera where initially the plane is small and increases in size as it comes closer... I’m using Yolov8 segmentation. What are your opinions on how to annotate the dataset and which metric should I use to determine the efficiency of the model? You can find [the slides from the talk here](https://voxel51.com/wp-content/uploads/2023/06/Sama-Computer-Vision-Meetup-June-8-2023.pdf). ## Redefining State-of-the-Art with YOLOv5 and YOLOv8 https://www.youtube.com/watch?v=rtLqqAtzp1U In recent years, object detection has been one of the most challenging and demanding tasks in computer vision. YOLO (You Only Look Once) has become one of the most popular and widely used algorithms for object detection due to its fast speed and high accuracy. YOLOv5 and YOLOv8 are the latest versions of this algorithm released by Ultralytics, which redefine what “state-of-the-art” means in object detection. In this talk, we will discuss the new features of YOLOv5 and YOLOv8, which include a new backbone network, a new anchor-free detection head, and a new loss function. These new features enable faster and more accurate object detection, segmentation, and classification in real-world scenarios. We will also discuss the results of the latest benchmarks and show how YOLOv8 outperforms the previous versions of YOLO and other state-of-the-art object detection algorithms. Finally, we will discuss the potential for this technology to “do good” in real-world scenarios and across various fields, such as autonomous driving, surveillance, and robotics. [Glenn Jocher](https://www.linkedin.com/in/glenn-jocher/) is founder and CEO of Ultralytics. In 2014 Glenn founded Ultralytics to lead the United States National Geospatial-Intelligence Agency (NGA) antineutrino analysis efforts, culminating in the miniTimeCube experiment and the world’s first-ever Global Antineutrino Map published in Nature. Today he’s driven to build the world’s best vision AI as a building block to a future AGI, and YOLOv5, YOLOv8, and Ultralytics HUB are the spearheads of this obsession. Q&A from the talk included: - Is the detection performance/speed noticeably faster or better for custom datasets, eg, person only, between Yolo V4, V5, and V8? - What is more preferable while segmenting small objects? Box annotations or polygon annotations? - Is it possible to ensemble different YOLO models trained on different sets of data for achieving a credible mAP? What are the methods or ways we can do that? - Any tips for combining YOLOv8 and RetinaNet? - Is it possible to use YOLOv8 with other programming languages besides Python, with C++ for example? - What line of thinking inspired you to make changes to the architecture in V8, for example not use anchor boxes? - Do you suggest anything else for tracking besides ByteTrack? ## Join the Computer Vision Meetup! \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Computer Vision Meetup membership has grown to more than [4,000 members](https://www.meetup.com/pro/computer-vision-meetups/) in just under a year! The goal of the Meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. Join one of the 13 Meetup locations closest to your timezone. - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) We have exciting speakers already signed up over the next few months! Become a member of the [Computer Vision Meetup closest to you](https://www.meetup.com/pro/computer-vision-meetups/), then register for the Zoom. ## What’s Next? \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Up next on July 13 at 10 AM Pacific we have a very special vector search-themed Computer Vision Meetup happening featuring speakers from [Zilliz/Milvus](https://zilliz.com/), [Weaviate](https://weaviate.io/), [Qdrant](https://qdrant.tech/), [LanceDB](https://lancedb.com/), and Voxel51. - **Unleashing the Potential of Visual Data: Vector Databases in Computer Vision**\- [Filip Haltmayer](https://www.linkedin.com/in/filiphaltmayer/) (Zilliz) - **Computer Vision Applications at Scale with Vector Databases** \- [Zain Hasan](https://www.linkedin.com/in/zainhas/) (Weaviate) - **Reverse Image Search for Ecommerce Without Going Crazy** \- [Kacper Łukawski](https://www.linkedin.com/in/kacperlukawski/) (Qdrant) - **Fast and Flexible Data Discovery & Mining for Computer Vision at Petabyte Scale** \- [Jai Chopra](https://www.linkedin.com/in/jaichopra/) (LanceDB) - **How-To Build Scalable Image and Text Search for Computer Vision Data using Pinecone and Qdrant** \- [Jacob Marks](https://www.linkedin.com/in/jacob-marks/) (Voxel51) Register for the Zoom [here](https://voxel51.com/computer-vision-events/july-2023-computer-vision-meetup/?utm_source=blog). You can find a complete schedule of upcoming Meetups on [the Voxel51 Events page](https://voxel51.com/computer-vision-events/). ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [Diffusion models](https://voxel51.com/blog/tag/diffusion-models) [MS COCO](https://voxel51.com/blog/tag/ms-coco) [Sama-Coco](https://voxel51.com/blog/tag/sama-coco) [YOLOv5](https://voxel51.com/blog/tag/yolov5) [YOLOv8](https://voxel51.com/blog/tag/yolov8) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/803cb935ddffbb5b29b6d3c73104b3a1221ddbfc-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 3)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-3) [![](https://cdn.sanity.io/images/h6toihm1/production/98e839c6e81bb9c4fa3ad96bf0d5d1b77ee11f6c-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 1)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-1) [![](https://cdn.sanity.io/images/h6toihm1/production/de34ad70ef11fe3d7daf5b3c8ff6214e76c85522-2560x1440.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Giving YOLOv8 a Second Look (Part 2)\\ \\ Tutorials\\ \\ • \\ \\ Feb 22, 2023](https://voxel51.com/blog/giving-yolov8-a-second-look-part-2) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-298-lllmstxt|> ## Data-Centric AI Tooling [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Too Many Pixels, So Little Time: Why You Need Data-Centric Tooling in Your AI Stack Jun 15, 2023 • 6 min read Article content In this article [Evaluating How FiftyOne Fits Into Your AI Stack](https://voxel51.com/blog/too-many-pixels-so-little-time#06739e69fb08) [FiftyOne Enterprise: Leveling Up with Collaboration and Security](https://voxel51.com/blog/too-many-pixels-so-little-time#6fdddd4f226e) In this article [Evaluating How FiftyOne Fits Into Your AI Stack](https://voxel51.com/blog/too-many-pixels-so-little-time#06739e69fb08) [FiftyOne Enterprise: Leveling Up with Collaboration and Security](https://voxel51.com/blog/too-many-pixels-so-little-time#6fdddd4f226e) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _Now, you can easily try out FiftyOne Open Source and FiftyOne Enterprise in your browser and evaluate first hand where they fit into your AI Stack_ When I was a grad student in the early 2000s, there was a framed picture on the wall of my lab that read: "Too many pixels, so little time." I don't know where my advisor found that nugget of wisdom, but it has run increasingly true over the last two decades. In those days, our datasets numbered in the hundreds of images or a few videos. And, we painstakingly combed through their pixels to map patterns in the data to specific capabilities and failures in our models while building deployable AI systems. Fast-forward a decade and we found datasets increasing in size by two or three orders of magnitude; it became quite difficult to continue that practice. Yet, it did not stop being important, or even critical, in developing highly performant AI systems. Unstructured data is exceptional. It significantly outnumbers structured data on the internet, e.g., video alone accounted for [65% of internet traffic in 2021](https://www.sandvine.com/press-releases/sandvines-2023-global-internet-phenomena-report-shows-24-jump-in-video-traffic-with-netflix-volume-overtaking-youtube). It drives recent trends in areas like autonomous driving, large language models, and generative AI. Yet, in each case, understanding the nuances of a collection of unstructured data to unlock its intrinsic value and build a better AI system is no easy task. Aggregate measures like accuracy, precision-recall, and others fail to close the loop when leveraging unstructured data in a task. They may provide a high-level overview of a model's performance but do not lead to a better understanding of the overall data distribution or to actionable insights on how to improve weak performance. Imagine you go to the dentist to have a cavity filled. If the dentist worked like this, then she'd take the drill, close her eyes, drill away, apply the filling, and then go take an X-ray to see how well she did. Ridiculous you say? Yes, very ridiculous. How would she actually work? She would constantly close the loop. She would look into your mouth, check the location, refer to the existing X-rays, etc. Over the last two decades, I've come to appreciate the need to work in a tight, closed-loop mindset in AI. This need led to the creation of the [open source FiftyOne toolset](https://github.com/voxel51/fiftyone). FiftyOne helps AI scientists and engineers close the loop while building systems that leverage unstructured data. With FiftyOne you can semantically search and explore your datasets using the straightforward Python query language or, now, with natural language through [VoxelGPT](https://voxel51.com/voxelgpt); you can visually analyze the distribution of your data both in aggregate, via, for example, embeddings, and in fine detail, via visualization at the individual sample-level; you can compare and evaluate multiple model runs on that data; and more. People in the FiftyOne community have shared a variety metaphors to capture the amazingness of FiftyOne, including this one: **FiftyOne is the (missing) debugger for the AI stack**. The rationale? Many computer scientists consider the debugger as the most important tool of programming, second only to the compiler/interpreter. And, now, you can more easily try out FiftyOne directly in your browser, without installing anything locally. Even more exciting, you can, for the first time, get a taste of [FiftyOne Enterprise](https://voxel51.com/enterprise/), our team-oriented Enterprise version of FiftyOne, in your browser too. > Get started instantly at [try.fiftyone.ai](https://try.fiftyone.ai/)! ## Evaluating How FiftyOne Fits Into Your AI Stack We released FiftyOne as [open source](https://github.com/voxel51/fiftyone) software in [August of 2020](https://medium.com/voxel51/introducing-fiftyone-a-tool-for-rapid-data-model-experimentation-73c85b8406e1). FiftyOne was a one of a kind tool upon its release: it concurrently solved the need to visually inspect unstructured datasets and model performance on those data while also providing a schema-free mechanism to construct sophisticated semantic queries on those data and models to drill into the unstructured datasets making it more easy to uncover failure modes and corner cases. This article is not meant to be a comprehensive overview of FiftyOne. For that you can check out this recent playback of the [Getting Started with FiftyOne Workshop](https://www.youtube.com/playlist?list=PLuREAXoPgT0SJLKsgFzKxffMApbXp90Gi), register for an [upcoming live Workshop](https://voxel51.com/computer-vision-events/), or go explore the [user guide](https://docs.voxel51.com/user_guide/index.html). Until now, there were two primary ways to try out FiftyOne. Both required some effort on the user. First, you could `pip install fiftyone` and then launch a quickstart, which would allow you to play in a locally deployed sandbox on a small dataset. Time to your first "Aha!" moment: 15 minutes. Second, you could git clone `https://github.com/voxel51/fiftyone` and install a locally functional developer environment. Time to your first "Aha!" moment: 4 hours. Even though these two routes have found wild success— `pip install fiftyone` has been executed 1,329,629 times as of this writing—we wanted to level up. For example, one limitation in these two approaches is that in order to try it on a dataset with more than a handful of samples requires learning the FiftyOne [Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html) , or importing [your own data](https://docs.voxel51.com/user_guide/dataset_creation/index.html). Although not uncommon, for example, similar SaaS tools in our space like Scale AI's Nucleus allow free trial usage for only a small number of data points. In all of these cases, it’s immensely difficult to get a real sense for how useful these tools will be for effective AI workflows without being able to evaluate them on larger datasets. With [try.fiftyone.ai](https://try.fiftyone.ai/) you instead can interact easily with larger datasets directly within your browser. No download required. Full evaluative, read-only functionality. You can try out various functionality like exploring embeddings, interacting with natural language queries through VoxelGPT, visually inspecting annotations, and more. See it for yourself right now! Simply visit [try.fiftyone.ai](http://try.fiftyone.ai/), click on the coco-2017-validation dataset as an example, open the embeddings tab, select clustering and color by uniqueness, and then lasso the most interesting looking cluster. I couldn't be more excited that it's now possible to directly evaluate FiftyOne in your browser! ## FiftyOne Enterprise: Leveling Up with Collaboration and Security You may not realize it, but with [try.fiftyone.ai](https://try.fiftyone.ai/) you’re actually getting a taste of FiftyOne Enterprise, the multiuser, collaborative version of FiftyOne. Open source FiftyOne was designed with the individual AI scientist in mind. It speeds up workflows and leads to measurably better models faster. It assumes all data is local. It assumes that this scientist is managing a specific computer instance and running FiftyOne there. In fact, this enabled us to create a unique bidirectional stateful relationship between a state of the art user interface and an underlying Python session. But, open source FiftyOne does not support multiple users. It has no notion of security. No notion of "sign on" let alone single sign-on. No notion of roles or usage rights. No notion of collaboration. AI teams need all of these capabilities to function in the Enterprise environment. What can AI teams do? FiftyOne Enterprise is the answer. We rebuilt FiftyOne's system layer from the ground up while maintaining full backwards compatibility to the data model and retaining the dataset browser experience. FiftyOne Enterprise adds enterprise security; single sign-on; role-based access control; collaboration functionality; sharing of datasets and models; full interactive export functionality; cloud-backed media; and more. By way of analogy, FiftyOne is git, and FiftyOne Enterprise is GitHub. You can read more about [FiftyOne Enterprise in the docs](https://docs.voxel51.com/enterprise/index.html). With [try.fiftyone.ai](https://try.fiftyone.ai/), you can now try certain aspects of these multi-user, collaboration experiences directly within your browser, such as the cloud-backed media, SSO, collaboration, and sharing. For example, when you connect to the site, you will use a single-sign-on functionality to connect a social login to allow access. Then, you will see the list of datasets shared with you. The demo site is currently read-only where datasets are set to ‘Can view’ for all users, so if you want to check out the full [user roles and permissions](https://docs.voxel51.com/teams/roles_and_permissions.html) experience within FiftyOne Enterprise, [reach out to us](https://voxel51.com/talk-to-sales/) and we’ll be in touch! [AI stack](https://voxel51.com/blog/tag/ai-stack) [data-centric AI tooling](https://voxel51.com/blog/tag/data-centric-ai-tooling) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne Enterprise](https://voxel51.com/blog/tag/fiftyone-enterprise) [try.fiftyone.ai](https://voxel51.com/blog/tag/try-fiftyone-ai) ![](https://cdn.sanity.io/images/h6toihm1/production/ad9fb967c5455e0f763411fb81956767d7f26482-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Jason Corso Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a4c2bee9ed053c5be2a1c161e5abf758c9a12ff8-1400x923.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.18 with App Performance Improvements, Sidebar Modes, and Custom Attributes\\ \\ Product & News\\ \\ • \\ \\ Nov 15, 2022](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes) [![](https://cdn.sanity.io/images/h6toihm1/production/e94f20fa81716294c7a6caccf7256e9106cb6e89-967x800.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.17 with Grouped Datasets, 3D, Geolocation, and Custom Plugins\\ \\ Product & News\\ \\ • \\ \\ Sep 21, 2022](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins) [![](https://cdn.sanity.io/images/h6toihm1/production/2be41b07bd86d7efc5916442ca5b94aa5115234d-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ The Greatest Hits of 2022: FiftyOne & Voxel51\\ \\ Product & News\\ \\ • \\ \\ Jan 10, 2023](https://voxel51.com/blog/the-greatest-hits-of-2022-fiftyone-voxel51) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-299-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – June 16, 2023 Jun 16, 2023 • 5 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#40f98c6d5c10) [Understanding the FiftyOne Brain’s compute\_hardness function](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#535c9ee4cdd7) [Custom annotation color schemes in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#11fd9cf09b37) [Using Dynamic Groups with a video dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#2776fac028c5) [Resolving FiftyOne MongoDB backend port conflicts](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#970daef310f2) [Support for multiple users in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#d5b2e884b9f4) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#012e58f1c815) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#40f98c6d5c10) [Understanding the FiftyOne Brain’s compute\_hardness function](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#535c9ee4cdd7) [Custom annotation color schemes in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#11fd9cf09b37) [Using Dynamic Groups with a video dataset](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#2776fac028c5) [Resolving FiftyOne MongoDB backend port conflicts](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#970daef310f2) [Support for multiple users in the FiftyOne App](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#d5b2e884b9f4) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023#012e58f1c815) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. As an open source community, the FiftyOne community is open to all. This means everyone is welcome to ask questions, and everyone is welcome to answer them. Continue reading to see the latest questions asked and answers provided! ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## Understanding the FiftyOne Brain’s compute\_hardness function Community Slack member Ananthu asked: _“How does the [compute\_hardness](https://docs.voxel51.com/api/fiftyone.brain.html?highlight=compute_hardness#fiftyone.brain.compute_hardness) function work under the hood? An abstract explanation would be nice!”_ Community member Joy gave an excellent response: The `compute_hardness` function computes the entropy, or the amount of uncertainty, in the predicted class probabilities of a model's output. _“A classification model predicts which class a certain data point belongs to. The "raw" output of the model is often in the form of ["logits"](https://en.wikipedia.org/wiki/Logit), or log-odds. Each logit corresponds to a score for a specific class. The higher the score, the more likely the model thinks the data point belongs to that class. However, these logits are not in a very interpretable form. So, they are usually transformed into probabilities using a function like [softmax](https://en.wikipedia.org/wiki/Softmax_function). The softmax function takes a vector of logits and squashes them into a range of \[0, 1\] such that the entire vector sums to 1.0. This way, each element in the softmax output can be interpreted as the probability of the data point belonging to a specific class._ _Now, entropy is a concept borrowed from information theory. In this context, it's used to quantify the "uncertainty" or "surprise" of a probability distribution. A uniform distribution, where all outcomes are equally likely, has the highest entropy because it is the most uncertain or surprising - you have no idea which outcome is going to occur. Conversely, a distribution where one outcome is certain to happen has an entropy of zero, because there is no surprise or uncertainty._ _So, when you calculate the entropy of the softmax output, you're calculating the uncertainty in the model's predictions. If the entropy is low, it means the model is very confident in its predictions. If the entropy is high, it means the model is less certain about its predictions.”_ Learn more about the `compute_hardness` function and the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html) in the Docs. ## Custom annotation color schemes in the FiftyOne App Community Slack member ZKW asked: _“Is it possible for a detection/segmentation label with a different class to be visualized in different colors, in the same field instead of all in the same color? If possible, I’d also like to change the bbox color in the same way.”_ Community member Joy provided the following explanation: _“In the latest version of FiftyOne you have a lot of control over the colors of your annotations. Check out the Color Scheme widget in the App toolbar.”_ ![](https://cdn.sanity.io/images/h6toihm1/production/09a4bc867e748e355ac7d2a074ee98f833610412-914x703.gif?auto=format&dpr=2&fit=max&q=75&w=914) Learn more about [color schemes](https://docs.voxel51.com/user_guide/app.html#app-color-schemes) and other new features in the latest [FiftyOne 0.21.0](https://docs.voxel51.com/release-notes.html#fiftyone-0-21-0) release. ## Using Dynamic Groups with a video dataset Community Slack member Amy asked: _“I would like to visualize a set of point clouds recorded in the same location over time. I have visualized the individual point clouds in FiftyOne, but is there a way to iterate through the dataset like a video clip?”_ FiftyOne’s [dynamic groups](https://docs.voxel51.com/user_guide/using_views.html#grouping) functionality supports this type of use case -- you can `group_by` a `scene_id`, ordering by a frame number for instance. If you are including stereo images, this is similar to our [quickstart-groups](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#dataset-zoo-quickstart-groups) dataset. Here's my sample code for loading this dataset and creating mock `scene_id` and `frame` variables: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset('quickstart-groups') 5 6scene_id = ['scene1','scene2','scene3','scene4']*50 7frame = range(len(dataset)) 8dataset.set_values('scene_id',scene_id) 9dataset.set_values('frame',frame) 10 11session = fo.launch_app(dataset,auto=False) ``` After launching the App, you can use the Dynamic Groups tool to group by the `scene_id`, ordered by `frame`. There will be one group in the grid for each scene, and clicking enters a display that is paginated over the frames in that group. Learn more about how to [dynamically group](https://docs.voxel51.com/user_guide/app.html#app-dynamic-groups) samples in a collection by a specified field in the Docs. ## Resolving FiftyOne MongoDB backend port conflicts Community Slack member ZKW asked: _“Can anyone recommend a method for how to deal with a MongoDB port conflict when deploying 2 Docker containers on the same Linux machine? Both of them are running FiftyOne and both have this mount point:_ _-v /opt/.fiftyone/var/lib/mongo:/root/.fiftyone/var/lib/mongo”_ Community member Joy provided the following explanation: MongoDB explicitly prohibits running two Mongo instances on the same data directory. More info in this [StackOverflow post](https://stackoverflow.com/questions/52660659/cant-run-multiple-mongodb-docker-container-with-same-shared-volume). ## Support for multiple users in the FiftyOne App Community Slack member Manoharan asked: _“I am running FiftyOne App, but I have a problem accessing the application with multiple users. Two users tried to open the same application and change the loaded dataset. The App automatically reloaded for the second person when a change is made by the first person. Could you suggest how multiple users can access the same FiftyOne deployment and dataset?”_ The open source distribution of FiftyOne does not support multiple users because it was designed with the individual AI scientist in mind. It is a single-user, locally mounted data experience. If you need support for multiple users, cloud-based data, role-based security, and data versions, we recommend you check out [FiftyOne Teams](https://voxel51.com/fiftyone-teams/), which was created to meet the needs of AI teams collaborating on datasets and models. ![](https://cdn.sanity.io/images/h6toihm1/production/a9c2ec647ee4a0aa813551da999aef02262e598c-400x400.jpg?auto=format&dpr=2&fit=max&q=75&w=400) ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,700+ [FiftyOne Slack](https://slack.voxel51.com/) members - 3,100+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 4,400+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 300+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [compute\_hardness](https://voxel51.com/blog/tag/compute_hardness) [custom color schemes](https://voxel51.com/blog/tag/custom-color-schemes) [dynamic groups](https://voxel51.com/blog/tag/dynamic-groups) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [MongoDB](https://voxel51.com/blog/tag/mongodb) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0ecb0645c4938217bcade4d3d80cf59f7b05329b-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision View Stages Tips and Tricks – Jan 20, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4d37703d72d4b83a85bda19eb1999d5247915fc0-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 02, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 3, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/a17b9ee7620741f8c3225d0256d174c857a575d5-1200x672.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Oct 14, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Oct 15, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-14-2022) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-300-lllmstxt|> ## Introducing VoxelGPT Plugins [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps), [Plugins](https://voxel51.com/blog/category/plugins) Webinar Recap: Introducing VoxelGPT and Building Custom Plugins Jun 20, 2023 • 4 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#799f7ac3369b) [What Is FiftyOne?](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#e7157055bd4e) [VoxelGPT: Your AI Assistant for Computer Vision](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#0f32e3da2731) [Plugins & Operators](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#23cf13d0592f) [Bringing It All Together: Wait, VoxelGPT Is a Plugin?!](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#d76b699cf2c7) [Live Demo Time!](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#4e8c87ef1596) [Other Notes](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#1baaed76b7da) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#799f7ac3369b) [What Is FiftyOne?](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#e7157055bd4e) [VoxelGPT: Your AI Assistant for Computer Vision](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#0f32e3da2731) [Plugins & Operators](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#23cf13d0592f) [Bringing It All Together: Wait, VoxelGPT Is a Plugin?!](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#d76b699cf2c7) [Live Demo Time!](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#4e8c87ef1596) [Other Notes](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins#1baaed76b7da) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We recently released two new products: [FiftyOne 0.21](https://voxel51.com/blog/announcing-fiftyone-0-21/), which is packed with exciting new features, and [VoxelGPT](https://voxel51.com/voxelgpt/), a revolutionary new plugin that turns natural language queries into actual Python code that can filter, sort, semantically slice and reveal insights about the images and videos in your dataset, all without having to write a single line of code. Voxel51 Co-Founder and CTO [Brian Moore](https://www.linkedin.com/in/brimoor/) walked us through the new capabilities in a live webinar, with plenty of live demos and code examples, so that you can see all the awesomeness in action. Specifically, the new goodies that Brian demonstrated in the webinar include: - **VoxelGPT**: Your AI assistant for computer vision, powered by GPT-3.5 and FiftyOne’s flexible computer vision query language - **Operators & Plugins**: New capabilities in FiftyOne 0.21 that make VoxelGPT and any custom features you can dream up to add to your own FiftyOne App Watch the video playback on [YouTube](https://www.youtube.com/watch?v=F-2M37NFavU), take a look at the [slides](https://docs.google.com/presentation/d/1YN4OoSLPNsiaCPLPm1lNNMu-LUEQPOBL-sF9yijacl0/edit?usp=sharing), read the [transcript](https://www.rev.com/transcript-editor/shared/WjIwu105_oi8SgUmGAdUWEPWBvP5wVJW02IaEmNNpONTDfo7Y3ueDWTJFou9OkL1Uaw5BMlLSK4XOEJlyrO0ofnnxGY?loadFrom=SharedLink), and read the recap below for the highlights. Enjoy! https://www.youtube.com/watch?v=F-2M37NFavU ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave attendees the opportunity to help guide our monthly donation to charitable causes. This time, the votes were split evenly between two charities: [Global Empowerment Mission](https://www.globalempowermentmission.org/) and [Wildlife AI](https://www.wildlife.ai/). This means we’ll be sending $100 to each of these charities on behalf of the live audience and the FiftyOne community! ![](https://cdn.sanity.io/images/h6toihm1/production/9d7c97028e69f65ecb7be284a71446c78d0e8b04-1024x264.png?auto=format&dpr=2&fit=max&q=75&w=1024)![](https://cdn.sanity.io/images/h6toihm1/production/e80959f169a6a89e3cef73f9bb1bdd1750499fe7-770x146.png?auto=format&dpr=2&fit=max&q=75&w=770) ## What Is FiftyOne? Brian starts with a quick overview of what FiftyOne is for those who might be new to it. [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Here’s a snapshot of some of the computer vision workflows made possible by FiftyOne: - Visualize, query, and analyze computer vision datasets - Streamline data annotation workflows - Identify and correct labeling mistakes - Analyze model performance, both visually and programmatically - And dozens more workflows! ## VoxelGPT: Your AI Assistant for Computer Vision From [~2:31](https://www.youtube.com/watch?v=F-2M37NFavU&t=151s) to 13:44 in the webinar, Brian walks through what VoxelGPT is and gives a demo on how to use it. Brian explains that “VoxelGPT is a chat-based interface to allow you to ask questions about your datasets, the FiftyOne documentation, and even general computer vision knowledge directly in the FiftyOne App.” To demonstrate how to use VoxelGPT, Brian uses a dataset from COCO with ground truth and object detections, as well as model predictions from YOLO-NAS. Then Brian asks VoxelGPT these questions and shows the results: - _Can I export this data set in coco format?_ (Under the hood, VoxelGPT understands to look into the FiftyOne documentation.) - _Show me the samples with high confidence predictions that were evaluated as false positives._ (Under the hood, VoxelGPT understands to look into the dataset.) - _Explain computer vision in one sentence._ (Under the hood, VoxelGPT understands that you’re asking a general knowledge question and puts together an answer for that.) For those of you who are curious how we built VoxelGPT, Brian shares this information starting at ~8:29 in the video. Here’s the high level diagram: ![](https://cdn.sanity.io/images/h6toihm1/production/96a646393c4478cf98c2004ab673544e605d64c3-1474x824.png?auto=format&dpr=2&fit=max&q=75&w=1474) You can also learn more about how VoxelGPT was built in these two articles on Towards Data Science: - [How I Turned ChatGPT into an SQL-Like Translator for Image and Video Datasets](https://towardsdatascience.com/how-i-turned-chatgpt-into-an-sql-like-translator-for-image-and-video-datasets-7b22b318400a) - [What I Learned Pushing Prompt Engineering to the Limit](https://towardsdatascience.com/what-i-learned-pushing-prompt-engineering-to-the-limit-c40f0740641f) ## Plugins & Operators At ~13:44 in the webinar replay, Brian describes the upgrades FiftyOne 0.21 brings to the Plugin framework, as well as a new capability introduced in 0.21 – operators. [FiftyOne Plugins](https://docs.voxel51.com/plugins/index.html) provide a flexible mechanism for extending the functionality of the FiftyOne App: - Add new panels, visualizers, and other types - Trigger custom Python workflows - On-startup hooks - Custom icon placement - All components are optional - May contain 100% JavaScript (JS), 100% Python, or a mix [FiftyOne Operators](https://docs.voxel51.com/plugins/index.html#operators) allow you to trigger custom workflows directly from the App via custom forms - Rich type system to accept user inputs - Dynamic forms - Asynchronous streaming output - Can be implemented in 100% Python! ## Bringing It All Together: Wait, VoxelGPT Is a Plugin?! Yep, you read that right. At [~21:58](https://www.youtube.com/watch?v=F-2M37NFavU&t=1318s) in the webinar replay, Brian reveals: VoxelGPT is a plugin! Brian shows how it’s possible to build a full-fledged application like VoxelGPT using FiftyOne’s plugin architecture. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop And he walks through the code used to build VoxelGPT at [~25:12](https://www.youtube.com/watch?v=F-2M37NFavU&t=1512s) in the playback video, explaining which pieces of code perform which tasks to bring VoxelGPT to life as a plugin. You can see VoxelGPT instantly in your browser at [gpt.fiftyone.ai](https://gpt.fiftyone.ai/)! And because VoxelGPT is 100% open source, you can find everything you need to use it locally on [GitHub](https://github.com/voxel51/voxelgpt). ## Live Demo Time! Watch the live demo starting at [~29:58](https://www.youtube.com/watch?v=F-2M37NFavU&t=1798s) in the playback video. In this section, Brian first points our attention to a curated [collection of FiftyOne Plugins on GitHub](https://github.com/voxel51/fiftyone-plugins) that you can add to your local FiftyOne install (or your [FiftyOne Teams](https://voxel51.com/fiftyone-teams) deployment). From this list, Brian demonstrates two ways to use the I/O Plugin, showcasing these two operators: - add\_samples: Use this operator to add samples to an existing dataset by specifying a directory or glob pattern of media paths for which you want to create new samples. - export\_samples: You can use this operator to export your current dataset or view to disk in any supported format. Along the way, you’ll see examples of using the builtin radio group selector, the progress bar shown to the user, the builtin reload dataset operator, and more! ## Other Notes Open source software like FiftyOne doesn’t happen without an amazing community supporting it. Before closing the presentation, Brian explains that we’re always open to new contributions and contributors! Check out the [good first issue](https://github.com/voxel51/fiftyone/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22) label on GitHub. If you are interested in working on one of those, feel free to leave a comment there and we will be happy to help you out in any way. [custom plugins](https://voxel51.com/blog/tag/custom-plugins) [FiftyOne 0.21](https://voxel51.com/blog/tag/fiftyone-0-21) [operators](https://voxel51.com/blog/tag/operators) [plugins](https://voxel51.com/blog/tag/plugins) [VoxelGPT](https://voxel51.com/blog/tag/voxelgpt) [webinar](https://voxel51.com/blog/tag/webinar) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/99a6360654f227b8d75d13381fe015cc7983ee09-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Announcing FiftyOne 0.21 with Operators, Dynamic Groups, and Custom Color Schemes\\ \\ Product & News\\ \\ • \\ \\ Jun 1, 2023](https://voxel51.com/blog/announcing-fiftyone-0-21) [![](https://cdn.sanity.io/images/h6toihm1/production/fedd0c008df994c1834839ea35bc65b34e47fb5d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Double Trouble: Eliminate Image Duplicates with FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 14, 2023](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/bad75ba72dfae8cdefc3d9afe33a1ea9a9c4ec36-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Optical Character Recognition with PyTesseract\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 21, 2023](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-301-lllmstxt|> ## Visualize CVPR 2023 Datasets [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Datasets](https://voxel51.com/blog/category/datasets) Visualize CVPR 2023 Datasets at CVPR 2023! Jun 20, 2023 • 9 min read Article content In this article [How We Found The Datasets](https://voxel51.com/blog/visualize-cvpr-2023-datasets-at-cvpr-2023#68ccf66e60c3) [The Datasets Included in cvpr.fiftyone.ai](https://voxel51.com/blog/visualize-cvpr-2023-datasets-at-cvpr-2023#475bd71e4a3a) [Missed Out? Get Your Datasets into the Dataset Zoo Today!](https://voxel51.com/blog/visualize-cvpr-2023-datasets-at-cvpr-2023#6de57401121a) In this article [How We Found The Datasets](https://voxel51.com/blog/visualize-cvpr-2023-datasets-at-cvpr-2023#68ccf66e60c3) [The Datasets Included in cvpr.fiftyone.ai](https://voxel51.com/blog/visualize-cvpr-2023-datasets-at-cvpr-2023#475bd71e4a3a) [Missed Out? Get Your Datasets into the Dataset Zoo Today!](https://voxel51.com/blog/visualize-cvpr-2023-datasets-at-cvpr-2023#6de57401121a) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _Datasets from CVPR 2023 publications are now browsable through FiftyOne Teams at the Voxel51 booth and at [cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/)_ The annual IEEE/CVF [Conference on Computer Vision and Pattern Recognition](https://cvpr2023.thecvf.com/) (CVPR) is here! Known for its prestigious nature and exceptional creativity, researchers from all over the world share their published findings while continuing to push the boundaries of computer vision, machine learning and AI. Recently, [we published an interesting blog post analyzing the statistics behind this year's papers](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision/). From there, we learned that out of 9,155 submissions, 2,359 papers were accepted to this year’s conference. After some investigating, we found that 59 of the 2,359 accepted papers were associated with a new dataset. That’s 2.5% — impressive, right?! As a team who has built multiple datasets of our own over the last few years, we know full well the labor that goes into making a dataset. Since we are a team that fundamentally believes in the value of visualizing and analyzing during research and development in AI, we decided to make it possible to take a gander at some of these datasets for CVPR 2023 attendees in Vancouver. Over the last couple weeks, our team at Voxel51 added a handful of these datasets to [FiftyOne](https://voxel51.com/); today, we are publishing it live at [cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/). If you’re here in Vancouver in person, stop by Voxel51 Booth 1618 to browse in person. We are intensely grateful to the authors of the datasets who made their work easily accessible and hence includable in [cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/). ![](https://cdn.sanity.io/images/h6toihm1/production/0a4383d5a1693e9d9e3325b7e8bb23a4f8d8b821-1999x1040.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## How We Found The Datasets So, how did we find the 59 new datasets as part of CVPR’s publications? Further, how did we decide which datasets would get uploaded into [cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/)? First, we got the list of all 2,359 accepted conference papers. Once we had access to them, we looked into the abstracts of each paper to quickly determine if their paper involved creating and/or using a new dataset. Upon searching all of the papers this way, we found 59 papers that suggested they did. Although we did this manually without the help of an LLM, we acknowledge we may have missed some. From here, we proceeded to look up information on each dataset including but not limited to license information, dataset accessibility, and dataset type (image, video, mesh, etc.). This process allowed us to prune this list of 59 down to a manageable set of 20 or so compatible datasets from which we ultimately added 9. Some of this reduction is due to the classes of datasets that FiftyOne supports, e.g. whereas it supports images, videos, and point cloud, it does not support meshes. Let us explain more. While researching each of these datasets to upload, we looked at licensing as well as dataset availability. What was surprising to us was that while many of these datasets were marked as available and licensed to use, we were required to get permission from the authors to even download them. This may have been in the form of a Google form, requesting permission from a website to execute download, or even sending an email in some cases clearly outlining the way the dataset was to be used. Going through these datasets, **_we could only begin working with 5 that were licensed properly, given this is a research conference and available to download immediately_**. This led us to the question: How is it possible that out of 2,359 papers at CVPR 2023 and out of the 59 papers that include new datasets to share with the community, _we could only access 5 of them_? Following this realization, we followed up with authors on the other datasets that are compatible with FiftyOne. While many of them responded (and we were able to add several more datasets because of this), it slowed us down. Clearly, if you want your dataset easily adopted by the community, make it easily and openly accessible. Not a new lesson, but one that seems to have fallen flat in many cases. In the following section, you’ll learn a little more about each of the 9 datasets we were able to include in [cvpr.](https://cvpr.fiftyone.ai/) [fiftyone](https://cvpr.fiftyone.ai/) [.ai](https://cvpr.fiftyone.ai/). If you enjoy browsing these datasets, be sure to thank the authors who made their data so readily available to us that we could provide this opportunity to you! If you enjoy working with FiftyOne as a computer vision and AI research framework, you can thank us by [giving us a star on GitHub](https://github.com/voxel51/fiftyone). Further, if you find yourself publishing a dataset in the future, consider the benefits to making that data readily available for download with the publication! Add it to the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/api.html#adding-datasets-to-the-zoo). ## The Datasets Included in [cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/) Listed below are the datasets we’ve ingested into our FiftyOne Teams database to make it easy to explore. Our team of engineers here worked hard to make the data interesting to visualize, so read below on what each dataset contains and the best ways to start exploring it on your own! ### **Spring** The Spring dataset comes from the paper [Spring: A High-Resolution High-Detail Dataset and Benchmark for Scene Flow, Optical Flow, and Stereo by Lukas Mehl et al.](https://arxiv.org/abs/2303.01943) This computer-generated dataset serves as a benchmark for scene flow, optical flow, and stereo data, providing state-of-the-art visual effects within the realistic HD datasets. This dataset is 60x larger than the only scene flow benchmark - [KITTI 2015](https://www.cvlibs.net/datasets/kitti/eval_scene_flow.php?benchmark=stereo) ( [also available to browse in cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/datasets/kitti/samples)). ![](https://cdn.sanity.io/images/h6toihm1/production/b1a63ef4dc2d016162e932faa9b955c7eb273ff3-1999x957.png?auto=format&dpr=2&fit=max&q=75&w=1600) Within [cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/), you’ll find Spring loaded in as [groups](https://docs.voxel51.com/user_guide/app.html#grouping-samples) of images ready to browse by image, providing you with the opportunity to view each image with all of its associated data. ![](https://cdn.sanity.io/images/h6toihm1/production/6de47f99ff98e2e970350e07f43b0607158e5103-1999x1034.png?auto=format&dpr=2&fit=max&q=75&w=1600) **_Where to find Spring at CVPR: Tuesday afternoon, Poster #80_** ### **GeoNet** The GeoNet dataset comes from the paper [GeoNet: Benchmarking Unsupervised Adaptation across Geographies by Tarun Kalluri et al.](https://arxiv.org/abs/2303.15443) GeoNet strives to improve machine learning models by providing a dataset that is more inclusive of under-represented locations in the world. To accomplish this, the dataset contains benchmarks to represent diverse tasks such as scene recognition, image classification, and universal adaptation. This new large-scale dataset is curated from existing datasets by selecting images from Asia and the US and separating them into two domains. ![](https://cdn.sanity.io/images/h6toihm1/production/b8cae46914352231b7da2010e3f9be175b395f84-1999x965.png?auto=format&dpr=2&fit=max&q=75&w=1600) With [cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/), use the [map panel](https://docs.voxel51.com/user_guide/app.html?highlight=map#map-panel) to explore images by region and get a better feel for what is being represented in each location! ![](https://cdn.sanity.io/images/h6toihm1/production/21850a2c624fe1008dd86919d853146f746d0b82-1999x1043.png?auto=format&dpr=2&fit=max&q=75&w=1600) **_Where to find GeoNet at CVPR: Wednesday afternoon, Poster #287_** ### **Mobile-HDR** The Mobile-HDR dataset comes from the paper [Joint HDR Denoising and Fusion: A Real-World Mobile HDR Image Dataset by Shuaizheng Liu et al](https://drive.google.com/file/d/1EnFFwjnHGfKliRTnAMRGZIX-yBN8isx_/view) (linked from [GitHub](https://github.com/shuaizhengliu/Joint-HDRDN)) Using three mobile phones and varying lighting conditions, the authors developed the first HDR image dataset captured by using mobile phone cameras (instead of the typical DSLR cameras used in the daytime to capture images in existing HDR datasets). Their experiments showed that this approach provided an advantage over datasets captured with cameras not on mobile phones. ![](https://cdn.sanity.io/images/h6toihm1/production/e0572dd7c7b3ca115ca5b788565473fb4229103c-1999x956.png?auto=format&dpr=2&fit=max&q=75&w=1600) With [cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/), explore images with various camera settings and from different mobile devices! ![](https://cdn.sanity.io/images/h6toihm1/production/d2923b80e3eafa329cd2baa18d5572572584c0e1-3022x1726.gif?auto=format&dpr=2&fit=max&q=75&w=1600) **_Where to find Mobile-HDR at CVPR: Wednesday afternoon, Poster #154_** ### **MVImgNet** The MVImgNet dataset comes from the paper [MVImgNet: A Large-scale Dataset of Multi-view Images by Xianggang Yu et al](https://arxiv.org/abs/2303.06042). MVImgNet aims to fill the need for an ImgNet-type dataset for 3D vision by providing a large-scale dataset of multi-view images. The full dataset contains 6.5 million frames across 219,188 videos and covers 238 unique classes. With the inclusion of object masks, camera parameters, and point clouds, MVImgNet provides a soft bridge between 2D and 3D computer vision tasks. ![](https://cdn.sanity.io/images/h6toihm1/production/2b6c607aaa1c0ecf32828e9aa5ae66ce544e33df-1999x957.png?auto=format&dpr=2&fit=max&q=75&w=1600) With [cvpr.fiftyone.ai](http://cvpr.fiftyone.ai/), explore the various views of each object, just like you see with this fashionable clutch on a purple background! **_Where to find MVImgNet at CVPR: Wednesday morning, Poster #88_** ![](https://cdn.sanity.io/images/h6toihm1/production/d5111e99d2b5c94053e95b2af3afbe23a2651ad7-1999x1039.png?auto=format&dpr=2&fit=max&q=75&w=1600) ### **ImageNet-E** The ImageNet-E dataset comes from the paper [ImageNet-E: Benchmarking Neural Network Robustness via Attribute Editing by Xiaodan Li et al.](https://arxiv.org/abs/2303.17096) This dataset was created using the authors’ toolkit for object editing, controlling backgrounds, sizes, positions, and directions. The goal of this was to create a rigorous ImageNet benchmark (ImageNet-E) for evaluating image classifier robustness in terms of object attributes. In their research, they noticed that even a small change in background led to an average of 9.23% drop in accuracy of top-1 classification. ![](https://cdn.sanity.io/images/h6toihm1/production/c1275c96c88e41c045b4179160c263ea7b122f64-1999x1041.png?auto=format&dpr=2&fit=max&q=75&w=1600) To build their dataset, they selected the animal classes from ImageNet, leading to a total of 4,352 original images. They chose this subset of ImageNet because the animals tend to appear in nature without messy backgrounds. From here, they applied 11 separate transformations on each image, yielding 47,872 total images in the dataset. With [cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/), use the [saved views](https://docs.voxel51.com/user_guide/app.html#saving-views) to explore images by their transformation, classification, or image. ![](https://cdn.sanity.io/images/h6toihm1/production/4e9fd0a58cfe9ccc3b46d8daeddb0524737f9926-1999x1037.png?auto=format&dpr=2&fit=max&q=75&w=1600) **_Where to find ImageNet-E at CVPR: Thursday morning, Poster #370_** ### **JRDB-Pose** The JRDB-Pose dataset comes from the paper [JRDB-Pose: A Large-scale Dataset for Multi-Person Pose Estimation and Tracking by Edward Vendrow et al.](https://arxiv.org/abs/2210.11940v2) Captured from the social navigation robot JackRabbot, JRDB-Pose consists of challenging scenes with crowded indoor and outdoor locations with varying occlusion types. Importantly, the dataset contains human pose annotations with per-keypoint occlusion labels and track IDs consistent across a given scene. ![](https://cdn.sanity.io/images/h6toihm1/production/0aebd899bb64014d6abedc9f00a85387a8d0b23b-500x750.jpg?auto=format&dpr=2&fit=max&q=75&w=500)![](https://cdn.sanity.io/images/h6toihm1/production/2230f4cba103386247c7aabbaff0293ca8feac5f-1999x1494.png?auto=format&dpr=2&fit=max&q=75&w=1600)![](https://cdn.sanity.io/images/h6toihm1/production/d0824ad7f1a10e96bcd0c4f0eea05ca27b382731-1999x1020.png?auto=format&dpr=2&fit=max&q=75&w=1600) **_Where to find JRDB-Post at CVPR: Tuesday afternoon, Poster #64_**. Plus, explore it instantly at [cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/)! ### **ARKitTrack** The ARKitTrack dataset comes from the paper [ARKitTrack: A New Diverse Dataset for Tracking Using Mobile RGB-D Data by Haojie Zhao et al.](https://arxiv.org/abs/2303.13885) The purpose of ARKitTrack is to provide a new RGB-D tracking dataset for static and dynamic scenes captured with consumer-grade LiDAR scanners. The dataset comes equipped with bounding box annotations, frame-level attributes, and pixel-level target masks. ![](https://cdn.sanity.io/images/h6toihm1/production/7d953903d6d8485ba01ea3481494cf23993c27cd-1999x1026.png?auto=format&dpr=2&fit=max&q=75&w=1600) With [cvpr.](http://cvpr.fiftyone.ai/) [fiftyone](https://cvpr.fiftyone.ai/) [.ai](http://cvpr.fiftyone.ai/), watch the videos and check out the pixel-level target masks! ![](https://cdn.sanity.io/images/h6toihm1/production/5280e496985442d4cc020c5c0bfe8c6d00d4e462-1999x1026.png?auto=format&dpr=2&fit=max&q=75&w=1600) **_Where to find ARKitTrack_** **_at CVPR: Tuesday afternoon, Poster #94_** ### **LLCM** The LLCM dataset comes from the paper [Diverse Embedding Expansion Network and Low-Light Cross-Modality Benchmark for Visible-Infrared Person Re-identification by Yukang Zhang and Hanzi Wang](https://arxiv.org/abs/2303.14481). LLCM aims to capture low-light cross-modality images of different people captured across 9 different RGB (visible) and IR (infrared) cameras. This dataset was curated in order to more effectively benchmark the authors’ novel augmentation network DEEN (diverse embedding expansion network) against other state-of-the-art methods. ![](https://cdn.sanity.io/images/h6toihm1/production/aa7097ec671af7ad21da431cc228b8ab98f9917e-1999x1044.png?auto=format&dpr=2&fit=max&q=75&w=1600) With [cvpr.fiftyone.ai](https://cvpr.fiftyone.ai/), use [groups](https://docs.voxel51.com/user_guide/app.html#grouping-samples) to explore images by each individual and see the different images captured by the various RGB and IR cameras. ![](https://cdn.sanity.io/images/h6toihm1/production/b9587324a014a2f321ceaa784286b87ec68d57f1-1999x1034.png?auto=format&dpr=2&fit=max&q=75&w=1600) **_Where to find LLCM at CVPR: Tuesday morning, Poster #206_** ### **SynSL-120K** The SynSL-120K dataset comes from the paper [A New Benchmark: On the Utility of Synthetic Data with Blender for Bare Supervised Learning and Downstream Domain Adaptation by Hui Tang and Kui Jia.](https://arxiv.org/abs/2303.09165) This dataset consists of 120,000 synthetic images across 10 classes. The authors published this dataset as part of their work on bare supervised learning and downstream domain adaptation. The goal of this work was to present the first comprehensive study on synthetic data learning and demonstrate value to the transfer learning community as a new benchmark. ![](https://cdn.sanity.io/images/h6toihm1/production/12d2f5ec889d97087654271820e286ed9ed7005f-1999x1045.png?auto=format&dpr=2&fit=max&q=75&w=1600) **_Where to find SynSL-120K_** **_at CVPR: Wednesday afternoon, Poster #343_** ### **Other Datasets** For those who are less interested in the latest and greatest CVPR data, you can also view a hand-curated selection of dozens of datasets from the [FiftyOne Dataset Zoo](https://docs.voxel51.com/user_guide/dataset_zoo/index.html). But just because it isn’t new doesn’t mean it isn’t exciting—look closely enough and you’ll find the cat wearing a tie and the llama eating a carrot! ## Missed Out? Get Your Datasets into the Dataset Zoo Today! Don’t see your CVPR dataset in the dataset zoo? Want it to be accessible by the entire community? Great news—we’ve made it simple for you to get your dataset into the FiftyOne Dataset Zoo! First, [you can test it out locally by adding it to your local dataset zoo](https://docs.voxel51.com/user_guide/dataset_zoo/api.html#adding-datasets-to-the-zoo). Once you have it how you want it, you can submit a PR to the [FiftyOne GitHub repo](https://github.com/voxel51/fiftyone) for your dataset to be included with the next release. Want to follow an example? [Check out the PR submitted for the recently added SAMA-COCO dataset](https://github.com/voxel51/fiftyone/pull/2904)! [ARKitTrack](https://voxel51.com/blog/tag/arkittrack) [CVPR](https://voxel51.com/blog/tag/cvpr) [CVPR 2023](https://voxel51.com/blog/tag/cvpr-2023) [FiftyOne Teams](https://voxel51.com/blog/tag/fiftyone-teams) [GeoNet](https://voxel51.com/blog/tag/geonet) [ImageNet-E](https://voxel51.com/blog/tag/imagenet-e) [JRDB-Pose](https://voxel51.com/blog/tag/jrdb-pose) [LLCM](https://voxel51.com/blog/tag/llcm) [Mobile-HDR](https://voxel51.com/blog/tag/mobile-hdr) [MVImgNet](https://voxel51.com/blog/tag/mvimgnet) [Spring](https://voxel51.com/blog/tag/spring) [SynSL-120K](https://voxel51.com/blog/tag/synsl-120k) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/e6fca750f87dc5c6c8adc03d909d17706d81295b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How to Get the Most out of CVPR\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jun 20, 2023](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr) [![](https://cdn.sanity.io/images/h6toihm1/production/6f423b79aea6608b01ca2a85ae76ab401ed8e46a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ ControlNet Training: Improve YourControlNet Training Dataset\\ \\ Computer Vision, Datasets, Tutorials\\ \\ • \\ \\ Jun 29, 2023](https://voxel51.com/blog/conquering-controlnet) [![](https://cdn.sanity.io/images/h6toihm1/production/61b9a72c1362d209b3cb768c4ddc7f84cdf22452-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Automatically Set Up a New ML Project, Pain Free\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Feb 8, 2023](https://voxel51.com/blog/automatically-set-up-a-new-ml-project-pain-free) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) Visualize CVPR 2023 Datasets at CVPR 2023! - Voxel51 <|firecrawl-page-302-lllmstxt|> ## Maximize Your CVPR Experience [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) How to Get the Most out of CVPR Jun 20, 2023 • 8 min read Article content In this article [Be Selective](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr#7ded18f863c4) [Be Observant](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr#cacbfb428f74) [Be Friendly](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr#3595dafa28a4) [Be Active](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr#5291fb92c3f5) [Wrapping Up](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr#18bb9c53b705) In this article [Be Selective](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr#7ded18f863c4) [Be Observant](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr#cacbfb428f74) [Be Friendly](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr#3595dafa28a4) [Be Active](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr#5291fb92c3f5) [Wrapping Up](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr#18bb9c53b705) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _Unsolicited advice from a Professor, Founder, and 20-time CVPR-goer_ Last week, during our weekly one-on-one research meeting, one of my new PhD students innocently asked me a question: "How should I prepare for CVPR?" You see, this would be his first visit to any academic conference let alone CVPR, the prestigious one with [the highest h5 rank of any conference](https://scholar.google.com/citations?view_op=top_venues]) indexed by Google Scholar. I was caught off guard; having spent the last several years focusing on my startup company, [Voxel51](https://voxel51.com/) in an operational capacity (read: it took over all aspects of my life), I had not onboarded a new student in quite some time. So, the natural training a student gets from a mature lab was lost. I quickly responded with some ideas of what I thought he should do to prepare. But, I wasn't satisfied. Some of you in the CVPR community know me as a bit of a lifehacker, constantly trying to develop systems that support extracting the most value out of time spent in any activity. So, this short article is what I've come up with to help make your visit to CVPR 2023 in Vancouver a success. These four strategies capture how I'd recommend making the best of CVPR. I'll list them here and then describe each in more detail below. 1. Be selective: find a signal in the noise by being intentionally selective at what you focus on. 2. Be observant: consider macro trends at the conference, make an effort to summarize them. 3. Be friendly: remember that CVPR is a community gathering of people with similar interests to you. 4. Be active: the conference does not end when you board the flight home; actively recount your experience. The first conference I ever attended was SIGGRAPH 2001 in Los Angeles. I was delighted that my advisor happily supported his students to travel to their first conference without having a publication yet (a practice I do with my own students today). Not being naturally outgoing, I was nervous, really nervous. Of the more than 30,000 attendees, I only knew a couple people there, and was barely through my classes as a grad student to be mildly knowledgeable about the field. I imagine this is how many new students feel today at CVPR. Yet, I had such a fantastic time in LA that August. I was inspired by the IBM Research Augmented Reality ["Everywhere Displays"](https://history.siggraph.org/experience/everywhere-displays-by-pinhanez/) installation on the exhibit floor (yes in 2001). I recall [an amazing tutorial on Kalman Filters](https://history.siggraph.org/learning/an-introduction-to-the-kalman-filter-by-welch-and-bishop/) to which I still refer back today by Gary Bishop and Greg Welch. I vividly remember certain talks like [the one on representing irradiance maps](https://cseweb.ucsd.edu/~ravir/papers/envmap/envmap.pdf) by Ravi Ramamoorthi. But the thing I value most was a one-on-one conversation with David Breen, a reputable faculty member in the computer graphics community; he happily engaged with me; he listened to my thoughts about certain papers; he encouraged various angles of inquiry; but in the end, the simple fact that he was happy to engage with me was infinitely valuable. Since this first conference, I've attended dozens if not hundreds of conferences, each one learning a bit about the field and a bit more about the people. I hope these four strategies help you find value and connection during CVPR this week. I'll be there as well, stop me and say hello! ## Be Selective There are 2,359 papers being presented at the main part of CVPR this year, not including the dozens of workshops and tutorials. I don't know about you, but there is no way on earth I could consume this many ideas in an effective way in just a few days. However, when I first started attending conferences, I tried to do just that. Admittedly, two decades ago, the conferences were an order of magnitude smaller, so perhaps I was not actually crazy. But, no way. I would go to the conference; sit in every talk; stand at every poster; make an effort to try to understand each of the ideas. After even just one day of this, I was incapacitated. Granted, I'm a textbook introvert who expends significant energies around people. Nonetheless, this was a bad idea at its start. What I've learned to do is focus. Before the start of the conference or even before each day of the conference, I select a handful of papers, maybe 3 or 4 each day, a dozen in total perhaps. Pruning 2,359 papers down to 12 is no easy task. But, seeking perfection is not the goal: you can always read another paper after the conference. I mostly select the papers based on my current areas of interest (read: keyword search), allowing myself to add a couple outside of those areas because they might be especially exciting (read: what some of my esteemed friends recommend). Before each conference day, I review those papers to build a mental model of their key ideas. I make an effort to build a sufficient scaffolding along with a question or two that would help me flesh out that scaffolding. Then during the conference day, I focus on those papers from a technical perspective, going deep with them. This type of selectivity has allowed me to be effective at the conference working to extract significant value about those technical ideas that are closest to my own interests, while acknowledging that my attention is a limited resource. Being selective is listed first because I think it is most important: if you fail to be selective, it's unlikely you can engage successfully in the next two strategies effectively, yes including the one about being friendly (I find it easier to break the ice if there is a common technical ground to stand on). There is an exception to the Be Selective. Although I advocate for being very selective when it comes to what papers to go deep with, I also maintain a habit of “walking the floor” at the exhibit. I find it useful to see how technical ideas are mapped to practical applications. This may or may not resonate with you, but I recommend it. ## Be Observant There are a lot of ideas at CVPR. Yet, amongst these ideas are macro-trends that seem to be driving the lion's hare of attention. For the first few CVPRs I attended, these macro trends bounced between bags of features and probabilistic graphical models. Really?! Yes, there's no denying we as a community tend to get pretty gung ho about the latest fad. At this CVPR, one might expect diffusion or similar generative modeling ideas to be among the most popular trends in discussion, although interestingly the CVPR papers themselves were written long before the recent splashes from LLMs. Interesting times. This second strategy recommends making an effort to observe the trends while at the conference. These are high-level collections of similar works, solutions to a particular problem, a new dataset unlocking new directions, etc. It is not necessary (or plausible, see the Be Selective strategy) to understand each of these works in fine technical detail. And, often, I have found that these observations come in unexpected ways: sometimes, they’ll come when on the industry exhibit floor talking about the impact of some trend on a practical problem; other times they’ll come in one-on-one sidebar chats with the person you’re sitting next to after the question-period of an oral. You’ll be surprised. Make an effort to build mental bridges from recent ideas in the community to actual papers at the conference, acknowledging that so much of our community's output is on arXiv now that these trends are likely bigger than any one conference. Importantly, what I try to do is identify not the current biggest trends in the papers and ideas I observe being discussed, but what are the undercurrents that may likely become the trends next year or the year after that. It's not easy, but it is important. ## Be Friendly CVPR is people. The ideas, the papers presented are the work product of one or more individuals over weeks or months. For some this is a first paper; for others it is their hundredth. Sometimes these papers represent a second or even third revision by those people. Often these papers have limitations the individuals knowingly reserve for future work; other times these papers have flaws the individuals unwittingly canonized. None of this is trivial. Authoring an academic paper is participating in the community of scholars. It is putting your ideas out there to influence and contribute to the conversation. Authoring a paper at CVPR is not a touchdown or a home-run; it is not the conclusion of work; it is the continuance of the conversation. And a conference is probably the best venue to actively join the conversation. I mean, where else are you going to have thousands of like-minded, like-trained people to engage? Meet the people. Build connections. Help others. I often refer to CVPR as a sort of family reunion. The people at CVPR may likely be a part of your professional (and sometimes personal) life for decades. Build relationships. Some of the most amazing relationships I have today are the result of a willingness to connect at a conference. Oftentimes, in many of these cases, I cannot tell you when a certain relationship began. At some point, they just seem to have always existed. Start now. ## Be Active As the conference nears its end, you're feeling exhausted. I'm certainly so. Yet, as you board your flight home, it's critical to remain active in your effort to make your attendance at CVPR a success. How do you do that? Be active about your experience at CVPR. What does that mean? This is a different strategy than the other three. It happens after the conference. For each of the other three strategies, make an active effort to deepen what you've gathered. For Be Selective, I highly recommend doing something specific for each of the papers you focused on. For example, you could prepare a slide for each of the papers you focused on and give the presentation to your lab-mates. You could write up a half-page summary describing the paper in your own words. You could download the paper's code or data and work with it, if it exists. For Be Observant, I highly recommend revisiting the observations you made about trends. Consider the trends with your peers. Look in the literature for continued evidence of those trends. Summarize your views of those trends in slides or a narrative. Do you practice [Zettelkasten](https://en.wikipedia.org/wiki/Zettelkasten)? Add some cards! For Be Friendly, remain engaged with your new (and old) connections. Share your recounting on the other strategies above for their feedback. Invite them to visit your lab. Visit their lab. I'm not joking. Do it. It'll pay off in one way or another. ## Wrapping Up I'm looking forward to the week in Vancouver at CVPR 2023. I hope these strategies make your visit a good one. Looking back on the content above, I realize this is pretty general for many contemporary, large academic conferences. I'd love to hear from you if this has helped, or if you have other ideas on life hacking CVPR. Come meet me, and others from the Voxel51 team, at booth 1618 (plus experience all the awesomeness going on there!), or reach out to me on [LinkedIn](https://www.linkedin.com/in/jason-corso/). [CVPR](https://voxel51.com/blog/tag/cvpr) [CVPR 2023](https://voxel51.com/blog/tag/cvpr-2023) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/770b8cfdbd7944916b1195dc11e5b173dbab8e97-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 and the State of Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 18, 2023](https://voxel51.com/blog/cvpr-2023-and-the-state-of-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/ff87d65e5b4e5ef50c5732e905f3f16aff8b0a4e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ CVPR 2023 Survival Guide\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ May 25, 2023](https://voxel51.com/blog/cvpr-2023-survival-guide) [![](https://cdn.sanity.io/images/h6toihm1/production/5d9fe483cd6c19e4ee6246bac487e4716fe99fe7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ 5 Reasons to Visit Voxel51 at CVPR\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jun 9, 2023](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-303-lllmstxt|> ## ControlNet Training Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Datasets](https://voxel51.com/blog/category/datasets), [Tutorials](https://voxel51.com/blog/category/tutorials) ControlNet Training: Improve YourControlNet Training Dataset Jun 29, 2023 • 8 min read Article content In this article [Harness the Power of Diffusion Models with Higher-Quality Data](https://voxel51.com/blog/conquering-controlnet#921f08f83a40) [Setup \| ControlNet Training](https://voxel51.com/blog/conquering-controlnet#43c60a6aa40b) [Select Your ControlNet Training Dataset](https://voxel51.com/blog/conquering-controlnet#dbefdea158f1) [Download the Google Dataset](https://voxel51.com/blog/conquering-controlnet#3b5d79f6ef1e) [Load and Visualize Your ControlNet Training Data](https://voxel51.com/blog/conquering-controlnet#658bd7064d15) [Remove Corrupted Samples](https://voxel51.com/blog/conquering-controlnet#cf346514dcac) [Filter by Aspect Ratio](https://voxel51.com/blog/conquering-controlnet#15f023bfb284) [Filter by Resolution](https://voxel51.com/blog/conquering-controlnet#a18f7429d966) [Ensure Color Pallette](https://voxel51.com/blog/conquering-controlnet#b0006aa72c5a) [Deduplicate the Dataset](https://voxel51.com/blog/conquering-controlnet#00128f08467f) [Validate Image-Caption Alignment](https://voxel51.com/blog/conquering-controlnet#2304fb825494) [The Results: An Improved ControlNet Training Dataset](https://voxel51.com/blog/conquering-controlnet#4a7d47737be4) [What’s Next?](https://voxel51.com/blog/conquering-controlnet#2b71f680e030) In this article [Harness the Power of Diffusion Models with Higher-Quality Data](https://voxel51.com/blog/conquering-controlnet#921f08f83a40) [Setup \| ControlNet Training](https://voxel51.com/blog/conquering-controlnet#43c60a6aa40b) [Select Your ControlNet Training Dataset](https://voxel51.com/blog/conquering-controlnet#dbefdea158f1) [Download the Google Dataset](https://voxel51.com/blog/conquering-controlnet#3b5d79f6ef1e) [Load and Visualize Your ControlNet Training Data](https://voxel51.com/blog/conquering-controlnet#658bd7064d15) [Remove Corrupted Samples](https://voxel51.com/blog/conquering-controlnet#cf346514dcac) [Filter by Aspect Ratio](https://voxel51.com/blog/conquering-controlnet#15f023bfb284) [Filter by Resolution](https://voxel51.com/blog/conquering-controlnet#a18f7429d966) [Ensure Color Pallette](https://voxel51.com/blog/conquering-controlnet#b0006aa72c5a) [Deduplicate the Dataset](https://voxel51.com/blog/conquering-controlnet#00128f08467f) [Validate Image-Caption Alignment](https://voxel51.com/blog/conquering-controlnet#2304fb825494) [The Results: An Improved ControlNet Training Dataset](https://voxel51.com/blog/conquering-controlnet#4a7d47737be4) [What’s Next?](https://voxel51.com/blog/conquering-controlnet#2b71f680e030) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Harness the Power of Diffusion Models with Higher-Quality Data ControlNet has been one of the biggest success stories in ML in 2023. The project, which has racked up 21,000+ stars on [GitHub](https://github.com/lllyasviel/ControlNet), was all the rage at CVPR - and for good reason: it’s an easy, interpretable way to exert influence over the outputs of diffusion models. Rather than running the same diffusion model on the same prompt over and over again, hoping for a reasonable result, you can guide the model via an input map. Hence ControlNet’s cheeky tagline: “Let us control diffusion models!” There are distinct ControlNet models to ‘control’ the output via Canny edge maps, segmentation masks, pose keypoints, and even scribbles. One of the features that makes ControlNet so popular is its accessibility. In an era of hundred-billion parameter foundation models, [ControlNet models are just 1.45GB](https://huggingface.co/lllyasviel/ControlNet-v1-1/tree/main) (the same size as the underlying diffusion model). At a time when models like GPT-3.5 are being [trained on tens of thousands of GPUs](https://www.fierceelectronics.com/sensors/chatgpt-runs-10k-nvidia-training-gpus-potential-thousands-more) at a cost of hundreds of thousands, or even millions of USD, a ControlNet model can be trained at home on a single GPU in just 600 GPU hours! In other words, ControlNet training is so easy and quick _you_ can train your own ControlNet model. Despite ControlNet 1.0’s remarkable success, the model suffered from a few rather unfortunate bugs. Here’s an example: While for most inputs, the model produced stunning, realistic images, in some cases, such as the scenario above, the model’s output was _significantly oversaturated_. When ControlNet’s creator [Lvmin Zhang](https://lllyasviel.github.io/Style2PaintsResearch/lvmin) published ControlNet 1.1, which resolved these issues, the changes were so substantial that he created an [entirely new GitHub repository](https://github.com/lllyasviel/ControlNet-v1-1-nightly)! The craziest part: there were NO CHANGES to the model architecture. What changed? Data quality! That proves that your ControlNet training dataset is critical. It turns out that the data used to train ControlNet 1.0 had a few insidious flaws, including a group of grayscale people that was somehow duplicated thousands of times. The ControlNet 1.1 repo [explicitly mentions this and other problems](https://github.com/lllyasviel/ControlNet-v1-1-nightly#:~:text=in%20Depth%201.1%3A-,The%20training%20dataset%20of%20previous%20cnet%201.0%20has%20several%20problems%20including,the%20training%20dataset%20and%20should%20be%20more%20reasonable%20in%20many%20cases.,-The%20new%20depth). The lesson: **Data reigns supreme. State-of-the-art performance _requires_ high quality data.** In this blog post, I’ll show you how to clean and curate a high quality ControlNet training dataset so you can train your own state-of-the-art ControlNet model. All of the code required to follow along and curate your own image-caption ControlNet training dataset can be found [here](https://github.com/voxel51/fiftyone-examples/blob/master/examples/clean_conceptual_captions.ipynb). If you’re eager, you can jump straight to the highlights: - [Download the ControlNet training dataset](https://voxel51.com/blog/conquering-controlnet#download-dataset) - [Deduplicate the data](https://voxel51.com/blog/conquering-controlnet#deduplicate-data) - [Validate image-caption alignment](https://voxel51.com/blog/conquering-controlnet#validate-alignment) ## Setup \| ControlNet Training The only libraries we will need to clean and curate this data are [pandas](https://pandas.pydata.org/) (for tabular data) and [FiftyOne](https://github.com/voxel51/fiftyone) (for unstructured image data): Additionally, you will need [hashlib](https://docs.python.org/3/library/hashlib.html) for helper functions, and you will probably want [tqdm](https://github.com/tqdm/tqdm) to track progress while downloading images. You can import all of the required modules as follows: ## Select Your ControlNet Training Dataset According to the paper that introduced ControlNet, [Adding Conditional Control to Text-to-Image Diffusion Models](https://arxiv.org/pdf/2302.05543.pdf) (CVPR 2023), the original ControlNet models were trained on “3M image-caption pairs from the internet”. Unfortunately, Lvmin et al. [stop short of revealing](https://github.com/lllyasviel/ControlNet/issues/93#issuecomment-1436077011) precisely what data they use: > _“Given the current complicated situation outside research community, we refrain from disclosing more details about data. Nevertheless, researchers may take a look at that dataset project everyone know.”_ Lvmin Zhang That being said, the information they do reveal lines up closely with [Google’s Conceptual Captions Dataset](https://ai.google.com/research/ConceptualCaptions/): a dataset “consisting of ~3.3M images annotated with captions”. Regardless of whether this is the ControlNet dataset used used, Conceptual Captions will provide us with an illustrative example, and the dataset — when properly cleaned — should allow for training ControlNet models from scratch. ## Download the Google Dataset Google's proposed dataset download process is too cumbersome for my taste: first, you need to download a tab-separated variables (\`.tsv\`) file containing the captions and the urls where the corresponding images can be found, and then you need to download the images from their urls. Lucky for you, I’ve written this code so you don’t have to. Download the `tsv` file by clicking the “Download” button at the bottom of Google’s Conceptual Captions webpage, or by clicking on [this link](https://ai.google.com/research/ConceptualCaptions/download). We can load the `tsv` file as a pandas `DataFrame` in similar fashion to a `csv`, by passing in `sep=t` to specify that the separator is a tab. Give the columns of the `DataFrame` descriptive names: And then hash the url for each entry to generate a unique ID: The `DataFrame` looks like this: We will use these IDs to specify the download locations (filepaths) of images, so that we can associate captions to the corresponding images. If we want to download the images in batches, we can do so as follows: Here we download `batch_size` images starting from `start_index` into the folder `images`, with filename specified by the url hash we generated above. We use `curl` to execute the download operation, and set limits for the time spent attempting to download each image, because some of the links are no longer valid. To download a total of `num_images` images, run the following: ## Load and Visualize Your ControlNet Training Data Once we have the images downloaded into a `images` folder, we can load the images and their captions as a `Dataset` in FiftyOne: This code creates a `Dataset` named “gcc”, which is persisted to the underlying database, and then iterates through the first `num_images` rows of the pandas `DataFrame`, creating a `Sample` with the appropriate filepath and caption. For this walkthrough, I downloaded the first roughly 310,000 images. The first step we should take when inspecting a new computer vision dataset is to visualize it! We can do this by launching the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html): \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## Remove Corrupted Samples When we look at the data, we can immediately see that some of the images are not valid. This may be due to links which are no longer working, interruptions during downloading, or some other issue entirely. Fortunately, we can filter out these invalid images easily. In FiftyOne, the `compute_metadata()` method computes media-type-specific metadata for each sample. For image-based samples, this includes image width, height, and size in bytes. When the media file is nonexistent or corrupted, the metadata will be left as null. We can thus filter out the corrupted images by running `compute_metadata()` and matching for samples where the metadata exists: ## Filter by Aspect Ratio A next step we may want to take is filtering out samples with unusual aspect ratios. If our goal is to control the outputs of a diffusion model, we will likely only be working with images within a certain range of reasonable aspect ratios. We can do this using FiftyOne’s `ViewField`, which allows us to apply arbitrary expressions to attributes of our samples, and then filter based on these. For instance, if we want to discard all images that are more than twice as large in either dimension as they are in the other dimension, we can do so with the following code: For the sake of clarity, this is what the discarded samples look like: If you so choose, you can use a more or less stringent aspect ratio filter! ## Filter by Resolution In a similar vein, we might want to remove the low resolution images. We want to generate _stunning_, _photorealistic_ images, so there is no sense including low resolution images in your ControlNet training dataset. This filter is similar to the aspect ratio filter. If we select 300 pixels as our lowest allowed width and height, the filter takes the form: Once again, you can choose whatever thresholds you like. For clarity, here is a representative view of the discarded images: ## Ensure Color Pallette Looking at the low resolution images, we also might be reminded that some of the images in our dataset are greyscale. We likely want to generate images that are as vibrant as possible, so we should discard the black-and-white images. In FiftyOne, one of the attributes logged in image metadata is the number of channels: color images have three channels (RGB), whereas grayscale images only have one channel. Removing grayscale images is as simple as matching for images with three channels! ## Deduplicate the Dataset Our next task in our data curation quest is to remove duplicate images. When an image is exactly or approximately duplicated in a training dataset, the resulting model may be biased by this small set of overrepresented samples - not to mention the added training costs. We can find approximate duplicates in our dataset by using a model to generate embeddings for our images (we will use a [CLIP model](https://github.com/openai/CLIP) for illustration): Then we create a similarity index based on these embeddings: Finally, we can set a numerical threshold at which point we will consider images approximate duplicates (here we choose 0.3), and only retain one representative from each group of approximate duplicates: ## Validate Image-Caption Alignment Okay, now you’re in luck, because we saved the coolest step for last! Google’s Conceptual Captions Dataset consists of image-caption pairs from the internet. More precisely, “the raw descriptions are harvested from the Alt-text HTML attribute associated with web images”. This is great as an initial pass, but there are bound to be some low-quality captions in there. We may not be able to ensure that all of our captions perfectly describe their images, but we can certainly filter out some poorly aligned image-captions pairs! We will do so using [CLIPScore](https://arxiv.org/pdf/2104.08718.pdf), which is a “reference-free evaluation metric for image captioning”. In other words, you just need the image and the caption. CLIPScore is easy to implement. First, we use Scipy’s cosine distance method to define a _cosine similarity_ function: Then we define a function which takes in a `Sample`, and computes the CLIPScore between image embedding and caption embedding, stored on the samples: Essentially, this expression just lower bounds the score at zero. The scaling factor 100 is the same as used by [PyTorch](https://torchmetrics.readthedocs.io/en/stable/multimodal/clip_score.html). We can then compute the CLIPScore - our measure of alignment between images and captions - by adding the fields to our dataset and iterating over our samples: If we want to see the “least aligned” samples, we can sort by “clip\_score”. To see the most aligned samples, we can do the same, but passing in `reverse=True`: We can then set a CLIPScore threshold depending on how aligned we demand the image-caption pairs are. To my taste, a threshold of 21.8 seemed good enough: The second line clones the view into a new persistent `Dataset` named “gcc\_clean”. ## The Results: An Improved ControlNet Training Dataset After our ControlNet training dataset cleaning and curation, we have turned a relatively mediocre initial dataset of more than 310,000 samples into a high quality dataset with 83,181 samples. The fruits of our labor look like this: We surely haven’t created a _perfect_ dataset — a perfect dataset does not exist. What we have done is addressed all of the data quality issues that plagued ControlNet 1.0, plus a few more, just for good measure. Now you are ready to train your own state-of-the-art ControlNet model! _Note: this post is adapted from a flash session that I presented at CVPR last week!_ ## What’s Next? If you enjoyed this blog post, you may also find the following blog posts interesting: - [AI Telephone — A Battle of Multimodal Models](https://towardsdatascience.com/ai-telephone-a-battle-of-multimodal-models-282b01daf044) - [CVPR 2023 and the State of Computer Vision](https://medium.com/voxel51/cvpr-2023-and-the-state-of-computer-vision-a80315d9b535) - [How I Turned My Company’s Docs into a Searchable Database with OpenAI](https://medium.com/towards-data-science/how-i-turned-my-companys-docs-into-a-searchable-database-with-openai-4f2d34bd8736) [caption](https://voxel51.com/blog/tag/caption) [CLIP](https://voxel51.com/blog/tag/clip) [controlnet](https://voxel51.com/blog/tag/controlnet) [CVPR](https://voxel51.com/blog/tag/cvpr) [CVPR 2023](https://voxel51.com/blog/tag/cvpr-2023) [Diffusion models](https://voxel51.com/blog/tag/diffusion-models) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Google](https://voxel51.com/blog/tag/google) [google conceptual captions](https://voxel51.com/blog/tag/google-conceptual-captions) [multimodal](https://voxel51.com/blog/tag/multimodal) ![](https://cdn.sanity.io/images/h6toihm1/production/d58692baec7c64699806d60d25a0d14f534105fa-300x300.png?auto=format&dpr=2&fit=max&q=75&w=42) Jacob Marks Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/25ed1525cb4a1797abb9edcdf5ac3d7b8239eaf3-2805x1581.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Concept Traversal Plugin for FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Oct 19, 2023](https://voxel51.com/blog/computer-vision-concept-traversal-plugin-for-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/0b2ac80144f5813c10886f51c8a1c34370d01a9f-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visualize CVPR 2023 Datasets at CVPR 2023!\\ \\ Computer Vision, Datasets\\ \\ • \\ \\ Jun 20, 2023](https://voxel51.com/blog/visualize-cvpr-2023-datasets-at-cvpr-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/7ac3933208d45ce7ee8d08eb935a43eadfa2475e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Teaching Androids to Dream of Sheep\\ \\ Computer Vision, Datasets, Product & News, Vector Search\\ \\ • \\ \\ Aug 7, 2023](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-304-lllmstxt|> ## FiftyOne Community Update [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) FiftyOne Computer Vision Community Update – July 2023 Jul 6, 2023 • 5 min read Article content In this article [Community Spotlights](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#d59298eb8fd0) [Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#aa88a58e27b8) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#cb95155c28a0) [Community Contributions](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#835ecea9e8c7) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#7fe0b8be126e) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#846082069763) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#4d3295020293) [Upcoming Computer Vision Events](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#c48ada924bd3) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#d1a318516a25) [Blogs](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#1bce053258ac) [Videos](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#ba947ec2bd78) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#eddeebf22a1d) In this article [Community Spotlights](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#d59298eb8fd0) [Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#aa88a58e27b8) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#cb95155c28a0) [Community Contributions](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#835ecea9e8c7) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#7fe0b8be126e) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#846082069763) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#4d3295020293) [Upcoming Computer Vision Events](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#c48ada924bd3) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#d1a318516a25) [Blogs](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#1bce053258ac) [Videos](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#ba947ec2bd78) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023#eddeebf22a1d) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the monthly blog series where we bring you up to speed on recent happenings in the FiftyOne community and celebrate noteworthy milestones. 🙌 🚀 ## Community Spotlights We love hearing how FiftyOne helps you solve challenges and reach new heights! Curious what sorts of use cases are possible with [FiftyOne](https://voxel51.com/fiftyone/)? Here are just a few highlights from what people in the community have to say. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop [Qinecsa](https://qinecsa.com/) is a leading provider of technology-led, end-to-end pharmacovigilance solutions and is a trusted partner to global life science companies. Qinesca brings together best-in-class technology and scientific expertise to connect life science companies to the right safety solutions. > _“Qinesca’s Pv-trace project assists with object tracking inside warehouses. FiftyOne is used to create annotations of boxes and build models for edge devices.”_ > > — Bhageeratha Tanedar, IT Analyst ![](https://cdn.sanity.io/images/h6toihm1/production/20da7d9cb02d2e0457d19fe610af4b613d71bb25-1024x560.png?auto=format&dpr=2&fit=max&q=75&w=1024) [Taranis](https://www.taranis.com/) crop intelligence solutions use deep agronomic expertise and the most advanced AI and submillimeter image technologies to automatically detect crop threats and provide farmers with leaf-level insights. Taranis serves more than 100 agribusinesses, thousands of farmers, and millions of acres around the world, and is growing rapidly. > _"FiftyOne has been super effective in helping me to quickly understand the issues in my datasets, models, and latent representations."_ > > — Dan Erez, AI Tech Lead ## Community Rewards ![](https://cdn.sanity.io/images/h6toihm1/production/e3ed352e497d0e500ff4a1484b8422b3c9bef5cb-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) Is your organization using FiftyOne to solve interesting computer vision problems? [Share your success story](https://voxel51.com/fiftyone-computer-vision-success-story-submission/) and claim a box of community rewards as a thank you! ## Product Releases Last week we released [FiftyOne v0.21.1.1](https://docs.voxel51.com/release-notes.html#fiftyone-0-21-1) to augment the recent [FiftyOne 0.21](https://voxel51.com/blog/announcing-fiftyone-0-21/) release which featured the introduction of operators, dynamic groups, and custom color schemes. v0.21.1 contains 20+ enhancements and fixes. This point release includes several App optimizations, including [filtering with indexes](https://docs.voxel51.com/user_guide/app.html#leveraging-indexes-while-filtering). ## Community Contributions A quick shoutout to the following community members who contributed to the FiftyOne project with the v0.21.1 release. - “Only match .txt files when reading YOLO labels” by @democat3457 in [#3127](https://github.com/voxel51/fiftyone/pull/3127) - “Replace set with addFields in aggregations py” by @AaronDou in [#3111](https://github.com/voxel51/fiftyone/pull/3111) ## FiftyOne on GitHub GitHub is home to the open source FiftyOne project. Here’s the latest snapshot of what’s happening in the [FiftyOne GitHub repo](https://github.com/voxel51/fiftyone): - Total stars: 3,800+ - Total contributors: 63 - Total used by: 328 - Total forks: 371 - Total issues closed so far: 814 ## FiftyOne Community Slack The FiftyOne Community [Slack channel](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) is where you can join more than 1,800 machine learning engineers and data scientists using FiftyOne to improve the quality of their computer vision data and build better models. Last month alone we had almost 150 new members join. Ask questions, answer questions, or simply follow along with the discussion! To make it easy to catch the highlights, every Friday we recap interesting questions and answers from Slack in [Tips & Tricks blog series](https://voxel51.com/blog/category/tips-tricks/). Recent posts include: - [FiftyOne Computer Vision Tips and Tricks – June 16](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-16-2023/) - [FiftyOne Computer Vision Tips and Tricks – June 2](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-june-2-2023/) - [FiftyOne Computer Vision Embeddings Tips and Tricks – May 26](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-26-2023/) - [FiftyOne Computer Vision Tips and Tricks – May 19](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-may-19-2023/) ## Computer Vision Meetups \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Voxel51 sponsors 13 virtual [Computer Vision Meetups](https://www.meetup.com/pro/computer-vision-meetups/) around the world. (To join, visit the Meetup [link](https://www.meetup.com/pro/computer-vision-meetups/) and scroll down to find the location friendliest to your time zone.) The Computer Vision Meetups are geared towards data scientists, machine learning engineers, and open source enthusiasts who want to expand their knowledge of computer vision and complementary technologies. We put an emphasis on open source software, and speakers who are computer vision practitioners or academics doing research in the field. This month’s Meetups include: ### **July ’23 Computer Vision Meetup (Americas and EU)** ![](https://cdn.sanity.io/images/h6toihm1/production/180cb62c446076425bbf75bebdb895e5f3a31b25-1024x576.png?auto=format&dpr=2&fit=max&q=75&w=1024) - July 13, 10AM PDT (1:00PM EDT) - Unleashing the Potential of Visual Data: Vector Databases in Computer Vision - _Filip Haltmayer (Zilliz)_ - Computer Vision Applications at Scale with Vector Databases - _Zain Hasan (Weaviate)_ - Reverse Image Search for Ecommerce Without Going Crazy - _Kacper Łukawski (Qdrant)_ - Fast and Flexible Data Discovery & Mining for Computer Vision at Petabyte Scale - _Jai Chopra (LanceDB)_ - How to Build Scalable Image and Text Search for Computer Vision Data Using Pinecone & Qdrant - _Jacob Marks (Voxel51)_ - [Register for the Zoom](https://voxel51.com/computer-vision-events/july-2023-computer-vision-meetup/) ### **July ’23 Computer Vision Meetup (APAC)** ![](https://cdn.sanity.io/images/h6toihm1/production/73f73bd527e50c14cff997b9ccc419b56fb9a9c1-1024x576.png?auto=format&dpr=2&fit=max&q=75&w=1024) - July 20, 2023 – 10AM IST / 04:30 UTC - MARLIN: Masked Autoencoder for facial video Representation LearnINg – _Zhixi Cai (Monash University)_ - Unleashing the Potential of Visual Data: Vector Databases in Computer Vision – _Filip Haltmayer (Zilliz)_ - [Register for the Zoom](https://voxel51.com/computer-vision-events/july-20-meetup-apac/) ### **Recapping June’s Meetups** ![](https://cdn.sanity.io/images/h6toihm1/production/2260b433f7da8171c8f20164bf89e715d04cb7a6-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=1200) If you missed the last Meetup, make sure to check out [the recap blog](https://voxel51.com/blog/recapping-the-computer-vision-meetup-june-8-2023/) and watch the playbacks! - [Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation](https://youtu.be/0dAUPNWQIEg) - [Re-Annotating MS COCO, an Exploration of Pixel Tolerance](https://youtu.be/LPj9oXRNBZ8) - [Redefining State-of-the-Art with YOLOv5 and YOLOv8](https://youtu.be/rtLqqAtzp1U) ## Upcoming Computer Vision Events In addition to meetups, we invite you to join us for one or more of these upcoming [events](https://voxel51.com/computer-vision-events/): - July 26 - [Getting Started with FiftyOne Workshop (Americas & EMEA)](https://voxel51.com/computer-vision-events/fiftyone-workshop-7-26-2023/) - Aug 24 - [Computer Vision Meetup (APAC)](https://voxel51.com/computer-vision-events/august-24-meetup-apac/) - Aug 30 - [Getting Started with FiftyOne Workshop (Americas)](https://voxel51.com/computer-vision-events/fiftyone-workshop-8-30-2023/) - Sept 14 - [Computer Vision Meetup (Americas & EMEA)](https://voxel51.com/computer-vision-events/september-14-meetup/) ## New Docs, Blogs, Videos, and Tutorials We want everyone to be successful with FiftyOne, and one of the ways we try to do that is by publishing resources that you might find helpful and handy. Here’s a list of some of the new [documentation](https://docs.voxel51.com/), [blogs](https://voxel51.com/blog/), [videos](https://www.youtube.com/@voxel51/videos), [tutorials](https://docs.voxel51.com/tutorials/index.html), [integrations](https://docs.voxel51.com/integrations/index.html), and [cheat sheets](https://docs.voxel51.com/cheat_sheets/index.html) that you may want to check out. ## Blogs - [VoxelGPT: Your AI Assistant for Computer Vision](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision/) - [Announcing FiftyOne 0.21 with Operators, Dynamic Groups, and Custom Color Schemes](https://voxel51.com/blog/announcing-fiftyone-0-21/) - [Conquering ControlNet](https://voxel51.com/blog/conquering-controlnet/) - [How to Get the Most out of CVPR](https://voxel51.com/blog/how-to-get-the-most-out-of-cvpr/) - [Visualizing the Datasets at CVPR 2023!](https://voxel51.com/blog/visualize-cvpr-2023-datasets-at-cvpr-2023/) - [Webinar Recap: Introducing VoxelGPT and Building Custom Plugins](https://voxel51.com/blog/introducing-voxelgpt-building-custom-plugins/) - [Too Many Pixels, So Little Time: Why You Need Data-Centric Tooling in Your AI Stack](https://voxel51.com/blog/too-many-pixels-so-little-time/) ## Videos - [Introducing VoxelGPT & Building Custom FiftyOne Plugins](https://www.youtube.com/watch?v=F-2M37NFavU) - [Use VoxelGPT to get answers to complex computer vision and machine learning questions](https://www.youtube.com/watch?v=0g_PVdCSOPU) - [Use VoxelGPT to generate insightful views about your computer vision data](https://www.youtube.com/watch?v=rSJAAy5uCg0) - [Use VoxelGPT to search FiftyOne Docs, tutorials and user guides](https://www.youtube.com/watch?v=_CfBYMMdTHk) - [Use VoxelGPT to explore, curate and discover insights about your data](https://www.youtube.com/watch?v=MDeG6lM7hUg) ## Voxel51’s Commitment to Open Source and Community Open source, transparency, and giving back to the computer vision community is what we are all about! Whether it’s developing the open source [FiftyOne computer vision toolset](https://github.com/voxel51/fiftyone) to help engineers and data scientists build high-quality datasets and models, sponsoring [Meetups](https://www.meetup.com/pro/computer-vision-meetups/) to help members boost their computer vision knowledge, or [giving to charitable causes](https://voxel51.com/charitable-giving/) on behalf of the community, Voxel51 is committed to bringing transparency and clarity to the world’s data. [Community Update](https://voxel51.com/blog/tag/community-update) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [open source](https://voxel51.com/blog/tag/open-source) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/869a02098d1898869a250f4a5a23648c479af7da-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – April ‘23\\ \\ Product & News\\ \\ • \\ \\ Apr 6, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/338b38d41e6072dd11af86f21f5309337c52f36b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – May ‘23\\ \\ Product & News\\ \\ • \\ \\ May 5, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4a03c488b86f54b773a6dbe9d9fe21ca6c7da882-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – August 2023\\ \\ Product & News\\ \\ • \\ \\ Aug 3, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-305-lllmstxt|> ## Computer Vision for Vector Search [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news), [Vector Search](https://voxel51.com/blog/category/vector-search) The Computer Vision Interface for Vector Search Jul 12, 2023 • 4 min read Article content In this article [Search through a billion images with a single line of code](https://voxel51.com/blog/the-computer-vision-interface-for-vector-search#e349726d7dba) In this article [Search through a billion images with a single line of code](https://voxel51.com/blog/the-computer-vision-interface-for-vector-search#e349726d7dba) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## _Search through a billion images with a single line of code_ \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop There’s too much data. Data lakes and data warehouses; vast pastures of pixels and oceans teeming with text. Finding the _right_ data is like searching for a needle in a haystack! Vector search engines solve this problem by transforming complex data (raw pixel values of an image, characters in a text document) into entities called embedding vectors. These numerical vectors are then _indexed_, so that you can efficiently search against the raw data. It’s no surprise that vector search engines like [Qdrant](https://qdrant.tech/), [Pinecone](https://www.pinecone.io/), [LanceDB](https://lancedb.com/), and [Milvus](https://milvus.io/) have become essential components in almost any new AI application. If you’re working with image or video data and you want to incorporate vector search into your workflows, there can be quite a bit of overhead: - _How do you implement cross-modal retrieval like searching for images with text?_ - _How do you incorporate traditional search filters like confidence thresholds or class labels?_ - _What about searching over the objects (people, cats, dogs, cars, bikes, …) within your images?_ These are just a few of the many challenges you will encounter. Wait. Stop. Hold your horses. There’s a better way… [FiftyOne](https://fiftyone.ai/) is _the_ computer vision interface for vector search. The FiftyOne open source toolkit now features native integrations with [Qdrant](https://docs.voxel51.com/integrations/qdrant.html), [Pinecone](https://docs.voxel51.com/integrations/pinecone.html), [LanceDB](https://docs.voxel51.com/integrations/lancedb.html), and [Milvus](https://docs.voxel51.com/integrations/milvus.html) so you can use your preferred vector search engine to efficiently search your visual data in a single line of code. Want to find 25 images most similar to the second sample in your dataset with one click? Want to find images of traffic that contain at least one person and one bicycle with a click? You can! ### How Does It Work? 1\. Load your dataset. For the purposes of illustration, we’ll load a subset of the MS COCO validation split. ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4from fiftyone import ViewField as F 5 6dataset = foz.load_zoo_dataset( 7 "coco-2017", 8 split='validation', 9 max_samples = 1000 10) 11session = fo.launch_app(dataset) ``` 2\. Generate the similarity index. In order to search against our media, we need to _index_ the data. In FiftyOne, we can do this via the `compute_similarity()` function. Specify the model you want to use to generate the embedding vectors, and what vector search engine you want to use on the backend. You can also give the similarity index a name, which is useful if you want to run vector searches against multiple indexes. ```python 1## setup lancedb 2pip install lancedb 3 4## generate a similarity index 5## with default model embeddings 6## using LanceDB backend 7fob.compute_similarity( 8 dataset, 9 brain_key="lancedb_index", 10 backend="lancedb", 11) 12 13## setup milvus 14## download and start docker container + 15pip install pymilvus 16 17## generate a similarity index 18## with CLIP model embeddings 19## using Milvus backend 20fob.compute_similarity( 21 dataset, 22 brain_key="milvus_clip_index", 23 backend="milvus", 24 metric="dotproduct" 25) ``` 3\. Search against the index. Now you can run image searches across your entire dataset with a single line of code using the `sort_by_similarity()` method. To find the 25 most similar images to the second image in our dataset, we can pass in the ID of the sample, the number of results we want returned, and the name of the index we want to search against: ```python 1## get ID of first sample 2query = dataset.skip(1).first().id 3 4## find 25 most similar images with LanceDB backend 5sim_view = dataset.sort_by_similarity( 6 query, 7 k=25, 8 brain_key="lancedb_index" 9) 10 11## display results 12session = fo.launch_app(sim_view) ``` You can also do this entirely via UI in the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html): ![](https://cdn.sanity.io/images/h6toihm1/production/1b940b593de97bb349ccc2f24ed0914dc21271e8-2560x1363.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ### Semantic Search Made Simple Gone is the hassle of handling multimodal data. If you want to _semantically search_ your images using natural language, you can use the exact same syntax! Use a multimodal model like CLIP to create your index embeddings, and then pass in a text query instead of a sample ID: ```python 1## semantic query 2query = "kites flying in the sky" 3 4## find 30 most similar images with Milvus backend 5kites_view = dataset.sort_by_similarity( 6 query, 7 k=30, 8 brain_key="milvus_clip_index" 9) 10 11## display results 12session = fo.launch_app(kites_view) ``` This can be especially useful in unstructured data exploration, and digging deeper into your data than existing labels would otherwise allow. This, as well, can be executed entirely in the FiftyOne App: ![](https://cdn.sanity.io/images/h6toihm1/production/a9d0f569e19f7e21ad9db24060bb6bfb833ba072-3452x1828.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ### Pass Along the Prefilters Running vector searches on specific subsets of your data typically involves writing complicated _prefilters_: filters which are passed into the vector search engine to be applied to the dataset prior to vector search. FiftyOne’s [vector search integrations](https://docs.voxel51.com/integrations/index.html) take care of these details for you! If you want to find images that look like “traffic”, but only want this search to be applied to images with a person _and_ a bicycle, you can do this by calling `sort_by_similarity()` on the filtered view: ```python 1## create filtered view 2view = dataset.match_labels(F("label").is_in(["person", "bicycle"])) 3 4## search against this view 5traffic_view = view.sort_by_similarity( 6 "traffic", 7 k=25, 8 brain_key="milvus_clip_index" 9) 10session = fo.launch_app(traffic_view) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/fc44138fe457f55a94dd322d670e3db8fab0e337-3448x1830.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ### Get Your Things in Order All of the aforementioned functionality also works out of the box with object detection patches! When generating a similarity index, all you need to do is pass in the `patches_field` argument — naming the label field where the “objects” can be found — and `compute_similarity()` will generate embedding vectors for each object across all of your images. The vector database indexes these patch embeddings so that you can sort these detections by similarity to a reference object, or a natural language query: ```python 1## setup qdrant 2# pull and start docker container + 3pip install qdrant-client 4 5## create a similarity index for ground truth patches 6## with CLIP model, indexed with Qdrant vector database 7fob.compute_similarity( 8 dataset, 9 patches_field="ground_truth", 10 model="clip-vit-base32-torch", 11 brain_key="qdrant_gt_index", 12 backend="qdrant" 13) 14 15## Search for the object that looks most like a tennis racket 16tennis_view = dataset.to_patches("ground_truth").sort_by_similarity( 17 "tennis racket", 18 k = 25, 19 brain_key= "qdrant_gt_index" 20) 21 22session = fo.launch_app(tennis_view) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/19a971670034fcf643940bd515607d192d84479e-3448x1828.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ### Conclusion No matter how many images or videos you have, you need to be using vector search. [FiftyOne’s native vector search integrations](https://docs.voxel51.com/integrations/index.html) will make your life easier. With FiftyOne, similarity searches are as straightforward as applying more traditional filter and query operations. Mix and match vector search queries with metadata queries to your heart’s content. ### Next Steps ![](https://cdn.sanity.io/images/h6toihm1/production/180cb62c446076425bbf75bebdb895e5f3a31b25-1024x576.png?auto=format&dpr=2&fit=max&q=75&w=1024) If you’re interested in vector search and computer vision, come to the Virtual Computer Vision Meetup on July 13th at 10AM PT, which will be focused entirely on vector search! You can [register here](https://voxel51.com/computer-vision-events/july-2023-computer-vision-meetup/). For comprehensive guides on working with each vector search backend, check out FiftyOne’s integration docs: - [Qdrant integration](https://docs.voxel51.com/integrations/qdrant.html) - [Pinecone integration](https://docs.voxel51.com/integrations/pinecone.html) - [LanceDB integration](https://docs.voxel51.com/integrations/lancedb.html) - [Milvus integration](https://docs.voxel51.com/integrations/milvus.html) For general information about vector search in FiftyOne, check out [sorting by similarity in the FiftyOne App](https://docs.voxel51.com/user_guide/app.html#sorting-by-similarity), and the [FiftyOne Brain User Guide on similarity](https://docs.voxel51.com/user_guide/brain.html#brain-similarity). If you like the open source machine learning library FiftyOne, show your support by giving the project a ⭐ on [GitHub](https://github.com/voxel51/fiftyone) (3,900 stars and counting!) Thanks to the Qdrant and Pinecone teams for contributing integrations to [FiftyOne 0.20](https://medium.com/voxel51/a-google-search-experience-for-computer-vision-data-voxel51-a9ee41390986), and thank you to Ayush Chaurasia and the LanceDB team, and Filip Haltmayer and the Milvus team for contributing vector search engine integrations to the FiftyOne ecosystem in [Fiftyone 0.21.3](https://docs.voxel51.com/release-notes.html#fiftyone-0-21-3)! [CLIP](https://voxel51.com/blog/tag/clip) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [integrations](https://voxel51.com/blog/tag/integrations) [lanceDB](https://voxel51.com/blog/tag/lancedb) [milvus](https://voxel51.com/blog/tag/milvus) [Pinecone](https://voxel51.com/blog/tag/pinecone) [Qdrant](https://voxel51.com/blog/tag/qdrant) [semantic search](https://voxel51.com/blog/tag/semantic-search) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Loading related posts... [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-306-lllmstxt|> ## Vector Search Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps), [Vector Search](https://voxel51.com/blog/category/vector-search) Recapping the Vector Search-Themed Computer Vision Meetup — July 13, 2023 Jul 17, 2023 • 7 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#d9b0bea8c864) [Unleashing the Potential of Visual Data: Vector Databases in Computer Vision](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#6140002330fd) [Computer Vision Applications at Scale with Vector Databases](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#b9098b8fdd9a) [Reverse Image Search for Ecommerce Without Going Crazy](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#7a5198cadd63) [Fast and Flexible Data Discovery & Mining for Computer Vision at Petabyte Scale](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#a3c2a5a24218) [How to Build Scalable Image and Text Search for Computer Vision Data Using Pinecone & Qdrant](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#d454729bf767) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#7ab0c50e16c2) [What’s Next?](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#a8178a385f12) [Get Involved!](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#9dea8b9a09b1) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#d9b0bea8c864) [Unleashing the Potential of Visual Data: Vector Databases in Computer Vision](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#6140002330fd) [Computer Vision Applications at Scale with Vector Databases](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#b9098b8fdd9a) [Reverse Image Search for Ecommerce Without Going Crazy](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#7a5198cadd63) [Fast and Flexible Data Discovery & Mining for Computer Vision at Petabyte Scale](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#a3c2a5a24218) [How to Build Scalable Image and Text Search for Computer Vision Data Using Pinecone & Qdrant](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#d454729bf767) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#7ab0c50e16c2) [What’s Next?](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#a8178a385f12) [Get Involved!](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023#9dea8b9a09b1) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We just wrapped up the July 13, 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/), and if you missed it or want to revisit it, here’s a recap! In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide our monthly donation to charitable causes. The charity that received the highest number of votes this month was [Drink Local Drink Tap](https://drinklocaldrinktap.org/), an international non-profit focused on solving water equity and quality issues through education, advocacy, and affordable, safe clean water sources. We are sending this event’s charitable donation of $200 to Drink Local Drink Tap on behalf of the computer vision community. ![](https://cdn.sanity.io/images/h6toihm1/production/1f04f4fbdc0e326700331f47aa5bbfcad923c988-1024x397.png?auto=format&dpr=2&fit=max&q=75&w=1024) Missed the Meetup? No problem. Here are playbacks and talk abstracts from the event. ## Unleashing the Potential of Visual Data: Vector Databases in Computer Vision https://youtu.be/TNMYdL6mW6M Discover the game-changing role of vector databases in computer vision applications. These specialized databases excel at handling unstructured visual data, thanks to their robust support for embeddings and lightning-fast similarity search. Join us as we explore advanced indexing algorithms and showcase real-world examples in healthcare, retail, finance, and more using the FiftyOne engine combined with the Milvus vector database. See how vector databases unlock the full potential of your visual data. Speaker: [Filip Haltmayer](https://www.linkedin.com/in/filiphaltmayer/) is a Software Engineer at [Zilliz](https://zilliz.com/) working in both software and community development. ### **Resource links** - [What is Milvus](https://zilliz.com/what-is-milvus) - [VectorDBBench](https://github.com/zilliztech/VectorDBBench) - [Open Source Projects](https://zilliz.com/product/open-source-vector-database) - [Milvus + FiftyOne Integration](https://docs.voxel51.com/integrations/milvus.html) - [Slides](https://voxel51.com/wp-content/uploads/2023/07/Voxel51.pdf) ### **Q&A from the talk** - _In the example where you searched to find similar shoes. In this case do you have embeddings generated for the individual objects (I think "patches" in FiftyOne) too? Not just image-wide embeddings?_ - _Does FiftyOne's similarity checker work on Milvus?_ ## Computer Vision Applications at Scale with Vector Databases https://youtu.be/YTIDj7jeRbs Vector Databases enable semantic search at scale over hundreds of millions of unstructured data objects. In this talk Zain will introduce how you can use multi-modal encoders with the Weaviate vector database to semantically search over images and text. This will include demos across multiple domains including e-commerce and healthcare. Speaker: [Zain Hasan](https://www.linkedin.com/in/zainhas/) is a senior developer advocate at [Weaviate](https://weaviate.io/), an open source vector database. ### **Q&A from the talk** - _In regards to multimodal applications, can you discuss a bit more how audio and image can be embedded in the same vector space, e.g. how do you make a lion's roar be in proximity to the image of a lion? Or is a text label used to connect them?_ - _What do you mean by short term recommendation and long term recommendations?_ - _Is it possible to embed movements in a database like Weaviate, e.g. search through security video to find someone hitting someone else?_ - _Can you explain how the symbolic graphs are used in recommendation systems?_ - _What is the typical response time?_ - _What are some of the challenges you faced while trying to combine embeddings from different modalities?_ - _Can you suggest some popular models for fashion and garments use cases?_ ## Reverse Image Search for Ecommerce Without Going Crazy https://youtu.be/OTf-zIFP2o0 Traditional full-text-based search engines have been on the market for a while and we are all currently trying to extend them with semantic search. Still, it might be more beneficial for some ecommerce businesses to introduce reverse image search capabilities instead of relying on text only. However, both semantic search and reverse image may and should coexist! You may encounter common pitfalls while implementing both, so why don't we discuss the best practices? Let's learn how to extend your existing search system with reverse image search, without getting lost in the process! Speaker: [Kacper Łukawski](https://www.linkedin.com/in/kacperlukawski/) is a Developer Advocate at [Qdrant](https://qdrant.tech/). ### **Resource link** - [Slides](https://speakerdeck.com/kacperlukawski/reverse-image-search-for-ecommerce-without-going-crazy) ### **Q&A from the talk** - _Can we use vector embeddings to track objects moving from one frame to another? If there are n objects in one frame and their position has moved in the next frame, can we compare the embeddings of these objects and find the exact location of these objects in the new frame?_ - _How do you rank between textual and vector search? In other words, can we boost one search over the other?_ - _Do you think there would be some benefits of using some kind of RLHF or RL(KPI)F to decide when/how to update the embedding models?_ ## Fast and Flexible Data Discovery & Mining for Computer Vision at Petabyte Scale https://youtu.be/2Ie2RB91Dfg Improving model performance requires methods to discover computer vision data, sometimes from large repositories, whether its similar examples to errors previously seen, new examples/scenarios or more advanced techniques such as active learning and RLHF. LanceDB makes this fast and flexible for multi-modal data, with support for vector search, SQL, Pandas, Polars, Arrow, and a growing ecosystem of tools that you're familiar with. We'll walk through some common search examples and show how you can find needles in a haystack to improve your metrics! Speakers: [Jai Chopra](https://www.linkedin.com/in/jaichopra/) founding product lead and [Ayush Chaurasia](https://www.linkedin.com/in/ayushchaurasia/) founding engineer from [LanceDB](https://lancedb.com/). ### **Resource links** - Get started with LanceDB by reading the [LanceDB Documentation](https://lancedb.github.io/lancedb/) - Join the LanceDB community on [Discord](https://discord.com/invite/zMM32dvNtd) - Read the [LanceDB blog](https://blog.lancedb.com/) for product updates and technical discussions - Check out [YoloExplorer](https://github.com/lancedb/yoloexplorer) for practical examples for querying computer vision data - Read the [Voxel51 LanceDB Integration documentation](https://docs.voxel51.com/integrations/lancedb.html) for how to use LanceDB with Voxel51 ### **Q&A from the talk** - _In regards to autonomous vehicles, it's probably one of the "edge computing" cases of an MV environment, where one has no and very little cloud resources. Are there any strategies you can recommend in narrowing down the tech stack or algorithms to account for this?_ ## How to Build Scalable Image and Text Search for Computer Vision Data Using Pinecone & Qdrant https://youtu.be/5ArqV9bPBP4 Have you ever wanted to find the images most similar to an image in your dataset? What if you haven’t picked out an illustrative image yet, but you can describe what you are looking for using natural language? And what if your dataset contains millions, or tens of millions of images? In this talk Jacob will show you step-by-step how to integrate all the technology required to enable search for similar images, search with natural language, plus scaling the searches with Pinecone and Qdrant. He’ll dive-deep into the tech and show you a variety of practical examples that can help transform the way you manage your image data. Speaker: [Jacob Marks](https://www.linkedin.com/in/jacob-marks/) is a Machine Learning Engineer and Developer Evangelist at Voxel51. ### **Resource links** - [Try FiftyOne](https://try.fiftyone.ai/), no install required! - [“How I Turned My Company’s Docs into a Searchable Database with OpenAI”](https://towardsdatascience.com/how-i-turned-my-companys-docs-into-a-searchable-database-with-openai-4f2d34bd8736) - [“The Computer Vision Interface for Vector Search”](https://medium.com/voxel51/the-computer-vision-interface-for-vector-search-55c14e8a82ac) - [How-to: Sorting by Similarity](https://docs.voxel51.com/user_guide/app.html#sorting-by-similarity) and [Similarity with the FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html#brain-similarity) ### **Q&A from the talk** - _Can the Fiftyone platform be used for non-computer vision applications?_ - _Is there an LLM behind VoxelGPT?_ - _The vector embeddings from an image could be way different from the vector embeddings of an audio that is related/same context as the image. How do those two vectors end up close to each other in the embedding space?_ ## Join the Computer Vision Meetup! Computer Vision Meetup membership has grown to almost [5,000 members](https://www.meetup.com/pro/computer-vision-meetups/) in just under a year! The goal of the Meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. Join one of the 13 Meetup locations closest to your timezone. - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) We have exciting speakers already signed up over the next few months! Become a member of the [Computer Vision Meetup closest to you](https://www.meetup.com/pro/computer-vision-meetups/), then register for the Zoom. ## What’s Next? ![](https://cdn.sanity.io/images/h6toihm1/production/bcba759f93a4a9a16186c72a58238b8f0ed2b958-1024x576.png?auto=format&dpr=2&fit=max&q=75&w=1024) Up next on July 20 at 10 AM IST we have a great line up speakers including: - **MARLIN: Masked Autoencoder for facial video Representation LearnINg –** [Zhixi Cai](https://www.linkedin.com/in/zhixi-cai-b5b042259/), _Monash University_ - **Unleashing the Potential of Visual Data: Vector Databases in Computer Vision** – [Filip Haltmayer](https://www.linkedin.com/in/filiphaltmayer/), _Zilliz_ - **DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data** – [Stephanie Fu](https://www.linkedin.com/in/stephanie-fu/) and [Shobhita Sundaram](https://www.linkedin.com/in/shobsund/), _MIT and_ [Netanel Tamir](https://www.linkedin.com/in/netanel-yakir-tamir-4a9691167/), _Weizmann Institute of Science_ Register for the Zoom [here](https://voxel51.com/computer-vision-events/july-20-meetup-apac/). You can find a complete schedule of upcoming Meetups on [the Voxel51 Events page](https://voxel51.com/computer-vision-events/). ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping him over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [meetup](https://voxel51.com/blog/tag/meetup) [vector database](https://voxel51.com/blog/tag/vector-database) [vector search engines](https://voxel51.com/blog/tag/vector-search-engines) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Loading related posts... [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-307-lllmstxt|> ## Computer Vision Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup — July 20, 2023 Jul 21, 2023 • 6 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#b3a8adda9064) [DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#08f91cc09042) [NIGHTS Synthetic Dataset and FiftyOne Demo](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#df8f63a68282) [MARLIN: Masked Autoencoder for facial video Representation LearnINg](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#fac40141bf9b) [Unleashing the Potential of Visual Data: Vector Databases in Computer Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#b91b59198454) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#eb899b8643f0) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#d5295a5bdfca) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#4a4f8c52fa9f) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#b3a8adda9064) [DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#08f91cc09042) [NIGHTS Synthetic Dataset and FiftyOne Demo](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#df8f63a68282) [MARLIN: Masked Autoencoder for facial video Representation LearnINg](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#fac40141bf9b) [Unleashing the Potential of Visual Data: Vector Databases in Computer Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#b91b59198454) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#eb899b8643f0) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#d5295a5bdfca) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023#4a4f8c52fa9f) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We just wrapped up the July 20, 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/), and if you missed it or want to revisit it, here’s a recap! In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide a $200 donation to charitable causes. There was a two-way tie for the highest number of votes received, so we’ll be making donations of $100 to each of these organizations! ![](https://cdn.sanity.io/images/h6toihm1/production/1f04f4fbdc0e326700331f47aa5bbfcad923c988-1024x397.png?auto=format&dpr=2&fit=max&q=75&w=1024) [Drink Local Drink Tap](https://drinklocaldrinktap.org/) is an international non-profit focused on solving water equity and quality issues through education, advocacy, and affordable, safe clean water sources. ![](https://cdn.sanity.io/images/h6toihm1/production/d20a3e8f5bb9a0d2ef0a52fa7e5f5d04a56a783f-576x284.png?auto=format&dpr=2&fit=max&q=75&w=576) [Education Development Center](https://www.edc.org/) advances lasting solutions to the most pressing educational, health, and workforce challenges across the globe. Missed the Meetup? No problem. Here are playbacks and talk abstracts from the event. ## DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data https://www.youtube.com/watch?v=poB4kU69Ksc Current perceptual similarity metrics compare images in terms of their low-level colors and textures, but fail to capture mid-level similarities in image layout, object pose, and semantic content. To address this gap, we introduce NIGHTS, a synthetic image dataset labeled with human similarity judgments, and DreamSim, a metric tuned to better align with human perception. We analyze how our metric is affected by different visual attributes, and show that it outperforms prior learned metrics and recent large vision models on retrieval and reconstruction tasks. [Stephanie Fu](https://www.linkedin.com/in/stephanie-fu/) recently graduated from MIT with bachelor's degrees in computer science and music and an M.Eng in computer science. Her research interests include computer vision, representation learning, and the connections between human and machine perception. [Shobhita Sundaram](https://www.linkedin.com/in/shobsund/) is a PhD student at MIT in computer science. She is interested in computer vision, particularly generative models and representation learning. Previously she obtained her bachelors in computer science and mathematics from MIT while researching biologically-inspired models for computer vision. [Netanel Tamir](https://www.linkedin.com/in/netanel-yakir-tamir-4a9691167/) is an MSc student at the Weizmann Institute of Science, studying computer science. He’s interested in computer vision, representation learning and psychophysics. ### **Resource links** - [DreamSIM GitHub repository](https://github.com/ssundaram21/dreamsim) - [NIGHTS dataset](https://github.com/ssundaram21/dreamsim/tree/main/dataset) - [DreamSim on ArXiv](https://arxiv.org/abs/2306.09344) ## NIGHTS Synthetic Dataset and FiftyOne Demo https://youtu.be/bqSbNjMZDp8 In this impromptu demo we'll explore NIGHTS, a synthetic image dataset labeled with human similarity judgments using the open source FiftyOne computer vision toolset. [Jacob Marks](https://www.linkedin.com/in/jacob-marks/) is a Machine Learning Engineer and Developer Evangelist at Voxel51. ### **Resource links** - [Explore NIGHTS right now](https://try.fiftyone.ai/datasets/nights/) instantly in your browser! - Look out for a blog deep dive coming soon ## MARLIN: Masked Autoencoder for facial video Representation LearnINg https://youtu.be/TrEWfARH5hI This talk proposes a self-supervised approach to learn universal facial representations from videos, that can transfer across a variety of facial analysis tasks such as Facial Attribute Recognition (FAR), Facial Expression Recognition (FER), DeepFake Detection (DFD), and Lip Synchronization (LS). Our proposed framework, named MARLIN, is a facial video masked autoencoder, that learns highly robust and generic facial embeddings from abundantly available non-annotated web crawled facial videos. As a challenging auxiliary task, MARLIN reconstructs the spatio-temporal details of the face from the densely masked facial regions which mainly include eyes, nose, mouth, lips, and skin to capture local and global aspects that in turn help in encoding generic and transferable features. Through a variety of experiments on diverse downstream tasks, we demonstrate MARLIN to be an excellent facial video encoder as well as feature extractor, that performs consistently well across a variety of downstream tasks including FAR (1.13% gain over supervised benchmark), FER (2.64% gain over unsupervised benchmark), DFD (1.86% gain over unsupervised benchmark), LS (29.36% gain for Frechet Inception Distance), and even in low data regime. [Zhixi Cai](https://www.linkedin.com/in/zhixi-cai-b5b042259/) is a Ph.D. student in the Data Science and Artificial Intelligence Department of Monash University IT Faculty, supervised by Dr. Munawar Hayat, Dr. Kalin Stefanov, and Dr. Abhinav Dhall. ### **Resource links** - [MARLIN GitHub repository](https://github.com/ControlNet/MARLIN) - [MARLIN on ArXiv](https://arxiv.org/abs/2211.06627) - [MARLIN on CVPR](https://openaccess.thecvf.com/content/CVPR2023/html/Cai_MARLIN_Masked_Autoencoder_for_Facial_Video_Representation_LearnINg_CVPR_2023_paper) ## Unleashing the Potential of Visual Data: Vector Databases in Computer Vision https://www.youtube.com/watch?v=EHttz\_aO7F8 Discover the game-changing role of vector databases in computer vision applications. These specialized databases excel at handling unstructured visual data, thanks to their robust support for embeddings and lightning-fast similarity search. Join us as we explore advanced indexing algorithms and showcase real-world examples in healthcare, retail, finance, and more using the FiftyOne engine combined with the Milvus vector database. See how vector databases unlock the full potential of your visual data. [Filip Haltmayer](https://www.linkedin.com/in/filiphaltmayer/) is a Software Engineer at [Zilliz](https://zilliz.com/) working in both software and community development. His contributions mainly revolve around the Milvus and Towhee projects, helping develop both applications and helping grow their respective user bases through client interaction, integrations, and technical talks. ### **Resource links** - [What is Milvus](https://zilliz.com/what-is-milvus) - [VectorDBBench](https://github.com/zilliztech/VectorDBBench) - [Open Source Projects](https://zilliz.com/product/open-source-vector-database) - [Milvus + FiftyOne Integration](https://docs.voxel51.com/integrations/milvus.html) - [Slides](https://voxel51.com/wp-content/uploads/2023/07/Voxel51.pdf) ## Join the Computer Vision Meetup! Computer Vision Meetup membership has grown to almost [5,000 members](https://www.meetup.com/pro/computer-vision-meetups/) in just under a year! The goal of the Meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. Join one of the 13 Meetup locations closest to your timezone. - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) We have exciting speakers already signed up over the next few months! Become a member of the [Computer Vision Meetup closest to you](https://www.meetup.com/pro/computer-vision-meetups/), then register for the Zoom. ## What’s Next? Up next on Aug 10 at 10 AM pacific we have a great line up speakers including: \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - **Neural Congealing: Aligning Images to a Joint Semantic Atlas –** Dolev Ofri-Amar at Weizmann Institute of Science - **Advancing Personalized Medicine and Radiotherapy through AI-Enabled Computer Vision –** Roushanak Rahmat, PhD, AI Researcher - **A Practical Approach to Deep Learning for Computer Vision with Tensorflow 2 –** Folefac Martins at Vinsight and ML Instructor Register for the Zoom [here](https://voxel51.com/computer-vision-events/august-10-meetup/). You can find a complete schedule of upcoming Meetups on [the Voxel51 Events page](https://voxel51.com/computer-vision-events/). ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to me, Meetup co-organizer Jimmy Guerrero on Meetup.com, or ping me over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. — _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [DreamSim](https://voxel51.com/blog/tag/dreamsim) [facial video representation](https://voxel51.com/blog/tag/facial-video-representation) [human perception](https://voxel51.com/blog/tag/human-perception) [MARLIN](https://voxel51.com/blog/tag/marlin) [milvus](https://voxel51.com/blog/tag/milvus) [NIGHTS dataset](https://voxel51.com/blog/tag/nights-dataset) [synthetic data](https://voxel51.com/blog/tag/synthetic-data) [vector database](https://voxel51.com/blog/tag/vector-database) [vector search](https://voxel51.com/blog/tag/vector-search) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/af656fd506d11b555a019950b688830000b62f30-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Recapping the Vector Search-Themed Computer Vision Meetup — July 13, 2023\\ \\ Event Recaps, Vector Search\\ \\ • \\ \\ Jul 17, 2023](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/25ed1525cb4a1797abb9edcdf5ac3d7b8239eaf3-2805x1581.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Concept Traversal Plugin for FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Oct 19, 2023](https://voxel51.com/blog/computer-vision-concept-traversal-plugin-for-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/923449db6dd88d2affc40af7d04463e3a7de43be-1024x576.webp?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the AI, Machine Learning and Data Science Meetup — May 2, 2024\\ \\ Event Recaps\\ \\ • \\ \\ May 3, 2024](https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-may-2-2024) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-308-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – July 21, 2023 Jul 21, 2023 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#fe323569f891) [More effective FiftyOne source code searches](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#38cf780228c1) [Renaming semantic mask paths](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#2a5433a1ea22) [Working with cloud-based notebooks using proxy\_url](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#389e0ae34b04) [Exporting and sharing dataset views](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#dd63ce7e9b67) [Making label boxes “invisible”](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#567230287380) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#8ae42087e403) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#fe323569f891) [More effective FiftyOne source code searches](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#38cf780228c1) [Renaming semantic mask paths](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#2a5433a1ea22) [Working with cloud-based notebooks using proxy\_url](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#389e0ae34b04) [Exporting and sharing dataset views](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#dd63ce7e9b67) [Making label boxes “invisible”](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#567230287380) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023#8ae42087e403) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. As an open source community, the FiftyOne community is open to all. This means everyone is welcome to ask questions, and everyone is welcome to answer them. Continue reading to see the latest questions asked and answers provided! ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## More effective FiftyOne source code searches Community Slack member ZKW asked: _“I can't find the source code of the `dataset.export(...)` definition in the FiftyOne GitHub repo when searching for the keywords ' `def export('`. Is there another method for searching code in this repo?_ Yes! Check out the Docs on [custom exporters](https://docs.voxel51.com/recipes/custom_exporter.html#Writing-Custom-Dataset-Exporters)), the [exporters.py](https://github.com/voxel51/fiftyone/blob/78408e6edbdfab19e6bdf2551f812d6723f35d02/fiftyone/utils/data/exporters.py#L1270) file and the results of [this search](https://github.com/search?q=repo%3Avoxel51%2Ffiftyone+dataset.export&type=code). There are two additional places you can perform searches besides GitHub: - You can use [VoxelGPT](https://voxel51.com/voxelgpt/) to conduct searches using prompts. You can try it out without having to install or configure anything at [try.fiftyone.ai](https://try.fiftyone.ai/). - Plus, there is always the [FiftyOne documentation’s search bar](https://docs.voxel51.com/). https://www.youtube.com/watch?v=\_CfBYMMdTHk ## Renaming semantic mask paths Community Slack member Guillaume asked: _“When I try to rename the semantic mask paths using the code below, I get ‘None’.”_ ```python 1old_root = "/old/path/to/data" 2new_root = "/new/path/to/data" 3print(dataset.first().groundtruth_segmentation.mask_path) 4# /old/path/to/data/masks/0001.png 5view = dataset.set_field("groundtruth_segmentation.mask_path", F("groundtruth_segmentation.mask_path").replace(old_root, new_root)) 6print(view.first().groundtruth_segmentation.mask_path) 7# None ``` “To me, the `mask_path` is a `StringField`, so I don’t understand what is wrong.” A few solutions: Replace `F("groundtruth_segmentation.mask_path")` with `F("$groundtruth_segmentation.mask_path")`. Using the `$` prefix is interpreted with respect to the root of the sample. You could also use `F("mask_path")` instead. The key is that in this operation: ```python 1dataset.set_field("foo.bar.spam.eggs", expr) ``` ..the `expr` is applied one level up from the leaf. In other words, it is applied to `foo.bar.spam`. Learn more about the [`mask`](https://docs.voxel51.com/api/fiftyone.core.labels.html#fiftyone.core.labels.Segmentation.mask) attribute in the FiftyOne Docs. ## Working with cloud-based notebooks using proxy\_url Community Slack member ZKW asked: _“What is an example use case for the new feature that gives you the ability to specify a `proxy_url` (which overrides the server URL) in the App config?”_ The most common use case where this feature would be useful is when working in a cloud-based Jupyter notebook (ideally in a VPN so the URL is not publicly accessible). In order for App cells to display, the session must have this proxy URL for iframes to connect successfully as the session is not running locally. For example, you might have a notebook server running on localhost:5151 that is accessible on a corporate VPN through https://company-notebooks.local/5151 Learn more about the [`proxy_url`](https://docs.voxel51.com/user_guide/config.html#highlight=proxy_url:~:text=in%20notebook%20cells.-,proxy_url,-FIFTYONE_APP_PROXY_URL) setting in the FiftyOne Docs. ## Exporting and sharing dataset views Community Slack member Namrata asked: _“Is there a way to export/share a dataset view with someone else?”_ Yes! [FiftyOne Teams](https://docs.voxel51.com/teams/index.html) is specifically designed to make collaborating a simple and secure operation. With FiftyOne Teams, admins can assign one of four roles to a dataset: Admin, Member, Collaborator, or Guest. ![](https://cdn.sanity.io/images/h6toihm1/production/7faef1a5d4e483d6aa94d652e3f41e27db2c5575-2290x1122.png?auto=format&dpr=2&fit=max&q=75&w=1600) When a user is granted a _Collaborator_ role, it means that they’ll have access to datasets to which they have been invited, and their access will be _Can view_ or _Can edit_ access to datasets. Collaborators cannot create new datasets, clone existing datasets, or view other users of the deployment. Collaborators may export datasets to which they’ve been granted access. ![](https://cdn.sanity.io/images/h6toihm1/production/4eef7adce476ea7eba8963d226079d0b58342f88-1360x1044.png?auto=format&dpr=2&fit=max&q=75&w=1360) Learn more about all of the [FiftyOne Teams roles](https://docs.voxel51.com/teams/roles_and_permissions.html#roles-and-permissions) in the FiftyOne Docs, or dive straight into the [_Collaborator_ role](https://docs.voxel51.com/teams/roles_and_permissions.html#collaborator). ## Making label boxes “invisible” Community Slack member Vagif asked: _“By default, when the FiftyOne App loads a dataset, all the label boxes appear visible on the images. How can I make all the label boxes invisible until their checkboxes are ticked in the sidebar? In other words, I am looking for the equivalent of the "clear shown labels" button, but invoked through code.”_ ![](https://cdn.sanity.io/images/h6toihm1/production/770959406bb9bf5a92af641062f01b3ab8387b26-572x188.png?auto=format&dpr=2&fit=max&q=75&w=572) One way to accomplish this is to tag the samples you are interested in seeing the bounding boxes for and then creating a view for them that you can switch between, easily. Tagging can happen manually in the App or in code. You can use the `tag_samples()` and `untag_samples()` methods to add or remove sample tags from the samples in a view. You can find more information on [tagging contents](https://docs.voxel51.com/user_guide/using_views.html#tagging-contents) in the FiftyOne Docs. ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,800+ [FiftyOne Slack](https://slack.voxel51.com/) members - 3,900+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 4,800+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 350+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [open source](https://voxel51.com/blog/tag/open-source) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/b6b0253fd2feeb31250b038d3be4a7a387d75ea1-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – July 28, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jul 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/79d00d175a8098516cb2f4a7711131cbf322d01a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Finding and Correcting Mistakes – FiftyOne Tips and Tricks – Aug 18, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Aug 18, 2023](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/395ba1a1dacb511782456902b224aa4fa8552dc0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Understanding Grouped Datasets – FiftyOne Tips and Tricks – Sep 1, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 1, 2023](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-309-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – July 28, 2023 Jul 28, 2023 • 5 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#7795cae5e8fa) [Listing Dataset Views and exporting them](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#365a11249523) [Exporting subsets of images with tags](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#4bfaa286e6bd) [Importing datasets in custom formats](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#7a7c640b0992) [Finding FiftyOne Brain keys and deleting runs](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#9b1a6955e110) [Importing images with nested directories](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#4d222c1d1d90) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#aa2ff0f6211e) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#7795cae5e8fa) [Listing Dataset Views and exporting them](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#365a11249523) [Exporting subsets of images with tags](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#4bfaa286e6bd) [Importing datasets in custom formats](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#7a7c640b0992) [Finding FiftyOne Brain keys and deleting runs](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#9b1a6955e110) [Importing images with nested directories](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#4d222c1d1d90) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023#aa2ff0f6211e) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. As an open source community, the FiftyOne community is open to all. This means everyone is welcome to ask questions, and everyone is welcome to answer them. Continue reading to see the latest questions asked and answers provided! ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## Listing Dataset Views and exporting them Community Slack Sabrina Pereira asked: _“How can I see a list of saved views in a dataset? I made some views using the UI and now want to export them as datasets in the COCO format. How can I do that?_ `dataset.list_saved_views()` will list out your saved views. For context, if you find yourself frequently using/recreating certain views, you can use `save_view()` to save them on your dataset under a name of your choice: Then you can conveniently use `load_saved_view()` to load the view in a future session: Besides `list_saved_views()`, you can use `has_saved_view()`, and `delete_saved_view()` to manage your saved views. FiftyOne provides native support for exporting datasets to disk in a variety of common formats, and it can be easily extended to export datasets in custom formats. You can easily export entire datasets as well as arbitrary subsets of your datasets that you have identified by constructing a `DatasetView` into any format of your choice via the basic recipe below. Check out the FiftyOne Docs to learn more about [working with saved views](https://docs.voxel51.com/user_guide/using_views.html#saving-views) and [exporting datasets.](https://docs.voxel51.com/user_guide/export_datasets.html#exporting-fiftyone-datasets) ## Exporting subsets of images with tags Community Slack member Samuel asked: _“What is the recommended way to export images that I tag in the FiftyOne App (without exporting the whole dataset)? In my script (below), I have some code after `session.wait() ` that parses the dataset tags (by iterating through samples, and accessing `sample.tags`), then saves them as a JSON. However, after I tag samples in the App, they don't show up in the `dataset` object after the session. Invariably, `sample.tags` and `fo_dset.tags` are empty for all samples.”_ When you apply tags in the FiftyOne App, you may need to call `dataset.reload()` to force any in-memory samples in your main Python session to pull in the changes you made in the App. This would only be necessary if you are actually holding references to `Sample` objects in-memory in Python. Any `Dataset`-level methods like `count_sample_tags()` will always pull data from the database. So they’ll always be up-to-date without needing to call `reload()` first. In your case you could efficiently access the relevant data as follows: `tags, frame_ids = dataset.values([“tags”, “frame_id”])` Check out the FiftyOne Docs to learn more about working with [tags](https://docs.voxel51.com/user_guide/basics.html#tags) and [samples](https://docs.voxel51.com/user_guide/basics.html#samples). ## Importing datasets in custom formats Community Slack member Hendrik asked: _“I have two million images with JSON metadata attached. The metadata doesn’t follow a standard like COCO, but instead is in a custom format. Is it possible to import my images into FiftyOne and make them searchable by specific fields in the metadata?”_ Yes! The simplest and most flexible approach to loading your data into FiftyOne is to iterate over your data in a simple Python loop, create a `Sample` for each data + label(s) pair, and then add those samples to a `Dataset`. FiftyOne provides label types for common tasks such as classification, detection, segmentation, and many more. Here’s an image classification example to give you a sense of the basic workflow. Check out the FiftyOne Docs to learn more about [custom formats](https://docs.voxel51.com/user_guide/dataset_creation/index.html#custom-formats), [label types](https://docs.voxel51.com/user_guide/using_datasets.html#using-labels), and writing [custom dataset importers](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#custom-dataset-importer). ## Finding FiftyOne Brain keys and deleting runs Community Slack member Samuel asked: _“I'm having trouble wrangling FiftyOne Brain keys. I'm often recomputing embeddings and similarity, but don't want to have to specify a new name every time. So, I use a default brain key name like:_ _`fob.compute_similarity(dset, embeddings=embeddings, brain_key="img_embed_sim")`_ _However, this leads to frequent errors:_ _`"ValueError: Brain method run with key 'img_embed_sim' already exists"`_ _I was trying to delete the existing Brain method when this happens, but I can't find the method called " `img_embed_sim`" anywhere in my dataset. Any ideas on how I can resolve this? Also if I don't specify a Brain key, I don't seem to be able to sort by similarity in the App.”_ You must provide a `brain_key` in order for the run to be persisted permanently on your dataset and accessible in the FiftyOne App. So, that part of your question is by design. You can then use `dataset.list_brain_runs()` to see the existing Brain keys on your dataset to avoid duplicates. This key needs to be something descriptive of what the run contains so that you know what you’re choosing when you access the run from the App later. The relevant “delete” methods to consider are: - [delete\_brain\_run(brain\_key)](https://docs.voxel51.com/api/fiftyone.core.clips.html#fiftyone.core.clips.ClipsView.delete_brain_run): deletes the brain method run with the given key from this collection - [delete\_brain\_runs()](https://docs.voxel51.com/api/fiftyone.core.clips.html#fiftyone.core.clips.ClipsView.delete_brain_runs): deletes all brain method runs from this collection Check out the FiftyOne Docs to learn more about the capabilities of the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html?highlight=brain). ## Importing images with nested directories Community Slack member Maor asked: _“How do I upload 10k images where the path contains directories with directories with images?”_ Assuming you want the directories to be imported as tags, you can try: Check out the FiftyOne Docs to learn more about [loading datasets from disk](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#loading-datasets-from-disk). ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,900+ [FiftyOne Slack](https://slack.voxel51.com/) members - 3,900+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 4,900+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 350+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [open source](https://voxel51.com/blog/tag/open-source) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/7005eb6825eeb5d99277b093c5b8c62615dcbc38-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – July 21, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jul 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/79d00d175a8098516cb2f4a7711131cbf322d01a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Finding and Correcting Mistakes – FiftyOne Tips and Tricks – Aug 18, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Aug 18, 2023](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/395ba1a1dacb511782456902b224aa4fa8552dc0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Understanding Grouped Datasets – FiftyOne Tips and Tricks – Sep 1, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 1, 2023](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-310-lllmstxt|> ## FiftyOne Community Update [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) FiftyOne Computer Vision Community Update – August 2023 Aug 3, 2023 • 5 min read Article content In this article [Community Spotlights](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#0773607c46d1) [Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#7e01ef87d468) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#f29ab80043f9) [Community Integrations](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#eb2b1927c7a1) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#40ba8dc574e4) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#74cf15d291dc) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#a986a9ad1f39) [Upcoming Computer Vision Events](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#908520757fd6) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#b15b561da40c) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#92b80f7bae5e) In this article [Community Spotlights](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#0773607c46d1) [Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#7e01ef87d468) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#f29ab80043f9) [Community Integrations](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#eb2b1927c7a1) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#40ba8dc574e4) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#74cf15d291dc) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#a986a9ad1f39) [Upcoming Computer Vision Events](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#908520757fd6) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#b15b561da40c) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-august-2023#92b80f7bae5e) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the monthly blog series where we bring you up to speed on recent happenings in the FiftyOne community and celebrate noteworthy milestones. 🙌 🚀 ## Community Spotlights We love hearing how FiftyOne helps you solve challenges and reach new heights! Curious what sorts of use cases are possible with [FiftyOne](https://voxel51.com/fiftyone/)? Here’s a highlight from the open source FiftyOne community. ![](https://cdn.sanity.io/images/h6toihm1/production/0ea53409154dacf674cfabbec931e9f08426ff5c-422x130.png?auto=format&dpr=2&fit=max&q=75&w=422) > _“Fast Code AI relies on FiftyOne for managing our data lifecycle. Upon receiving initial data from our data labeling partner, IndiVillage, we load it into FiftyOne and scrutinize the labeling for any initial inconsistencies. We then provide feedback and obtain the next batch of data. This is where we train our baseline models and identify any outliers. We sort the data by their ground truth class confidence values in increasing order and display it in FiftyOne. This helps us easily spot and correct any incorrect labels. With approximately 30 attributes per image and over 40 possible values for each attribute, visualizing the data elsewhere is challenging. We greatly appreciate FiftyOne's contribution to our workflow, as it has significantly increased our efficiency. We cannot imagine managing our data without this tool, as it offers an unmatched level of customization and functionality.”_ > > — Arjun Jain, Founder and Chief Scientist, Fast Code AI ## Community Rewards ![](https://cdn.sanity.io/images/h6toihm1/production/e3ed352e497d0e500ff4a1484b8422b3c9bef5cb-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) Is your organization using FiftyOne to solve interesting computer vision problems? [Share your success story](https://voxel51.com/fiftyone-computer-vision-success-story-submission/) and claim a box of community rewards as a thank you! ## Product Releases In July there were a few point releases to both the open source and Teams versions of FiftyOne. Here are the links to the releases learn more: ### Open Source FiftyOne - [FiftyOne 0.21.4](https://docs.voxel51.com/release-notes.html#fiftyone-0-21-4-teams-sdk-0-13-4) \- July 14 - [FiftyOne 0.21.3](https://docs.voxel51.com/release-notes.html#fiftyone-0-21-3) \- July 12 - [FiftyOne 0.21.2](https://docs.voxel51.com/release-notes.html#fiftyone-0-21-2) \- July 3 ### FiftyOne Teams - [FiftyOne Teams 1.3.3](https://docs.voxel51.com/release-notes.html#fiftyone-teams-1-3-3) \- July 12 - [FiftyOne Teams 1.3.2](https://docs.voxel51.com/release-notes.html#fiftyone-teams-1-3-2) \- July 5 ## Community Integrations July brought us two new exciting vector search integrations! ### Milvus [Milvus](https://milvus.io/) is one of the most popular vector databases available, and we’ve made it easy to use Milvus’s vector search capabilities on your computer vision data directly from FiftyOne! FiftyOne provides an API to create Milvus collections, upload vectors, and run similarity queries, both [programmatically](https://docs.voxel51.com/integrations/milvus.html#milvus-query) in Python and via point-and-click in the App. Follow these [simple instructions](https://docs.voxel51.com/integrations/milvus.html#milvus-setup) to get started using Milvus + FiftyOne. Shout out to [Filip Haltmayer](https://www.linkedin.com/in/filiphaltmayer/) from Zilliz for creating the integration. ### LanceDB [LanceDB](https://www.lancedb.com/) is a serverless vector database with deep integrations with the Python ecosystem. It requires no setup and is free to use. FiftyOne provides an API to create LanceDB tables and run similarity queries, both [programmatically](https://docs.voxel51.com/integrations/lancedb.html#lancedb-query) in Python and via point-and-click in the App. Shout out to [Ayush Chaurasia](https://www.linkedin.com/in/ayushchaurasia/) from LanceDB for creating the integration. ## FiftyOne on GitHub GitHub is home to the open source FiftyOne project. Here’s the latest snapshot of what’s happening in the [FiftyOne GitHub repo](https://github.com/voxel51/fiftyone): - Total stars: 3,900+ - Total contributors: 65 - Total used by: 359 - Total forks: 391 - Total issues closed so far: 844 ## **FiftyOne Community Slack** The FiftyOne Community [Slack channel](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) is where you can join more than 1,900 machine learning engineers and data scientists using FiftyOne to improve the quality of their computer vision data and build better models. Last month alone we had almost 150 new members join. Ask questions, answer questions, or simply follow along with the discussion! To make it easy to catch the highlights, every Friday we recap interesting questions and answers from Slack in [Tips & Tricks blog series](https://voxel51.com/blog/category/tips-tricks/). Recent posts include: - [FiftyOne Computer Vision Tips and Tricks – July 28](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-28-2023/) - [FiftyOne Computer Vision Tips and Tricks – July 21](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-july-21-2023/) ## Computer Vision Meetups \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Voxel51 sponsors 13 virtual [Computer Vision Meetups](https://www.meetup.com/pro/computer-vision-meetups/) around the world. (To join, visit the Meetup [link](https://www.meetup.com/pro/computer-vision-meetups/) and scroll down to find the location friendliest to your time zone.) The Computer Vision Meetups are geared towards data scientists, machine learning engineers, and open source enthusiasts who want to expand their knowledge of computer vision and complementary technologies. We put an emphasis on open source software, and speakers who are computer vision practitioners or academics doing research in the field. This month’s Meetups include: ### August Computer Vision Meetup ![](https://cdn.sanity.io/images/h6toihm1/production/d6369675cafcd0764dd9a11b26c31b8570ee5a9b-1024x576.png?auto=format&dpr=2&fit=max&q=75&w=1024) - Neural Congealing: Aligning Images to a Joint Semantic Atlas – _[Dolev Ofri-Amar](https://www.linkedin.com/in/dolev-ofri/) at Weizmann Institute of Science_ - Advancing Personalized Medicine and Radiotherapy through AI-Enabled Computer Vision – _[Roushanak Rahmat](https://www.linkedin.com/in/roushanakrahmat/), PhD, AI Researcher_ - A Practical Approach to Deep Learning for Computer Vision with Tensorflow 2 – _[Folefac Martins](https://www.linkedin.com/in/folefac-martins-8a792617b/) at Vinsight and ML Instructor_ ### August Computer Vision Meetup - APAC ![](https://cdn.sanity.io/images/h6toihm1/production/e2f43b8d5085f67b9b26af5e21eb5d81a5b9d74d-1024x576.png?auto=format&dpr=2&fit=max&q=75&w=1024) - Removing Backgrounds Automatically or with a User’s Language – _[Jizhizi Li, PhD](https://www.linkedin.com/in/jizhizili/), University of Sydney_ - Self-Supervised Representative Learning for Action Recognition in Videos – _[Vidhya Vinay](https://www.linkedin.com/in/vidhya-vinay-2700824/), Co-Founder of Streamingo.ai_ - AI at the Edge: Optimizing Deep Learning Models for Real-World Applications _– [Raz Petel](https://www.linkedin.com/in/raz-petel-ai/), SightX_ ### Recapping July’s Meetups If you missed any of last month’s Meetups, you can get the executive summaries and links to the video playbacks here: - [Recapping the Computer Vision Meetup — July 20](https://voxel51.com/blog/recapping-the-computer-vision-meetup-july-20-2023/) - [Recapping the Vector Search-Themed Computer Vision Meetup — July 13](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023/) ## Upcoming Computer Vision Events In addition to Meetups, we invite you to join us for one or more of these upcoming [events](https://voxel51.com/computer-vision-events/): - Aug 9 - ​​ [Combining the Power of LLMs with Computer Vision](https://www.eventbrite.com/e/combining-the-power-of-llms-with-computer-vision-jacob-marks-voxel51-tickets-670960048567) - Aug 16 - [FiftyOne Community Office Hours & AMA](https://us02web.zoom.us/meeting/register/tZMtc-ivrjMiHdyXfN3K0SpAw69XQRZ0EdSD#/registration) - Aug 30 - [Getting Started with FiftyOne Workshop](https://voxel51.com/computer-vision-events/fiftyone-workshop-8-30-2023/) ## New Docs, Blogs, Videos, and Tutorials We want everyone to be successful with FiftyOne, and one of the ways we try to do that is by publishing resources that you might find helpful and handy. Here’s a list of some of the new [documentation](https://docs.voxel51.com/), [blogs](https://voxel51.com/blog/), [videos](https://www.youtube.com/@voxel51/videos), [tutorials](https://docs.voxel51.com/tutorials/index.html), [integrations](https://docs.voxel51.com/integrations/index.html), and [cheat sheets](https://docs.voxel51.com/cheat_sheets/index.html) that you may want to check out. ### Blogs - [The Computer Vision Interface for Vector Search](https://voxel51.com/blog/the-computer-vision-interface-for-vector-search/) ### Videos - [Introducing VoxelGPT & Building Custom FiftyOne Plugins](https://www.youtube.com/watch?v=F-2M37NFavU) - [Use VoxelGPT to get answers to complex computer vision and machine learning questions](https://www.youtube.com/watch?v=0g_PVdCSOPU) - [Use VoxelGPT to generate insightful views about your computer vision data](https://www.youtube.com/watch?v=rSJAAy5uCg0) - [Use VoxelGPT to search FiftyOne Docs, tutorials and user guides](https://www.youtube.com/watch?v=_CfBYMMdTHk) - [Use VoxelGPT to explore, curate and discover insights about your data](https://www.youtube.com/watch?v=MDeG6lM7hUg) ## Voxel51’s Commitment to Open Source and Community Open source, transparency, and giving back to the computer vision community is what we are all about! Whether it’s developing the open source [FiftyOne computer vision toolset](https://github.com/voxel51/fiftyone) to help engineers and data scientists build high-quality datasets and models, sponsoring [Meetups](https://www.meetup.com/pro/computer-vision-meetups/) to help members boost their computer vision knowledge, or [giving to charitable causes](https://voxel51.com/charitable-giving/) on behalf of the community, Voxel51 is committed to bringing transparency and clarity to the world’s data. [Community Update](https://voxel51.com/blog/tag/community-update) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [open source](https://voxel51.com/blog/tag/open-source) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/869a02098d1898869a250f4a5a23648c479af7da-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – April ‘23\\ \\ Product & News\\ \\ • \\ \\ Apr 6, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/338b38d41e6072dd11af86f21f5309337c52f36b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – May ‘23\\ \\ Product & News\\ \\ • \\ \\ May 5, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/7685a2b9b8681c3c641a1118ad1c4685f1af21b7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – July 2023\\ \\ Product & News\\ \\ • \\ \\ Jul 6, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-july-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-311-lllmstxt|> ## Teaching Androids to Dream [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Datasets](https://voxel51.com/blog/category/datasets), [Product & News](https://voxel51.com/blog/category/product-news), [Vector Search](https://voxel51.com/blog/category/vector-search) Teaching Androids to Dream of Sheep Aug 7, 2023 • 10 min read Article content In this article [DreamSim bridges the gap between human and machine perceptual similarity](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#9f5065c45291) [The Spectrum of Similarity](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#b60664809e67) [Dark Days: A History of Perceptual Similarity Metrics](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#d6b980fe47bb) [NIGHTS Time: A Benchmark Dataset for Perceptual Similarity](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#868ba277ce36) [What Dreams Are Made Of](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#ee056983af24) [DreamSim in Action](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#1ef3cb16f4c0) In this article [DreamSim bridges the gap between human and machine perceptual similarity](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#9f5065c45291) [The Spectrum of Similarity](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#b60664809e67) [Dark Days: A History of Perceptual Similarity Metrics](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#d6b980fe47bb) [NIGHTS Time: A Benchmark Dataset for Perceptual Similarity](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#868ba277ce36) [What Dreams Are Made Of](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#ee056983af24) [DreamSim in Action](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#1ef3cb16f4c0) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## DreamSim bridges the gap between human and machine perceptual similarity Have you ever heard the saying “one of these things is not like the other”? As a kid, I fondly remember watching the _Sesame Street_ segment of the same name. Each round, the viewer was presented with four objects: three matched, but one was always different. Children tuning in from around the world were tasked with identifying the imposter, so to speak — determining which object was _not_ like the others. Silly as it sounds, this children’s game speaks to something remarkably profound: humans often demonstrate a shared sense of visual similarity. The red that you see may not be the red that I see — especially so because I’m colorblind(!) — but individuals of all ages, raised in disparate environments, are often united in their innate sense of perceptual similarity. This ability to perceive and understand similarities is a cornerstone of human cognition. Computers, however, have long struggled at these types of tasks. For many years, the most successful approaches to computationally capturing notions of perceptual similarity involved comparing images on a pixel-by-pixel or patch-by-patch basis. But over the past half decade this picture has changed dramatically with the application of neural networks. Now, a team of researchers from MIT, the Weizmann Institute, and Adobe seems to have cracked the code on human perceptual similarity judgements. Their perceptual similarity metric [DreamSim](https://dreamsim-nights.github.io/) achieves state-of-the-art levels across a wide spectrum of features. To train DreamSim, they developed a new benchmark dataset called [NIGHTS](http://try.fiftyone.ai/datasets/nights/) (Novel Image Generations with Human-Tested Similarity). You can play around with the dataset in your browser for free at [try.fiftyone.ai/datasets/nights](http://try.fiftyone.ai/datasets/nights/)! This article is organized as follows: - [The Spectrum of Similarity](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#similarity-spectrum) - [History of Perceptual Similarity Metrics](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#similarity-history) - [NIGHTS Dataset](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#nights) - [DreamSim Metric](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#dreamsim) - [DreamSim in Action](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep#dreamsim-in-action) ## The Spectrum of Similarity The path to DreamSim was paved with many partial successes as well as many perceptual gaps. There is a broad spectrum encompassing what it means for two images to be similar. On one extreme, there are low-level distortions which affect the values of individual pixels. An example of this can be seen in the first image triplet in the above figure. These distortions are perceptual, and can be viewed as corruptions of the original image. On the other extreme, high-level distortions are more conceptual, as exemplified by the third triplet in the figure above. The middle of the spectrum — so-called “mid-level” distortions — are perceptual like low-level distortions. Unlike low-level distortions, however, mid-level distortions are more _semantic_. For an example of such a mid-level distortion, see the center triplet in the above figure. ## Dark Days: A History of Perceptual Similarity Metrics Before diving into NIGHTS and DreamSim, it’s worth reviewing some influential steps toward representing similarity. ### Peak Signal-to-Noise Ratio The first and simplest attempt at approximating similarity in images was to use Peak Signal-to-Noise Ratio (PSNR). This approach, which borrows heavily from the field of signal processing, treats one image as a reference, _uncorrupted_ image, and another as a corrupted or _distorted_ image. The difference between reference and distorted image is calculated on a pixel-wise basis, so that two images have lower PSNR if the values of pixels at the same locations are more similar. While PSNR can be effective in a limited range of scenarios (especially when working with compressed images), it generally fails to capture the complexities of human perceptual similarity. The similarity metric is sensitive to noise and high dynamic ranges in images, and it does not account for structural or semantic differences in images. _Note: PSNR is also intimately related to Mean Squared Error (MSE)_ ### Structural Similarity [Proposed in 2004](https://www.cns.nyu.edu/pub/lcv/wang03-preprint.pdf) as a measure for assessing image quality, structural similarity (SSIM) overcomes some of PSNR’s limitations, taking structural information into account. In particular, the structural similarity measure utilizes local [luminance](https://en.wikipedia.org/wiki/Luminance), contrast, and structure comparisons when assessing the quality of a given image with respect to a reference. When used as a perceptual similarity metric, SSIM outperforms PSNR. Nevertheless, it still operates on the level of pixels, failing to capture notions of higher level structure in images. As a result, SSIM doesn’t always align with human perceptual judgements. ### Learned Perceptual Image Patch Similarity Motivated by successes in using deep neural network features in the training loss for image synthesis tasks, in 2018 a group of researchers from UC Berkeley, Adobe Research, and OpenAI trained a neural network to learn human perceptual judgements. Their metric, [Learned Perceptual Image Patch Similarity](https://github.com/richzhang/PerceptualSimilarity#1-learned-perceptual-image-patch-similarity-lpips-metric) (LPIPS), goes beyond individual pixels to compare 64x64 patches within images. By encoding these patches in deep features, LPIPS is able to capture more nuanced notions of similarity. However, LPIPS is not designed to encode high level semantic information, so images with similar look and feel but highly disparate connotations can be regarded as “close”. ### Image-level Embeddings This, in turn, has led researchers to use image-level embeddings to capture semantic information. Models like CLIP and DINO have shown success at these image-to-image semantic tasks, but are not designed to distinguish fine-grained visual features within an image. CLIP, for instance, is trained to minimize the distance between the text embedding for a concept and an image that portrays the same concept. This places a premium at high level similarity, at the expense of low level similarity. DreamSim bridges the gap between low level measures of similarity like LPIPS and high level, conceptual similarity obtained from full image embeddings! ## NIGHTS Time: A Benchmark Dataset for Perceptual Similarity As is often the case in machine learning, data has been a limiting factor in human perceptual similarity tasks. Previous datasets have typically focused exclusively on either low level perceptual differences or high level conceptual content. For low-level similarity, the most popular dataset is the [Berkeley-Adobe Perceptual Patch Similarity](https://github.com/richzhang/PerceptualSimilarity#c-about-the-dataset) (BAPPS) dataset, which was introduced in concert with LPIPS. The BAPPS dataset consists of two types of human perceptual judgment. In the first set, dubbed two alternative forced choices (2AFC), human evaluators were given a reference patch and two distorted patches, and asked to select the distorted patch that most closely matched the reference. To validate the 2AFC judgments, the second set of experiments asked humans to make a split-second decision about whether two image patches (one reference, one distortion) were the same. This second variety is referred to as just noticeable differences (JND). For conceptual similarity in images, the gold standard is [THINGS](https://things-initiative.org/), which consists of 4.7 million image triplets spanning 1,854 object and concept categories. Each triplet was evaluated by a single human in odd-one-out fashion: in other words, one of these things is not like the others! To develop their state-of-the-art perceptual similarity metric, the DreamSim team created the most comprehensive perceptual similarity dataset to date: [NIGHTS](http://try.fiftyone.ai/datasets/nights/). The researchers used [stable diffusion](https://huggingface.co/spaces/stabilityai/stable-diffusion) to generate images in the label classes of common datasets like [ImageNet](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#imagenet-2012), and then generate variations for these base images across a variety of axes, including pose, perspective, and shape. Importantly, the base images and their variations share semantic commonality: the differences only span mid-level variations. Using this procedure, the team generated 100,000 image triplets. From there, the team collected 2AFC judgements from human evaluators on [MTurk](https://www.mturk.com/). Each triplet was given to multiple human evaluators, and only the triplets with unanimous human judgements were retained. These triplets are referred to as _cognitively impenetrable_, and only these are kept in an effort to capture something _shared_ across humans, automatic (requiring little, if any cognition) and stable over time. After filtering for only cognitively impenetrable triplets, the team was left with 20,019 high quality triplets. They also ran a number of JND evaluations and observed strong alignment between the two methods of attaining perceptual judgments. ## What Dreams Are Made Of With their novel NIGHTS dataset created and curated, the researchers were ready to revisit the task of computationally representing human perceptual similarity. As with LPIPS, they set out to learn a similarity metric. For each triplet, the team computed the cosine distance (a measure of directional agreement) between the model’s embedding for the reference image and each of the distorted images. These scalar distances were then passed into a hinge loss function — a standard loss function used for training binary classifiers like support vector machines — treating the human-preferred distortion as positive and the non-preferred distortion as the negative classification. Training a model then boiled down to minimizing this loss. As a starting point for DreamSim, the researchers took three of the most powerful vision foundation models: CLIP, OpenCLIP, and DINO. They then experimented with different fine-tuning strategies on these base models, and found tremendous success with a recent technique called [LoRA](https://arxiv.org/abs/2106.09685), which reduces the number of parameters to be fine tuned, and leverages the existing representational power of the base models. In the end, the team found that the best performing model was an ensemble obtained by concatenating the embedding vectors for fine-tuned versions of CLIP, OpenCLIP, and DINO. The resulting model, called DreamSim, coincides with human judgment more than 96% of the time on perceptual similarity tasks! ## DreamSim in Action To get started working with DreamSim, you can install the Python library on a CUDA-enabled GPU with: ```bash 1pip install dreamsim ``` You can then import the DreamSim model from the library with: ```python 1from dreamsim import dreamsim ``` If you have two images (in PIL format) and you want to compute the distance between them according to DreamSim — the cosine distance between their embeddings — you can do so by running: ```python 1from PIL import Image 2 3model, preprocess = dreamsim(pretrained=True) 4 5img1 = preprocess(Image.open("img1_path")).to("cuda") 6img2 = preprocess(Image.open("img2_path")).to("cuda") 7distance = model(img1, img2) ``` For a performance speedup, you can also use one of the single-branch versions of DreamSim (fine tuned CLIP, OpenCLIP, and DINO). To use the fine tuned DINO branch for instance, you can run the following: ```python 1dreamsim_dino_model, preprocess = dreamsim(pretrained=True, dreamsim_type="dino_vitb16") ``` One of the ancillary benefits of learned similarity metrics like DreamSim over traditional similarity metrics like PSNR and SSIM is that they allow for quick nearest neighbor search. If you want to find the most similar images to a reference image with SSIM, you need to compute the SSIM of the reference image with each candidate image separately. This can become unwieldy as the size of your dataset grows. With a learned similarity metric, however, you can circumvent this scaling problem by computing the embedding for each image once and then using vector search to find approximate nearest neighbors. With vector search plus DreamSim, finding similar images is finally possible and tractable! You can try this out in your browser for yourself, for free, at [try.fiftyone.ai/datasets/nights](https://try.fiftyone.ai/datasets/nights/samples). If you want to work with the NIGHTS dataset locally — including precomputed DreamSim embeddings — you can do so as follows: 1. From there, you can generate a new similarity index from the embeddings of any image model you’d like, and compare reverse image search results with DreamSim. The FiftyOne dataset comes with precomputed similarity indexes for ResNet50 (a pixels and patches model) and CLIP (a semantic model). As an illustrative example of DreamSim’s power, let’s look at how these three models fare at a reverse image search task starting from an image of a bowl of ramen. ResNet50 For a pixels-and-patches models like ResNet50, the images returned as most similar are good textural and color matches with the reference image, but semantically there are some significant differences: pork, steak, and chicken are not that close to ramen conceptually. CLIP On the other extreme, for a semantic similarity model like CLIP, the results returned by the similarity search are conceptually quite close: each image has a bowl with either ramen or rice and vegetables, and most even have a fried egg. However, there are some textural differences: some of the images have a wooden backdrop whereas others have a slate backdrop.DreamSim DreamSim combines the best of both worlds. It isn’t perfect — there are a few images with wooden backdrops in the top 25 results — but DreamSim achieves a much better balance of low and mid-level similarity. 1. Download the raw images by running the script [here](https://try.fiftyone.ai/datasets/nights/samples) 2. Install FiftyOne with `pip install fiftyone` 3. Download the FiftyOne dataset info [here](https://drive.google.com/file/d/1BE4aRIBhdD7AoWabn0Wty0a_BZzTtcnZ/view?usp=sharing) and unzip the folder 4. Load the FiftyOne dataset from this file with `dataset = fo.Dataset.from_dir(“/path/to/downloaded/folder”)` 5. Adjust the image paths on the dataset’s samples ( `sample.filepath`) to match the location of the downloaded images [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [DreamSim](https://voxel51.com/blog/tag/dreamsim) [embeddings](https://voxel51.com/blog/tag/embeddings) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [LPIPS](https://voxel51.com/blog/tag/lpips) [NIGHTS](https://voxel51.com/blog/tag/nights) [open source](https://voxel51.com/blog/tag/open-source) [perceptual metrics](https://voxel51.com/blog/tag/perceptual-metrics) [similarity](https://voxel51.com/blog/tag/similarity) [similarity search](https://voxel51.com/blog/tag/similarity-search) ![](https://cdn.sanity.io/images/h6toihm1/production/d58692baec7c64699806d60d25a0d14f534105fa-300x300.png?auto=format&dpr=2&fit=max&q=75&w=42) Jacob Marks Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/dbcb8b06a25ad72d4e091f48db8502f84be8d6fc-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Build Custom Computer Vision Applications\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 7, 2023](https://voxel51.com/blog/build-custom-computer-vision-applications) [![](https://cdn.sanity.io/images/h6toihm1/production/c1022031003843a1ca2cfca462c25cafc52036d3-1200x675.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How to Cluster Images\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ Apr 10, 2024](https://voxel51.com/blog/how-to-cluster-images) [![](https://cdn.sanity.io/images/h6toihm1/production/aa69c7491887864b0a456481df907919a1394246-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Anomaly Detection with FiftyOne and Anomalib \| Tutorial\\ \\ Computer Vision, Datasets, Plugins\\ \\ • \\ \\ May 6, 2024](https://voxel51.com/blog/anomaly-detection-with-fiftyone-and-anomalib) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-312-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Computer Vision Tips and Tricks – Aug 4, 2023 Aug 4, 2023 • 3 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#aa6fdb0a5fba) [Adding new label fields and mapping them to super categories](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#feb7c2564895) [Choosing a dataset type for instance segmentation](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#07cdd4145ff2) [Reducing the number of images returned](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#1ab9b0ef72c0) [Adding keypoint skeletons based on sample attributes](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#44e44d3de76a) [Merging datasets and renaming labels](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#caf5b1f5af76) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#11fc4366afb8) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#aa6fdb0a5fba) [Adding new label fields and mapping them to super categories](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#feb7c2564895) [Choosing a dataset type for instance segmentation](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#07cdd4145ff2) [Reducing the number of images returned](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#1ab9b0ef72c0) [Adding keypoint skeletons based on sample attributes](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#44e44d3de76a) [Merging datasets and renaming labels](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#caf5b1f5af76) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023#11fc4366afb8) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. As an open source community, the FiftyOne community is open to all. This means everyone is welcome to ask questions, and everyone is welcome to answer them. Continue reading to see the latest questions asked and answers provided! ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## Adding new label fields and mapping them to super categories Community Slack member Sabrina asked: _“I have a COCO dataset with super categories. Is there an easy way to add a new label field which will just remap the original label to the super category? For example, I have annotations labeled by species, but want to remap them all to genus level. I was thinking of duplicating the current label field then `map_labels`, but I can't figure out how to duplicate `ground_truth`. Do I have to iterate through every sample?”_ An option is to use [`clone_sample_field()`](https://docs.voxel51.com/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.clone_sample_field), which clones the given sample field into a new field of the dataset. ## Choosing a dataset type for instance segmentation Community Slack member Namrata asked: _“I'm trying out instance segmentation in FiftyOne and I have a dataset of images in jpg format along with their masks and labels. What dataset type would be appropriate for this? `ImageSegmentationDirectory` seems to be the most appropriate, but there doesn't seem to be a way to upload the labels.”_ Outside of the `FiftyOneDataset` format, `COCODetectionDataset` is a good option. The `ImageSegmentationDirectory` format is for segmentation labels only (with respect to the entire image). It’s also worth checking out the instance segmentation with SAM part in the Jacob Marks' article, [“See What You Segment with SAM”](https://towardsdatascience.com/see-what-you-sam-4eea9ad9a5de) for reference. You can also checkout this [Football Player Segmentation](https://github.com/voxel51/fiftyone-examples/blob/master/examples/football_player_segmentation.ipynb) notebook that steps of loading the dataset using `COCODetectionDataset` and doing instance segmentation. ## Reducing the number of images returned Community Slack member Kornelia asked: _“I'm trying to investigate what is wrong with keypoint annotations for some images. I want to restrict my view of images to those for which I kept annotations. Is that possible? Currently, I see all images regardless of them having annotations. I tried to look for something in the FiftyOne App to hide them, but with no success.”_ You can try something like this: ```python 1ds = foz.load_zoo_dataset("quickstart", max_samples=10) 2ds.match(F("ground_truth.detections").length() > 0) ``` Check out the FiftyOne Docs to learn more about how to [load datasets](http://ofc/) and the available options. ## Adding keypoint skeletons based on sample attributes Community Slack member Kornelia asked: _“Is it possible to add different keypoint skeletons based on whether a picture was taken from a front or rear view?”_ If you're willing to store the keypoints with different skeletons in different label fields, then you can provide separate skeletons for each field like so: ```python 1dataset.skeletons["field1"] = skeleton1 2dataset.skeletons["field2"] = skeleton2 3dataset.save() ``` Check out the FiftyOne Docs to learn more about [storing kepoint skeletons](https://docs.voxel51.com/user_guide/using_datasets.html#storing-keypoint-skeletons). ## Merging datasets and renaming labels Community Slack member Aaditya asked: _“I’d like to merge two datasets and rename the label of one of those. How do I go about that?”_ Depending on your exact use case, two approaches to consider: First, you can use [`merge_samples()`](https://docs.voxel51.com/api/fiftyone.core.dataset.html?highlight=merge_samples#fiftyone.core.dataset.Dataset.merge_samples) to merge the samples from one dataset into another. Second, you could rename a field with [`rename_sample_field()`](https://docs.voxel51.com/api/fiftyone.core.dataset.html?highlight=merge_samples#fiftyone.core.dataset.Dataset.rename_sample_field). ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,900+ [FiftyOne Slack](https://slack.voxel51.com/) members - 3,900+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) (as of today!) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 360+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Instance segmentation](https://voxel51.com/blog/tag/instance-segmentation) [keypoint skeletons](https://voxel51.com/blog/tag/keypoint-skeletons) [Keypoints](https://voxel51.com/blog/tag/keypoints) [labels](https://voxel51.com/blog/tag/labels) [merging data](https://voxel51.com/blog/tag/merging-data) ![](https://cdn.sanity.io/images/h6toihm1/production/b447c3f47d7e0ddcf4272c2034a8431fc05a809f-300x300.png?auto=format&dpr=2&fit=max&q=75&w=42) Jimmy Guerrero Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4ac1a727dc192a21563cde51b6e345f620e09376-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/30c18878a02dd9a2458f82fbedac819f24880527-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks for Adding and Merging Data – Feb 17, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Feb 18, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-for-adding-and-merging-data-feb-17-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-313-lllmstxt|> ## FiftyOne Sample Fields Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) FiftyOne Sample Fields Tips and Tricks – Aug 11, 2023 Aug 11, 2023 • 3 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#b6fc33dfe82a) [Managing your sample fields](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#e46cca17b248) [Adding predictions to your samples](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#42930faf0a3b) [Adding strings or scalars to samples](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#dbe026d6f5c1) [How to map original labels to super categories](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#09b5fb8bb7de) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#961fde7d5f69) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#b6fc33dfe82a) [Managing your sample fields](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#e46cca17b248) [Adding predictions to your samples](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#42930faf0a3b) [Adding strings or scalars to samples](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#dbe026d6f5c1) [How to map original labels to super categories](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#09b5fb8bb7de) [Join the FiftyOne community!](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023#961fde7d5f69) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we recap interesting questions and answers that have recently popped up on [Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), [GitHub](https://github.com/voxel51/fiftyone), Stack Overflow, and Reddit. Recently, many users have been interested in learning more about fields so we will be shining a spotlight on them today! ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! ## **Managing your sample fields** One of the great things about FiftyOne datasets is that your data is more than just an image directory. With FiftyOne, you are able to store metadata, scalar fields, labels, or tags — all within a single sample. Let's start by viewing a sample to see what a default sample may look like. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4dataset = foz.load_zoo_dataset( 5 "coco-2017", 6 split="validation", 7 dataset_name="tips+tricks", 8) 9dataset.persistent = True 10 11sample = dataset.first() 12print(sample) ``` ```raw 1, 13 'ground_truth': 14 ... ``` We can see a several main fields right off the bat: filepath, tags, metadata, and our ground truth labels with all of our detections! Depending on the dataset you loaded, fields can change so it’s great to see how much is stored right off the bat by loading into the FiftyOne. The fun doesn’t stop there as there are plenty of options to add to our samples! ## **Adding predictions to your samples** One of the most useful examples of adding a field to your samples is to add predictions from your model to your samples. This will allow you to not only visualize your predictions next to your ground truths in the Fiftyone App, but also allow you to perform evaluations on your data to get key scores such as accuracy or mAP (mean average precision). A quick way to do this can be seen below or [here](https://docs.voxel51.com/tutorials/evaluate_detections.html#Add-predictions-to-dataset) in our docs! ```python 1from PIL import Image 2from torchvision.transforms import functional as func 3 4import fiftyone as fo 5 6# Get class list 7classes = dataset.default_classes 8 9# Add predictions to samples 10with fo.ProgressBar() as pb: 11 for sample in pb(predictions_view): 12 # Load image 13 image = Image.open(sample.filepath) 14 image = func.to_tensor(image).to(device) 15 c, h, w = image.shape 16 17 # Perform inference 18 preds = model(image) 19 labels = preds["labels"].cpu().detach().numpy() 20 scores = preds["scores"].cpu().detach().numpy() 21 boxes = preds["boxes"].cpu().detach().numpy() 22 23 # Convert detections to FiftyOne format 24 detections = [] 25 for label, score, box in zip(labels, scores, boxes): 26 detections.append( 27 fo.Detection( 28 label=classes[label], 29 bounding_box=box, 30 confidence=score 31 ) 32 ) 33 34 # Save predictions to dataset 35 sample["my_model"] = fo.Detections(detections=detections) 36 sample.save() 37 38session = fo.launch_app(predictions_view) 39 ``` ![](https://cdn.sanity.io/images/h6toihm1/production/ca6c52e8735860cac2680cfc5f3c0018aae1ca06-1600x782.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## **Adding strings or scalars to samples** With FiftyOne samples, you also have the flexibility to add several different basic data type fields to your sample easily. You can use this to keep track of where the data came from, who added it to the dataset, or why it is there. There are many ways to do this but here are two easy ways: 1. Add directly to the sample as we see in our int\_field example. 2. Update all samples with a new field of a specific type, then add the correct entry for that field on each sample. For a full list of basic field types, you can refer [here](https://docs.voxel51.com/user_guide/basics.html#fields) in the docs! ```python 1sample = dataset.first() 2 3## option 1 4sample["int_field"] = 51 5 6## option 2 7dataset.add_sample_field("location", fo.StringField) 8sample["location"] = "outdoor" 9 10sample.save() 11 ``` ## **How to map original labels to super categories** Another cool use case you can achieve easily with FiftyOne is adding something like super categories to your samples. Often users can get bogged down with different label types on a single sample. You can bring clarity to this by holding multiple labels on one sample. One way to tackle this challenge is to clone the sample field using clone\_sample\_field() to duplicate the original ground truths and then map the super categories to the new field. Here is an example! ```python 1import json 2 3with open('annotation.json') as f: 4    data = json.load(f) 5 6# Create a dictionary mapping category names to supercategory names 7class_to_supercategory = {} 8 9for category in data['categories']: 10    class_to_supercategory[category['name']] = category['supercategory'] 11 12# Duplicate ground_truth field 13clone_f = { 14    "ground_truth_detections": "gt_super", 15} 16 17dataset.clone_sample_fields(clone_f) 18 19# Remap the category names to supercategory names in the new field and save 20dataset.map_labels("gt_super", class_to_supercategory).save() ``` To learn more about fields, samples, and more FiftyOne features, head over to our [User Guide](https://docs.voxel51.com/user_guide/index.html) for more information! ## **Join the FiftyOne community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,900+ [FiftyOne Slack](https://slack.voxel51.com/) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 360+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [FAQ](https://voxel51.com/blog/tag/faq) [fields](https://voxel51.com/blog/tag/fields) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [sample fields](https://voxel51.com/blog/tag/sample-fields) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/4d37703d72d4b83a85bda19eb1999d5247915fc0-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Dec 02, 2022\\ \\ Tips & Tricks\\ \\ • \\ \\ Dec 3, 2022](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-dec-02-2022) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0ecb0645c4938217bcade4d3d80cf59f7b05329b-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision View Stages Tips and Tricks – Jan 20, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-314-lllmstxt|> ## Computer Vision Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup — August 10, 2023 Aug 14, 2023 • 5 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#0c9673a39de7) [Neural Congealing: Aligning Images to a Joint Semantic Atlas](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#8eba7996ee03) [Advancing Personalized Medicine and Radiotherapy through AI-Enabled Computer Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#3fb4d6f27eab) [A Practical Approach to Deep Learning for Computer Vision with Tensorflow 2](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#2548a6572469) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#6dc0a4ab13c7) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#4a19ea3bd2b0) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#c093f1ac15cd) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#0c9673a39de7) [Neural Congealing: Aligning Images to a Joint Semantic Atlas](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#8eba7996ee03) [Advancing Personalized Medicine and Radiotherapy through AI-Enabled Computer Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#3fb4d6f27eab) [A Practical Approach to Deep Learning for Computer Vision with Tensorflow 2](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#2548a6572469) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#6dc0a4ab13c7) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#4a19ea3bd2b0) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023#c093f1ac15cd) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We just wrapped up the August 10, 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/), and if you missed it or want to revisit it, here’s a recap! In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide a $200 donation to charitable causes. The charity that received the highest number of votes this month was [Coalition for Rainforest Nations](https://www.rainforestcoalition.org/), an organization on a mission to save the World’s last great rainforests to achieve environmental and social sustainability. We are sending this event’s charitable donation of $200 to Coalition for Rainforest Nations on behalf of the computer vision community! ![](https://cdn.sanity.io/images/h6toihm1/production/80c22f925f693246df7250581ea04e8716aa24f3-1024x129.png?auto=format&dpr=2&fit=max&q=75&w=1024) Missed the Meetup? No problem. Here are playbacks and talk abstracts from the event. ## Neural Congealing: Aligning Images to a Joint Semantic Atlas Presenting [Neural Congealing](https://neural-congealing.github.io/) — a zero-shot self-supervised framework for detecting and jointly aligning semantically-common content across a given set of images. Our approach harnesses the power of pre-trained DINO-ViT features to learn: (i) a joint semantic atlas — a 2D grid that captures the mode of DINO-ViT features in the input set, and (ii) dense mappings from the unified atlas to each of the input images. We derive a new robust self-supervised framework that optimizes the atlas representation and mappings per image set, requiring only a few real-world images as input without any additional input information (e.g., segmentation masks). We design our losses and training paradigm to account only for the shared content under severe variations in appearance, pose, background clutter or other distracting objects, and demonstrate results on a plethora of challenging image sets including sets of mixed domains (e.g., aligning images depicting sculpture and artwork of cats), sets depicting related yet different object categories (e.g., dogs and tigers), or domains for which large-scale training data is scarce (e.g., coffee mugs). [Dolev Ofri-Amar](https://www.linkedin.com/in/dolev-ofri/) is an MSc graduate from the Computer Science and Mathematics department at the Weizmann Institute of Science. Her interests are in computer vision and deep learning, mainly focused on image and video analysis and synthesis. **Resource links** - [Project page](https://neural-congealing.github.io/) - [ArXiv](https://arxiv.org/abs/2302.03956) - [GitHub Repo](https://github.com/dolev104/neural_congealing) ## Advancing Personalized Medicine and Radiotherapy through AI-Enabled Computer Vision In this talk, we will explore the transformative role of computer vision and AI in the realm of medical imaging, focusing on its applications in personalized medicine and radiotherapy. We will delve into cutting-edge research, open source tools, and real-world use cases that demonstrate the potential of AI to enhance diagnostics, treatment planning, and outcome prediction. The presentation will showcase how computer vision techniques coupled with deep learning algorithms are enabling precise and personalized care for patients, revolutionizing the field of healthcare. Speaker: [Roushanak Rahmat, PhD](https://www.linkedin.com/in/roushanakrahmat/) is an accomplished AI scientist with a PhD in AI from Heriot-Watt University. She specializes in computer vision, medical imaging, and data science, developing advanced algorithms and models for healthcare applications. With expertise in deep learning, she has contributed significantly to AI projects in the healthcare sector, particularly in improving radiotherapy treatment. Roushanak is passionate about using AI to revolutionize healthcare and actively shares her knowledge as a public speaker, Women Techmaker ambassador, and through her Medium blog and YouTube channel. ## A Practical Approach to Deep Learning for Computer Vision with Tensorflow 2 A detailed walkthrough of Neuralearn’s [Deep Learning for Computer Vision](https://www.freecodecamp.org/news/how-to-implement-computer-vision-with-deep-learning-and-tensorflow/) course. We shall discuss at a high level how modern deep learning algorithms can be used in solving computer vision tasks using tools like Tensorflow 2, Hugging Face, Onnx, FastAPI, Weights and Biases and Albumentations, going from the basics of Machine Learning to deploying working computer vision solutions. In this course, we lay much emphasis on practice, while explaining the theory behind the different algorithms we use. Learners from different backgrounds can easily follow along since efforts are made to explain every concept as clearly and concisely as possible. Because in recent times, deep learning is usually stereotyped as a math-heavy field, we explain in simple terms every math concept, so that learners who aren’t from a math-related background can start building real-world solutions easily. Given that our focus is mainly on practice, we work on several projects in this course including a car price predictor, a malaria disease classifier, a human emotions detector, an object detector, a digit generator, an image segmenter, a people counter and an image generator. [Folefac Martins](https://www.linkedin.com/in/folefac-martins-8a792617b/) is an MSc graduate from the Electrical and Telecoms Engineering department at the National Advanced School of Engineering, Polytechnique Yaounde. His interests are in deep learning, helping people realize their potential and entrepreneurship. **Resource links** - [Deep Learning for Computer Vision course](https://www.freecodecamp.org/news/how-to-implement-computer-vision-with-deep-learning-and-tensorflow/) on freeCodeCamp.org ## Join the Computer Vision Meetup! Computer Vision Meetup membership has grown to more than [5,000 members](https://www.meetup.com/pro/computer-vision-meetups/) in just one year! The goal of the Meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. Join one of the 13 Meetup locations closest to your timezone. - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) We have exciting speakers already signed up over the next few months! Become a member of the [Computer Vision Meetup closest to you](https://www.meetup.com/pro/computer-vision-meetups/), then register for the Zoom. ## What’s Next? Up next on Aug 24 at 12 PM AEST we have a great line up speakers including: ![](https://cdn.sanity.io/images/h6toihm1/production/e2f43b8d5085f67b9b26af5e21eb5d81a5b9d74d-1024x576.png?auto=format&dpr=2&fit=max&q=75&w=1024) - **Removing Backgrounds Automatically or with a User’s Language –** Jizhizi Li, PhD, University of Sydney - **Self-Supervised Representative Learning for Action Recognition in Videos –** Vidhya Vinay, Co-Founder of Streamingo.ai - **AI at the Edge: Optimizing Deep Learning Models for Real-World Applications –** Raz Petel, SightX Register for the Zoom [here](https://voxel51.com/computer-vision-events/august-24-meetup-apac/). You can find a complete schedule of upcoming Meetups on [the Voxel51 Events page](https://voxel51.com/computer-vision-events/). ## Get Involved! There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping me over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [deep learning](https://voxel51.com/blog/tag/deep-learning) [medical imaging](https://voxel51.com/blog/tag/medical-imaging) [meetup](https://voxel51.com/blog/tag/meetup) [Neural Congealing](https://voxel51.com/blog/tag/neural-congealing) [radiotherapy](https://voxel51.com/blog/tag/radiotherapy) [TensorFlow](https://voxel51.com/blog/tag/tensorflow) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/af656fd506d11b555a019950b688830000b62f30-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Recapping the Vector Search-Themed Computer Vision Meetup — July 13, 2023\\ \\ Event Recaps, Vector Search\\ \\ • \\ \\ Jul 17, 2023](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/a83a8877ea9e07c2c743561ed30acfe8d6748b5a-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — August 24, 2023\\ \\ Event Recaps\\ \\ • \\ \\ Aug 25, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/991b89d515d9f7fe4eea26c14396eb51116b4a8b-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup – April 27, 2023\\ \\ Event Recaps\\ \\ • \\ \\ Apr 28, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-315-lllmstxt|> ## First Week with FiftyOne [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Tutorials](https://voxel51.com/blog/category/tutorials) Spending My First Week With FiftyOne Aug 21, 2023 • 3 min read Article content In this article [Data Curation](https://voxel51.com/blog/spending-my-first-week-with-fiftyone#575e61b1b412) [Model Evaluation in an Instant](https://voxel51.com/blog/spending-my-first-week-with-fiftyone#d6fbcb1c4c8a) [Finding Mistakes in Your Data](https://voxel51.com/blog/spending-my-first-week-with-fiftyone#e749b4a44ba7) [Join the FiftyOne Community!](https://voxel51.com/blog/spending-my-first-week-with-fiftyone#e74b621891aa) In this article [Data Curation](https://voxel51.com/blog/spending-my-first-week-with-fiftyone#575e61b1b412) [Model Evaluation in an Instant](https://voxel51.com/blog/spending-my-first-week-with-fiftyone#d6fbcb1c4c8a) [Finding Mistakes in Your Data](https://voxel51.com/blog/spending-my-first-week-with-fiftyone#e749b4a44ba7) [Join the FiftyOne Community!](https://voxel51.com/blog/spending-my-first-week-with-fiftyone#e74b621891aa) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Forging your own path as a novice ML engineer is a tumultuous journey with numerous obstacles to overcome. Regardless of where you currently stand in this adventure, you're likely all too familiar with the feeling of investing countless hours into configuring data, creating bug-free evaluation scripts, or engaging in a relentless struggle to improve model accuracy by a mere one or two percent. However, after spending a week at [Voxel51](https://voxel51.com/) and taking my first step into the world of [FiftyOne](https://github.com/voxel51/fiftyone), the open source toolkit for building high-quality datasets and computer vision models, I was astonished to discover how many of those hours I could instantly reclaim. In just five minutes, I will demonstrate three easy ways FiftyOne can save you hours, if not days, of work within your computer vision ML workflow. ## **Data Curation** The best way to start your workflow is with the piece of mind that the data you’re working with is formatted correctly and not corrupted. Images sideways or upside down? Labels off by one? How often do ML engineers even look at their dataset before training, outside maybe the first one or two images? Loading data into FiftyOne and having the sanity check that the data is up to standard is a great first step in your ML workflow. With a single snippet of code, FiftyOne can present your data in a clean and intuitive way. FiftyOne leverages powerful data ingestors so you can import datasets in more than 28(!) [different formats](https://docs.voxel51.com/user_guide/dataset_creation/datasets.html#loading-datasets-from-disk) ranging over all the most popular classification, detection, segmentation, and 3D data types. The tool even allows you to [export datasets to any type](https://docs.voxel51.com/user_guide/export_datasets.html#supported-export-formats) for quick conversions from for example COCO to KITTI. With a quick spot check of your ground truths, FiftyOne can save immense amounts of time dealing with potential headaches down the line. ## **Model Evaluation in an Instant** There is no frustration like writing out a whole script only for the model or data type to change, forcing you to have to start all over again. Once again, I was blown away by how FiftyOne simplifies this entire process. After inferencing through your samples in your dataset, with FiftyOne you can easily add predictions to samples and evaluate them. FiftyOne provides a variety of builtin methods for [evaluating model predictions](https://docs.voxel51.com/tutorials/evaluate_detections.html), including regressions, classifications, detections, polygons, instance, and semantic segmentations. No more corralling detections or mass reformatting ground truth labels. Here you can see that evaluating detections, for example, with FiftyOne is as easy as four lines of code to get high quality evaluation results. Results can even be presented in several ways to get full insight into your evaluation. For example, want the mAP of the results? Just ask: Do you want the top 10 classes classification results? Easy: All of this and more can be achieved by adding just a few extra lines to your existing scripts. You can find additional resources on achieving your ideal [model evaluation](https://docs.voxel51.com/user_guide/evaluation.html) workflow in the docs. ## **Finding Mistakes in Your Data** This last highlight is one that every ML engineer can benefit from. When you are fighting for that last one or two percent of accuracy on your model, and you have tweaked every hyperparameter imaginable, it is important not to underestimate the impact of data quality. Sitting on thousands if not millions of samples, sprinkled in there are labeling mistakes or duplicates. By hand they would be impossible to catch. But with FiftyOne’s builtin [Brain](https://docs.voxel51.com/user_guide/brain.html) technology, entire datasets can be inspected for any troublesome samples. For example, to find duplicates, use FiftyOne Brain’s `compute_uniqueness()` method: \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Next if you want to find let's say detection mistakes, use FiftyOne Brain’s `compute_mistakenness()` method: It can not be understated how game changing this is. Any dataset can be transformed to the highest quality easily with no custom scripts. Poor annotations can even be tagged and automatically sent to CVAT (here’s a [FiftyOne+CVAT tutorial](https://docs.voxel51.com/tutorials/cvat_annotation.html) on how to do that) or another annotation tool of your choice in a matter of seconds after finding them. ## Join the FiftyOne Community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 2,000+ [FiftyOne Slack](https://slack.voxel51.com/) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 360+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [CVAT](https://voxel51.com/blog/tag/cvat) [data curation](https://voxel51.com/blog/tag/data-curation) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [model evaluation](https://voxel51.com/blog/tag/model-evaluation) [open source](https://voxel51.com/blog/tag/open-source) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/047b21a97f6c858334f9f35ed89fa7655ebf5767-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ State-of-the-Art Object Detection with YOLO-NAS & FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ May 4, 2023](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/79d00d175a8098516cb2f4a7711131cbf322d01a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Finding and Correcting Mistakes – FiftyOne Tips and Tricks – Aug 18, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Aug 18, 2023](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/dbcb8b06a25ad72d4e091f48db8502f84be8d6fc-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Build Custom Computer Vision Applications\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 7, 2023](https://voxel51.com/blog/build-custom-computer-vision-applications) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-316-lllmstxt|> ## OpenCV AI Competition 2023 [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) The Spirit of Competition – OpenCV AI Competition 2023 Aug 22, 2023 • 4 min read Article content In this article [Cortic Technology’s Winning OpenCV Project](https://voxel51.com/blog/opencv-ai-competition-2023#585c2bb36663) [How Open Source FiftyOne Can Help Enable Your OpenCV Submission](https://voxel51.com/blog/opencv-ai-competition-2023#8a0d6b83adb1) [Join the FiftyOne Community!](https://voxel51.com/blog/opencv-ai-competition-2023#e158f9679fd1) In this article [Cortic Technology’s Winning OpenCV Project](https://voxel51.com/blog/opencv-ai-competition-2023#585c2bb36663) [How Open Source FiftyOne Can Help Enable Your OpenCV Submission](https://voxel51.com/blog/opencv-ai-competition-2023#8a0d6b83adb1) [Join the FiftyOne Community!](https://voxel51.com/blog/opencv-ai-competition-2023#e158f9679fd1) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop As the summer days swiftly pass by, it heralds the commencement of yet another exciting event - the [OpenCV AI Competition](https://www.hackster.io/contests/opencv-ai-competition-2023)! Every year, OpenCV collaborates with industry leaders to present a challenge to machine learning engineers worldwide, encouraging them to showcase their cutting-edge AI models and algorithms. What sets this competition apart is its open-ended nature; participants are free to explore any task or prompt, as long as the solution harnesses the power of the [OpenCV Library](https://opencv.org/), easily one of the best computer vision toolkits available today. From robotics to agriculture, education to health and medical, and even sports, submissions are diverse and innovative. The competition culminates on November 30th, granting ample time for participants to conceive extraordinary projects such as a laser-guided weed killer or an autonomous driving RC helicopter - the sky's the limit! Over the years, OpenCV has played a pivotal role in inspiring teams to step up and innovate to address real problems within their communities. The spirit of competition has a unique way of uniting people for a greater cause, driving them to exceed their ordinary limits. Events like this foster an atmosphere of growth, propelling the field of computer vision forward and sparking brilliant new ideas. As the deadline approaches, submissions will pour in from across the globe, showcasing the potential for groundbreaking breakthroughs. Some winners may spin off into new startups or, in the case of team DeViTech, who built a laser-based optical 3D scanner, will finally be able to balance their bookcase. ![](https://cdn.sanity.io/images/h6toihm1/production/84624eda9f72725d4b82e05594803bf62a551021-797x621.png?auto=format&dpr=2&fit=max&q=75&w=797) One such winner of the OpenCV competition, Cortic Technology, took this problem first hand by developing an [AI platform for kids](https://www.cortic.ca/) to incorporate computer vision with Legos! In this post, I’ll share what I learned from talking to Ye Lu from Cortic, and how open source FiftyOne could help you streamline your OpenCV project! ## Cortic Technology’s Winning OpenCV Project ![](https://cdn.sanity.io/images/h6toihm1/production/fec7b23838b23cb94088f2ba8618011646ca544c-1203x528.png?auto=format&dpr=2&fit=max&q=75&w=1203) Cortic Technology, the triumphant team of the 2021 OpenCV Spatial AI competition, created a block programming language through a user-friendly interface, empowering young students to build their first computer vision applications. Their visionary founder and team lead, Ye Lu, aimed to minimize barriers for beginners in computer vision, enabling them to create incredible applications without prior expertise. When asked he explained that kids are hands-on learners and need to be immersed in the world of CV where they can see up close the results of their work. Cortic leverages things like Legos, Raspberry Pis, as well as mobile robots to construct a world where kids can be enchanted by modern AI developments. Also mentioned the need for powerful visualization tools, being able to see the detections, classifications, or other results in an interactive space can be extremely helpful for teaching kids. The OpenCV AI Competition stands as a beacon of innovation, rallying talented minds from across the globe to push the boundaries of AI. It serves as a testament to the power of community-driven initiatives in fostering education and inspiring the next generation of AI enthusiasts. With the continued efforts of organizations like OpenCV and the remarkable endeavors of past winners like Cortic Technology, we are taking significant strides towards creating a world where AI literacy and responsible usage are commonplace. ## How Open Source FiftyOne Can Help Enable Your OpenCV Submission Want to turbocharge your OpenCV submission? Use open source [FiftyOne](https://github.com/voxel51/fiftyone) to get the most out of your data! With the ability to store group data, different modalities, and present in intuitive and insightful ways, FiftyOne is a force multiplier to your project. Let’s go back to the autonomous driving RC helicopter idea from the opening paragraph. What if we were going to hover around our house to track migrating butterflies? Luckily with FiftyOne, data has never been easier to curate and visualize! We can pop a butterfly dataset into FiftyOne and voila! ![](https://cdn.sanity.io/images/h6toihm1/production/d3a2388ca3753f678d0098ccde502d096ae60128-1911x994.png?auto=format&dpr=2&fit=max&q=75&w=1600) FiftyOne can empower you to extract stunning insights from your dataset. You can [add detections](https://docs.voxel51.com/tutorials/evaluate_detections.html), [run evaluations](https://docs.voxel51.com/user_guide/evaluation.html), [compute embeddings](https://docs.voxel51.com/tutorials/image_embeddings.html), and much much more! Once your data has been uploaded to FiftyOne, a world of options open up for you. You can do [uniqueness sorting](https://docs.voxel51.com/tutorials/uniqueness.html) to find your most unique butterflies in your dataset like below. ![](https://cdn.sanity.io/images/h6toihm1/production/2e48837e2fa9ea579826cf2d7181abaf427d21dc-1815x963.png?auto=format&dpr=2&fit=max&q=75&w=1600) If you are a beginner to not just FiftyOne but all of computer vision, you can do simple operations as well such as finding your favorite butterfly type. To find more instructions on how you can gain control over your data, please head over to all of our tutorials [here](https://docs.voxel51.com/tutorials/evaluate_classifications.html)! Want help on how to unlock FiftyOne’s potential for your OpenCV AI solution? Join the FiftyOne Community Slack and find others taking initiative to tackle some of the world's most interesting Computer Vision problems. Best of luck to all entries to the competition! ![](https://cdn.sanity.io/images/h6toihm1/production/abad3ea7f3410ff83d15ccab154f7154384deea5-1324x710.gif?auto=format&dpr=2&fit=max&q=75&w=1324) ## Join the FiftyOne Community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 2,000+ [FiftyOne Slack](https://slack.voxel51.com/) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 370+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Competition](https://voxel51.com/blog/tag/competition) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [OpenCV](https://voxel51.com/blog/tag/opencv) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/b5ec2410f8c8844aa682fea044da0a14e3d9c5c7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ VoxelGPT: Your AI Assistant for Computer Vision\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jun 7, 2023](https://voxel51.com/blog/voxelgpt-your-ai-assistant-for-computer-vision) [![](https://cdn.sanity.io/images/h6toihm1/production/5068fe2d5a454e641a9ad3cc910a9b84d1dec0a6-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ NeurIPS 2023 and the State of AI Research\\ \\ Computer Vision, Event Recaps, Product & News\\ \\ • \\ \\ Dec 8, 2023](https://voxel51.com/blog/neurips-2023-and-the-state-of-ai-research) [![](https://cdn.sanity.io/images/h6toihm1/production/241965669b62a1c1b47c1ca0c5a91156608da907-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ NeurIPS 2023 Survival Guide\\ \\ Computer Vision, Event Recaps, Product & News\\ \\ • \\ \\ Dec 9, 2023](https://voxel51.com/blog/neurips-2023-survival-guide) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-317-lllmstxt|> ## FiftyOne Tips and Tricks [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) Finding and Correcting Mistakes – FiftyOne Tips and Tricks – Aug 18, 2023 Aug 18, 2023 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023#6f1944aca64c) [Finding and Removing Duplicate Images](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023#a20df70de11d) [Finding Classification Mistakes](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023#f73f6283c337) [Finding and Correcting Detection Mistakes](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023#cf7d2915361a) [Join the FiftyOne Community!](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023#836f81c1ca3c) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023#6f1944aca64c) [Finding and Removing Duplicate Images](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023#a20df70de11d) [Finding Classification Mistakes](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023#f73f6283c337) [Finding and Correcting Detection Mistakes](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023#cf7d2915361a) [Join the FiftyOne Community!](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023#836f81c1ca3c) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we cover interesting workflows and features of FiftyOne! ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! Also feel free to follow along in our [notebook](https://github.com/voxel51/fiftyone-examples/blob/mistakes-add/examples/Finding%20and%20Correcting%20Mistakes.ipynb) or on [YouTube](https://www.youtube.com/watch?v=WDl80g7_SBw)! ## Finding and Removing Duplicate Images Typical image datasets can contain upwards of tens of thousands if not millions of images. It is not uncommon to find duplicate images hidden amongst the masses. These duplicated images can harm the training of any model on this data, and can have serious consequences if not corrected. Leveraging FiftyOne, we can use built in functionality to find these duplicate images and remove them from our dataset, courtesy of the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html)! ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4# Load the CIFAR-10 test split 5# Downloads the dataset from the web if necessary 6dataset = foz.load_zoo_dataset("cifar10", split="test") 7 8session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/a70c92d151efabb9f708d75179cc605b0011ef2d-872x499.png?auto=format&dpr=2&fit=max&q=75&w=872) We are able to load in our data and take a look at it in the FiftyOne App. At first glance, the data looks normal. But with 10,000 images to look through, finding a duplicate can quickly turn into an all day affair. Luckily with the FiftyOne Brain, we are able to find duplicates in an instant with `compute_uniqueness()`! ```python 1import fiftyone.brain as fob 2 3fob.compute_uniqueness(dataset) 4 5# Sort in increasing order of uniqueness (least unique first) 6dups_view = dataset.sort_by("uniqueness") 7 8# Open view in the App 9session.view = dups_view ``` ![](https://cdn.sanity.io/images/h6toihm1/production/372c5b0f37e9a3b8d7fa3ffcdd197d3248c19ca8-1697x1091.png?auto=format&dpr=2&fit=max&q=75&w=1600) Here we are able to see duplicate images in CIFAR10! Next, we can click on each of these images individually and select them in the top left corner of the box. Once all your duplicate images are selected, they can be tagged with the following code. Make sure to leave one original image! ```python 1# Get currently selected images from App 2dup_ids = session.selected 3 4# Mark as duplicates 5dups_view = dataset.select(dup_ids) 6dups_view.tag_samples("dups") 7 8# Visualize duplicates-only in App 9session.view = dups_view ``` ![](https://cdn.sanity.io/images/h6toihm1/production/a92a7366fc91c4f31c0883f77d0add2ff4abadb8-1841x963.png?auto=format&dpr=2&fit=max&q=75&w=1600) After executing the code above, we are able to tell that duplicates have been properly tagged. Once you are confident that there are no more duplicates within your dataset, you can create a clean view that is ready for training. You can even export this view as a new and improved version of your dataset to use for future use. ```python 1from fiftyone import ViewField as F 2 3clean_view = dataset.sort_by("uniqueness").match_tags("dups", bool=False) 4 5export_dir = "/path/for/image-classification-dir-tree" 6 7label_field = "ground_truth"  # for example 8 9# Export the dataset 10clean_view.export( 11 export_dir=export_dir, 12 dataset_type=fo.types.ImageClassificationDirectoryTree, 13 label_field=label_field, 14) ``` ## Finding Classification Mistakes Another prevalent form of annotation mistakes is a classification mistake on the label of a sample. It can happen if the picture of your dog is labeled cat or vice versa and can really impact the learning capabilities of your models. FiftyOne has built in functionality using the FiftyOne Brain to catch these mistakes and help you fix them. In the following example, we will be taking a look at CIFAR10 again and purposely corrupt our dataset with incorrect labels. For a quick way to start this example, follow along in the linked notebook or check out the [docs](https://docs.voxel51.com/tutorials/classification_mistakes.html). Once you have corrupted your labels and had a trained model add its predictions to your dataset, we can begin. Let’s kick it off by using the FiftyOne Brain `compute_mistakenness()`! ```python 1import fiftyone.brain as fob 2 3# Get samples for which we added predictions 4h_view = dataset.match_tags("processed") 5 6# Compute mistakenness 7fob.compute_mistakenness(h_view, model_name, label_field="ground_truth", use_logits=True) 8 9# Sort by likelihood of mistake (most likely first) 10mistake_view = (dataset 11 .match_tags("processed") 12 .sort_by("mistakenness", reverse=True) 13) 14 15# Show only the samples for which we added label mistakes 16session.view = mistake_view ``` \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop After running our Brain function, we are able to discover all the mislabeled images we have in our dataset. With the mistakes having been found, we can tag and remove as we did previously. Alternatively, the files can also be removed from their filepaths or sent back to [annotation](https://docs.voxel51.com/tutorials/cvat_annotation.html) after being tagged as well. ## Finding and Correcting Detection Mistakes Detection mistakes can be especially difficult to find by hand. The problem can rapidly expand from checking 1000 images to 10,000 labels to make sure all the boxes are perfect. Having misplaced boxes or duplicate boxes is another heavy detriment to the training process of detection models. FiftyOne can help step in and take a load off for finding these mistakes in your dataset. Take a look below or in the [docs](https://docs.voxel51.com/tutorials/detection_mistakes.html) to see how: ```python 1dataset = foz.load_zoo_dataset("coco-2017", split="validation", max_samples=1000, overwrite=True, dataset_name="Find Mistakes") 2 3import fiftyone.utils.iou as foui 4 5#Calculate the overlaps of boxes within your dataset 6foui.compute_max_ious(dataset, "ground_truth", iou_attr="max_iou", classwise=True) 7 8print("Max IoU range: (%f, %f)" % dataset.bounds("ground_truth.detections.max_iou")) 9 10# Retrieve detections that overlap above a chosen threshold 11dups_view = dataset.filter_labels("ground_truth", F("max_iou") > 0.75) 12 13session.view = dups_view ``` ![](https://cdn.sanity.io/images/h6toihm1/production/50b31b3ef7228ebe74cd0f3f6a42ef0975419a23-874x721.png?auto=format&dpr=2&fit=max&q=75&w=874) In a few lines we were able to take 1000 samples and whittle down into 7 potential candidates for mistakes. Some are just two very close boxes like our first image of two baseball players. Some are truly mistakes as is the case with the man at the beach. To fix the mistake, we can open up the sample and tag the incorrect bounding box as a duplicate. ![](https://cdn.sanity.io/images/h6toihm1/production/b76977246106783a499307527aadb60901cbb606-874x721.png?auto=format&dpr=2&fit=max&q=75&w=874) After the label has been tagged as a duplicate it can be removed from the sample entirely, fixing the mistake in the sample with `dataset.delete_labels(tags="dups")`. Alternatively, you could send the tagged images for reannotation with one of FiftyOne’s native annotation integrations. One such integration with FiftyOne is [CVAT](https://docs.voxel51.com/tutorials/cvat_annotation.html), and can be accomplished quickly: ```python 1anno_key = "remove_dups" 2 3dups_view.annotate(anno_key, label_field="ground_truth", launch_editor=True) ``` Hopefully, these tips will help you find and correct these mistakes in your data to allow for you to create the best models you can! Good Luck! To learn more about fields, samples, and more FiftyOne features, head over to our [User Guide](https://docs.voxel51.com/user_guide/index.html) for more information! ## Join the FiftyOne Community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,900+ [FiftyOne Slack](https://slack.voxel51.com/) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 360+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [annotation](https://voxel51.com/blog/tag/annotation) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [CVAT](https://voxel51.com/blog/tag/cvat) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Labeling mistakes](https://voxel51.com/blog/tag/labeling-mistakes) [open source](https://voxel51.com/blog/tag/open-source) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/4ac1a727dc192a21563cde51b6e345f620e09376-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0fc50e89593dec2117ce5c1934761cf61b24f8c4-1920x1080.jpg?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Tips and Tricks for Accelerating Computer Vision Workflows – Mar 17, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Mar 18, 2023](https://voxel51.com/blog/fiftyone-tips-and-tricks-for-accelerating-computer-vision-workflows-mar-17-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/a3a918e30b0553723b9392ea90763379f98480a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – April 7, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Apr 7, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-318-lllmstxt|> ## Celebrating FiftyOne's Milestones [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Celebrating Three Years of FiftyOne! Aug 19, 2023 • 3 min read Article content In this article [FiftyOne 0.21.5 and 0.21.6 are here!](https://voxel51.com/blog/celebrating-three-years-of-fiftyone#e2e123b2a752) [Open source FiftyOne crossed 4000 stars on GitHub!](https://voxel51.com/blog/celebrating-three-years-of-fiftyone#7eca3396af9a) [FiftyOne turns three!](https://voxel51.com/blog/celebrating-three-years-of-fiftyone#66cab88adbf3) [Join the FiftyOne community!](https://voxel51.com/blog/celebrating-three-years-of-fiftyone#65ae73f4a60f) In this article [FiftyOne 0.21.5 and 0.21.6 are here!](https://voxel51.com/blog/celebrating-three-years-of-fiftyone#e2e123b2a752) [Open source FiftyOne crossed 4000 stars on GitHub!](https://voxel51.com/blog/celebrating-three-years-of-fiftyone#7eca3396af9a) [FiftyOne turns three!](https://voxel51.com/blog/celebrating-three-years-of-fiftyone#66cab88adbf3) [Join the FiftyOne community!](https://voxel51.com/blog/celebrating-three-years-of-fiftyone#65ae73f4a60f) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Last week the team at [Voxel51](https://voxel51.com/) got together to meet, greet, and work towards achieving our mission of bringing transparency and clarity to the world’s data. While we were away … [open source FiftyOne](https://github.com/voxel51/fiftyone) achieved three amazing milestones. In this post, we share and celebrate those milestones! ![](https://cdn.sanity.io/images/h6toihm1/production/ab610e6afc84e65ff078c3bd3a3c19fff6607543-3024x1888.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## FiftyOne 0.21.5 and 0.21.6 are here! Last week we announced FiftyOne 0.21.5 and 0.21.6 What's new? Here are some highlights: ### Models - [Segment Anything](https://docs.voxel51.com/user_guide/model_zoo/models.html#segment-anything-vith-torch) is now available in the [FiftyOne Model Zoo](https://docs.voxel51.com/user_guide/model_zoo/index.html#model-zoo)! - [DINOv2](https://docs.voxel51.com/user_guide/model_zoo/models.html#dinov2-vitl14) is now available in the [FiftyOne Model Zoo](https://docs.voxel51.com/user_guide/model_zoo/index.html#model-zoo)! - You can now [load any model from PyTorch Hub](https://docs.voxel51.com/integrations/pytorch_hub.html#pytorch-hub) and directly apply it to your FiftyOne datasets! ### Other goodies - Added support for controlling [field visibility in the grid independent of filtering](https://github.com/voxel51/fiftyone/pull/3248) - Added support for filtering by label tags in [individual label fields](https://github.com/voxel51/fiftyone/pull/3287) - Upgraded the [Labelbox integration](https://docs.voxel51.com/integrations/labelbox.html#labelbox-integration) to support the latest Labelbox API version - Added support for [gRPC connections](https://docs.voxel51.com/integrations/qdrant.html#qdrant-setup) when using the Qdrant similarity backend ### Bug fixes - Improved robustness when updating datasets in multiple processes concurrently - Improved handling of [group datasets](https://docs.voxel51.com/user_guide/groups.html) whose groups may contain missing samples for certain slices - Resolved bugs with [similarity queries](https://docs.voxel51.com/user_guide/brain.html#similarity) using the sklearn backend - Fixed text and checkbox attribute usage when using our [CVAT 2.5 integration](https://docs.voxel51.com/integrations/cvat.html) ### Community contributions Special thanks to these awesome community members for contributing to this release! [Rusteam](https://github.com/Rusteam) added Segment Anything to the model zoo [#3019](https://github.com/voxel51/fiftyone/pull/3019) [timmermansjoy](https://github.com/timmermansjoy) added support for using MPS devices when running Torch models on macOS [#2843](https://github.com/voxel51/fiftyone/pull/2843) [smidm](https://github.com/smidm) fixed a bug when exporting keypoints with NaN coordinates in COCO format [#3316](https://github.com/voxel51/fiftyone/pull/3316) [Sa-Schmi](https://github.com/Sa-Schmi) fixed a bug with custom Visualizers in the App [#3357](https://github.com/voxel51/fiftyone/pull/3357) [NeoKish](https://github.com/NeoKish) squashed a number of documentation bugs [#3283](https://github.com/voxel51/fiftyone/pull/3283), [#3289](https://github.com/voxel51/fiftyone/pull/3289), [#3290](https://github.com/voxel51/fiftyone/pull/3290) [glenn-jocher](https://github.com/glenn-jocher) updated the YOLOv5 exporter to support Ultralytics' latest dataset format [#3393](https://github.com/voxel51/fiftyone/pull/3393) [mys007](https://github.com/mys007) added bazel support for the App [#3338](https://github.com/voxel51/fiftyone/pull/3338) [helioshe4](https://github.com/helioshe4) updated the model zoo to officially support torchvision>=0.15.0 [#3348](https://github.com/voxel51/fiftyone/pull/3348) Check out [the release notes](https://docs.voxel51.com/release-notes.html) for a full rundown of the new features! ## Open source FiftyOne crossed 4000 stars on GitHub! If you’re like me, you star repos as a way to show support for open source projects you love, or to bookmark a repo you visit frequently or want to dive into later. For us at Voxel51, stars are one way we feel the love that the capabilities we’re building in the open source project are meaningful to members of the community. Last week we crossed 4000 stars! Thank you to everyone who’s been a part of the amazing journey with us so far, and we look forward to continuing to grow and many more milestones ahead! \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## FiftyOne turns three! Three years ago, we [launched](https://voxel51.com/press/fiftyone-open-source-launch/) open source FiftyOne, the world’s first (and today’s most prolific!) open source tool for building high-quality datasets and computer vision models. Not yet familiar? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s been an awesome ride so far and we wanted to take a quick walk back down memory lane and celebrate key accomplishments in the FiftyOne community. ### A quick trip down memory lane: key dates **October 18, 2018**: Started Voxel51 Inc. to enable developers, scientists, and organizations to build high-quality datasets and computer vision models **August 8, 2019**: Announced Seed Funding **June 1, 2020**: Released FiftyOne 0.1 to a few dozen private-beta users **August 11, 2020**: Open sourced FiftyOne, making it the open source tool for building high-quality datasets and computer vision models **Late July, 2021**: Began working with dozens of startups and Fortune 500 enterprises as early adopters of FiftyOne Teams **September 21, 2022**: Announced Series A funding and the public availability of FiftyOne Teams ... And **today (in honor of August 11, 2023 last week)**: We celebrate 3 years of open source FiftyOne! \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 1,900+ [FiftyOne Slack](https://slack.voxel51.com/) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 360+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne Community](https://voxel51.com/blog/tag/fiftyone-community) [open source](https://voxel51.com/blog/tag/open-source) [Voxel51 milestone](https://voxel51.com/blog/tag/voxel51-milestone) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0f4cab7e19991301a33a2691bbfeb2d9a22024fc-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Four Years of Open Source FiftyOne!\\ \\ Product & News\\ \\ • \\ \\ Aug 12, 2024](https://voxel51.com/blog/four-years-of-open-source-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/b2e0ca21c16fec18362fd532541417ccd9969a8c-1400x1127.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Turns One!\\ \\ Product & News\\ \\ • \\ \\ Aug 12, 2021](https://voxel51.com/blog/fiftyone-turns-one) [![](https://cdn.sanity.io/images/h6toihm1/production/29fa790b13637c569efda3fa1a797042c04d6bad-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – Feb ‘23\\ \\ Product & News\\ \\ • \\ \\ Feb 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-feb-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-319-lllmstxt|> ## Create Your AI Art Gallery [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Plugins](https://voxel51.com/blog/category/plugins), [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks), [Tutorials](https://voxel51.com/blog/category/tutorials) Build Your Own AI Art Gallery Aug 25, 2023 • 5 min read Article content In this article [Add Stable Diffusion and DALL-E2 Images Directly to Your Dataset](https://voxel51.com/blog/build-your-own-ai-art-gallery#b03e0aaa5dad) [AI Art Gallery 🤖🎨🖼️](https://voxel51.com/blog/build-your-own-ai-art-gallery#8ec6a0a279c6) [Plugin Overview & Functionality](https://voxel51.com/blog/build-your-own-ai-art-gallery#cc8cbf4d0adc) [Installing the Plugin](https://voxel51.com/blog/build-your-own-ai-art-gallery#203815cba3d7) [Lessons Learned](https://voxel51.com/blog/build-your-own-ai-art-gallery#767f366a13b3) [Conclusion](https://voxel51.com/blog/build-your-own-ai-art-gallery#a045f8ab6abb) In this article [Add Stable Diffusion and DALL-E2 Images Directly to Your Dataset](https://voxel51.com/blog/build-your-own-ai-art-gallery#b03e0aaa5dad) [AI Art Gallery 🤖🎨🖼️](https://voxel51.com/blog/build-your-own-ai-art-gallery#8ec6a0a279c6) [Plugin Overview & Functionality](https://voxel51.com/blog/build-your-own-ai-art-gallery#cc8cbf4d0adc) [Installing the Plugin](https://voxel51.com/blog/build-your-own-ai-art-gallery#203815cba3d7) [Lessons Learned](https://voxel51.com/blog/build-your-own-ai-art-gallery#767f366a13b3) [Conclusion](https://voxel51.com/blog/build-your-own-ai-art-gallery#a045f8ab6abb) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) _**Editor's note**: The plugin featured in this post enables you to add images to your dataset directly from text prompts, and it has recently been significantly updated, with a ton of new image generation models added! Check out the [GitHub repo](https://github.com/jacobmarks/text-to-image) to see all the models this plugin now supports._ ## Add Stable Diffusion and DALL-E2 Images Directly to Your Dataset Welcome to week one of _Ten Weeks of Plugins_. For the next ten weeks, we will be building a FiftyOne plugin (or multiple!) each week and sharing the lessons learned! If you’re new to them, FiftyOne Plugins provide a flexible mechanism for anyone to extend the functionality of their FiftyOne App. You may find the following resources helpful: - [FiftyOne Plugins Repo](https://github.com/voxel51/fiftyone-plugins) - [FiftyOne Plugin Docs](https://docs.voxel51.com/plugins/index.html#downloading-plugins) - Plugins Channel in the [FiftyOne Community Slack](https://slack.voxel51.com/) Ok, let’s dive into this week’s FiftyOne Plugin! ## AI Art Gallery 🤖🎨🖼️ ![](https://cdn.sanity.io/images/h6toihm1/production/2112efbb160cc5810782a3c6ff6fb7f153fad3e6-1600x848.gif?auto=format&dpr=2&fit=max&q=75&w=1600) Have you ever generated stunning imagery with DALL-E2 or Stable Diffusion, only to ask yourself: “How do I catalog this??”. Do you find yourself manually downloading AI generated images and moving them into the desired folders? How do you handle metadata like which model and scheduler was used to generate a specific image? Consider those problems solved! ## Plugin Overview & Functionality For the first week of _10 Weeks of Plugins_, I built an [AI Art Gallery plugin](https://github.com/jacobmarks/ai-art-gallery). This plugin allows you to select your text-to-image model, generate images from text prompts within the FiftyOne App, and add the image directly to your dataset — your “AI art gallery” — in one foul swoop. Out of the box, the plugin supports three models: - [DALL-E2](https://openai.com/dall-e-2) - [Stable Diffusion](https://replicate.com/stability-ai/stable-diffusion) - [Feed-forward VQGAN-CLIP](https://replicate.com/mehdidc/feed_forward_vqgan_clip) Depending on which text-to-image model you choose, you can configure everything from the size of the image to be generated, all the way to the number of inference steps. After generating an image, the sample grid in the FiftyOne App will refresh and the latest image will appear in the bottom right of the grid. You can then filter your art gallery by text prompt, model name, date and time generated, or any other attributes used in the generation process. ### Adding Creator’s Notes Want to add a note to a specific piece of artwork describing your inspiration? How about jotting down observations about how the wording of your prompt affects the image generation process? Add a tag! ![](https://cdn.sanity.io/images/h6toihm1/production/d23b3add2453b902a4a9c29e30a35b17c76c89f9-2560x1354.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ### Cleaning House When you’re making art, you may not love every single piece of artwork. Part of curation is deciding what makes the cut and what doesn’t. You can delete a piece of artwork by selecting the image (the checkbox in the upper left corner of the image), pressing the backtick (" `` ` ``”) to pull up your list of [operators](https://docs.voxel51.com/plugins/index.html#operators) (functions which execute Python code from the UI), and choosing `delete_selected_samples`. Hit `Execute` and you’re done! ![](https://cdn.sanity.io/images/h6toihm1/production/4d96c976c2b954d6b98a95d8d33da4d7d7bfd352-1600x849.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## Installing the Plugin If you haven’t already done so, install FiftyOne: ```bash 1pip install fiftyone ``` Then you can download this plugin from the command line with: ```bash 1fiftyone plugins download https://github.com/jacobmarks/ai-art-gallery ``` Refresh the FiftyOne App, and you should see the `txt2img` operator in your operators list when you press the “ `` ` ``” key. To keep the plugin’s core code simple, this plugin uses text-to-image models that are accessible via API endpoint, with [OpenAI](https://openai.com/) (DALL-E2) and [Replicate](https://replicate.com/) (Stable Diffusion and VQGAN). As such, you will need to have an account with at least one of these services in order to use the plugin out of the box. Make sure that you have your API info in environment variables: ```bash 1export OPENAI_API_KEY=... 2export REPLICATE_API_TOKEN=... ``` You do not need both to use the plugin — the operator checks your environment variables and only shows as options models accessible via the corresponding APIs. If you want to use a different text-to-image model, local or via API, it should be easy to extend this code by writing a `Text2Image` subclass for the model you are interested in. ## Lessons Learned I built the AI Art Gallery plugin as a Python Plugin, so I didn’t have to worry about writing Typescript/React code. The plugin consisted of three files: - `__init__.py`: defining the operator - `fiftyone.yml`: making the plugin _register_ for download and installation - `README.md`: describing the plugin ### Creating a Responsive Input Form I managed to build the plugin with a single operator: `txt2img`, and decided to lean heavily on the operator’s _inputs_, which are [described in](https://github.com/jacobmarks/ai-art-gallery/blob/a5858fc6b8ee1aeacc5a035d99a8a67a807f1d88/__init__.py#L177) the `resolve_input()` method (this is the bulk of the code!). Not every FiftyOne Plugin will be so input heavy. As you can see in the GIF at the top of the blog, the options displayed in the input form change as I change the selected model. I was able to do this as follows: - [Define an input object](https://github.com/jacobmarks/ai-art-gallery/blob/a5858fc6b8ee1aeacc5a035d99a8a67a807f1d88/__init__.py#L188), `radio_choices`, as a collection of radio buttons — an enumeration of discrete choices. The user can then select a single one of these. - [Give the user’s choice a “name”](https://github.com/jacobmarks/ai-art-gallery/blob/a5858fc6b8ee1aeacc5a035d99a8a67a807f1d88/__init__.py#L194C1-L194C1), in this case `model_choices`. - Retrieve the value selected by the user from the `params` dictionary in the operator’s context: `ctx.params.get("model_choices", False)` and perform different blocks of logic depending on the value. Another key to building a responsive input form was making extensive use of the FiftyOne [Python Operator API docs](https://docs.voxel51.com/api/fiftyone.operators.types.html#module-fiftyone.operators.types). There are a ton of operator types to choose from, and within this input form I used `Dropdown`, `RadioGroup`, and `SliderView`. ### Custom Component Properties While I only scratched the surface of component customization in this plugin, I learned that as a plugin creator, you have immense control over the look and feel of your plugin components, even when just working with Python operators! When building the slider to set the number of inference steps, one of the front end engineers at Voxel51, [Ibrahim](https://www.linkedin.com/in/ibrahimmanjra/), let me in on a little secret: you can pass in a dictionary of `componentProps` as an argument to any view. This essentially puts the power of [Material UI Components](https://mui.com/material-ui/getting-started/) in your hands. I stayed pretty basic in this plugin, using `componentProps` to specify the minimum and maximum allowed values for the slider, as well as a step size: ```python 1inference_steps_slider = types.SliderView( 2 label="Num Inference Steps", 3 componentsProps={"slider": {"min": 1, "max": 500, "step": 1}}, 4) ``` But moving forward, I definitely want to explore this more deeply! ### Plugins and Environment Variables While it isn’t shown in the GIF, what appears in the input form will be different depending on what packages you have installed and API connections you have set up. The key enabler here is that plugins have access to your environment variables. To make this plugin work in a variety of conditions, I did the following: 1. [Use the FiftyOne utils](https://github.com/jacobmarks/ai-art-gallery/blob/main/__init__.py#L19C7-L19C7) `lazy_import()` method so that a package is only imported when it is needed. Otherwise, it could cause an import error unnecessarily. 2. Use importlib’s `find_spec` [to check](https://github.com/jacobmarks/ai-art-gallery/blob/a5858fc6b8ee1aeacc5a035d99a8a67a807f1d88/__init__.py#L55C1-L65C78) if a package matching a certain specification has been installed. 3. [Check](https://github.com/jacobmarks/ai-art-gallery/blob/a5858fc6b8ee1aeacc5a035d99a8a67a807f1d88/__init__.py#L59C6-L59C6) whether a certain variable name is in the environment variables with `os.environ`. 4. Execute different logic blocks based on these conditions. ## Conclusion Whether you are building your own AI art gallery, or using text-to-image models to generate a synthetic dataset, this plugin will shorten the feedback cycle and empower you to curate your visual data better. But this plugin is only the beginning. With FiftyOne Plugins, the sky’s the limit on how you can extend the FiftyOne computer vision toolkit to meet the needs of your data and model workflows. Stay tuned over the next ten weeks while we pump out a killer lineup of plugins! You can track our journey in our [ten-weeks-of-plugins repo](https://github.com/jacobmarks/ten-weeks-of-plugins) — and I encourage you to fork the repo and join me on this journey! [AI art](https://voxel51.com/blog/tag/ai-art) [DALLE2](https://voxel51.com/blog/tag/dalle2) [genAI](https://voxel51.com/blog/tag/genai) [generative AI](https://voxel51.com/blog/tag/generative-ai) [OpenAI](https://voxel51.com/blog/tag/openai) [plugins](https://voxel51.com/blog/tag/plugins) [replicate](https://voxel51.com/blog/tag/replicate) [Stable Diffusion](https://voxel51.com/blog/tag/stable-diffusion) [text-to-image](https://voxel51.com/blog/tag/text-to-image) [vqgan](https://voxel51.com/blog/tag/vqgan) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/947ffda95306b167a53fa0c1dc0f05ec32b71873-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Ask Your Images Anything\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 1, 2023](https://voxel51.com/blog/ask-your-images-anything) [![](https://cdn.sanity.io/images/h6toihm1/production/fedd0c008df994c1834839ea35bc65b34e47fb5d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Double Trouble: Eliminate Image Duplicates with FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 14, 2023](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/bad75ba72dfae8cdefc3d9afe33a1ea9a9c4ec36-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Optical Character Recognition with PyTesseract\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 21, 2023](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-320-lllmstxt|> ## Computer Vision in Healthcare [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Industry Solutions](https://voxel51.com/blog/category/industry-solutions), [Product & News](https://voxel51.com/blog/category/product-news) How Computer Vision Is Changing Healthcare Aug 31, 2023 • 21 min read Article content In this article [Healthcare industry overview](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#18cb5c4b8092) [Applications of computer vision in healthcare](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#04cc149405e3) [Companies at the cutting edge of computer vision in healthcare](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#3b381c1a13e3) [Healthcare industry datasets and challenges](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#2b01590684dc) [Healthcare industry models and frameworks](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#a94b88c0808f) [Wait, What’s FiftyOne?](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#63746af1ce8d) [Join the FiftyOne community!](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#8a543034aefa) In this article [Healthcare industry overview](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#18cb5c4b8092) [Applications of computer vision in healthcare](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#04cc149405e3) [Companies at the cutting edge of computer vision in healthcare](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#3b381c1a13e3) [Healthcare industry datasets and challenges](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#2b01590684dc) [Healthcare industry models and frameworks](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#a94b88c0808f) [Wait, What’s FiftyOne?](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#63746af1ce8d) [Join the FiftyOne community!](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare#8a543034aefa) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the third installment of [Voxel51](https://voxel51.com/)’s computer vision industry spotlight blog series. In this series, we highlight how different industries — from construction to climate tech, from retail to robotics, and more — are using computer vision, machine learning, and artificial intelligence to drive innovation. We’ll dive deep into the main computer vision tasks being put to use, current and future challenges, and companies at the forefront. In this edition, we’ll focus on healthcare! Read on to learn about computer vision in healthcare and medicine. ## Healthcare industry overview Key facts and figures: - Globally, the healthcare market is [projected to eclipse $11 trillion by 2025](https://www.supportivecareaba.com/statistics/healthcare-industry). - Around the world, there are [approximately 59 million healthcare workers](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5299814/#:~:text=There%20are%20approximately%2059%20million%20healthcare%20workers%20worldwide.). - Roughly [14% of all workers in the United States](https://www.census.gov/library/stories/2021/04/who-are-our-health-care-workers.html#:~:text=There%20were%2022%20million%20workers,American%20Community%20Survey%20(ACS).) (totaling 22 million) are employed in healthcare. - In 2021, healthcare spending [accounted for 18.3%](https://www.cms.gov/Research-Statistics-Data-and-Systems/Statistics-Trends-and-Reports/NationalHealthExpendData/NationalHealthAccountsHistorical) of the USA’s GDP. - According to Grand View Research, in 2022 the global AI in healthcare market was [valued at $15.4 billion](https://www.grandviewresearch.com/industry-analysis/artificial-intelligence-ai-healthcare-market). This market is expected to grow 37.5% per year until 2030. Before we dive into several popular applications of computer vision-based AI technologies in healthcare, here are some of the industry’s key challenges. Key industry challenges in healthcare: - **Rising costs of healthcare**: in real terms, in the United States healthcare costs have [risen by 290% since 1980](https://www.brookings.edu/research/a-dozen-facts-about-the-economics-of-the-u-s-health-care-system/). Costs are so high that in a recent [Kaiser Family Foundation](https://www.kff.org/polling/) poll, 43% of respondents said that a family member had put off or postponed necessary health care as a result. - **Shortage of physicians**: despite the millions of workers employed in the healthcare industry, 132 countries are experiencing shortages of physicians, with an estimated [12.8 million more doctors needed](https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(22)00532-3/fulltext) worldwide to alleviate this problem. The United States is expected to face a [shortage of up to 124,000 physicians by 2034](https://www.ama-assn.org/practice-management/sustainability/doctor-shortages-are-here-and-they-ll-get-worse-if-we-don-t-act). - **Time-intensive EHR**: according to a study involving 155,000 US physicians, on average, physicians spend [more than 16 minutes per patient ecounter](https://pubmed.ncbi.nlm.nih.gov/31931523/) using electronic health records (EHR). This time was split between chart review, documentation, and ordering. Continue reading for some ways computer vision in healthcare is enabling people to live longer, healthier lives. ## Applications of computer vision in healthcare ### Computer-aided detection and diagnosis _Computer-aided detection applied to mammograms. Image courtesy of [Siemens Healthineers](https://www.siemens-healthineers.com/en-iq/mammography/news/history-of-computer-aided-detection.html)._ In healthcare, computer aided detection (CADe) and computer-aided diagnosis (CADx) refer to any applications in which a computer assists a doctor or in understanding and evaluating medical data. Within the context of computer vision, this typically means images coming from MRI, CT, X-ray, or another form of diagnostic imaging technique. However, CAD can even be (and has been) [applied to camera images of faces](https://pubmed.ncbi.nlm.nih.gov/28991753/). Computers can assist doctors in detection by identifying suspicious regions in images and alerting the physician to these concerning areas. In some cases, CADe amounts to predicting high confidence regions of interest and projecting bounding boxes for these regions onto the original image. In other cases, region of interest identification is followed by instance segmentation. In still other cases, anomaly detection may be applied to catch abnormalities. Computer-aided diagnosis takes images and other patient record information as input, and outputs evaluates the likelihood that the patient has a given disease or condition. In cancer diagnosis, for instance, this often means predicting whether a tumor is benign or malignant. Over the past few years, computer vision systems integrating detection and diagnosis have [achieved remarkable performance](https://www.sciencedirect.com/science/article/abs/pii/S1386505618302880). Nevertheless, there are still [challenges to be overcome](https://pubmed.ncbi.nlm.nih.gov/31445285/), from high false positive detection rates to adoption of these technologies into existing systems and workflows. Here are some papers on using computer vision for detection and diagnosis: - [Computer aided detection (CAD): an overview](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC1665219/) - [Computer-Aided Diagnosis in the Era of Deep Learning](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7293164/) - [Computer aids and human second reading as interventions in screening mammography](https://www.sciencedirect.com/science/article/abs/pii/S0959804908001287) - [Computer-aided detection (CADe) and diagnosis (CADx) system for lung cancer with likelihood of malignancy](https://biomedical-engineering-online.biomedcentral.com/articles/10.1186/s12938-015-0120-7) - [Computer-Aided Diagnosis with Deep Learning Architecture](https://www.nature.com/articles/srep24454) And here are a few papers showing the promise transformer models hold for aiding medical professionals in diagnostic processes: - [Task Transformer Network for Joint MRI Reconstruction and Super-Resolution](https://arxiv.org/abs/2106.06742) - [Unsupervised MRI Reconstruction via Zero-Shot Learned Adversarial Transformers](https://arxiv.org/abs/2105.08059) - [TED-net: Convolution-free T2T Vision Transformer-based Encoder-decoder Dilation network for Low-dose CT Denoising](https://arxiv.org/abs/2106.04650) - [Eformer: Edge Enhancement based Transformer for Medical Image Denoising](https://arxiv.org/abs/2109.08044) ### Monitoring disease progression _Computer vision helps doctors monitor the progression of diseases and detect onset earlier. Image courtesy of Accuray on Unsplash._ In addition to detecting abnormalities and diagnosing diseases, computer vision can be used to precisely monitor the progression of a disease over time. Just as a doctor measures patient weight, height, and blood pressure during an exam, and compares these to the patient’s past measurements to assess patient health, computer vision models allow doctors to track various diseases by comparing markers in biomedical images taken at different points in time. In some applications, progression can be monitored by precisely tracking the size of objects like cavities or lesions. More generally, computer vision applications in disease progression monitoring are characterized by a deep learning model assigning a numerical score to an image or video. These scores allow physicians to quantify the severity and time of onset for a specific patient. By precisely monitoring progression, computer vision models can help physicians to detect the onset of a disease more rapidly. For glaucoma, [a 2018 study concluded](https://www.sciencedirect.com/science/article/abs/pii/S000293941830271X) that deep learning accelerated time to detection by more than a year. While computer vision applications in disease monitoring primarily involve applying machine learning models to images, there are also applications involving video data. [In a 2022 paper](https://www.sciencedirect.com/science/article/pii/S2666521221000223), researchers found that by evaluating the movement patterns of Parkinson’s disease patients in the act of rising from a chair, the researchers could robustly estimate disease severity. To interpret patient movement, the researchers estimated patient pose in each frame, and then quantified the velocity and smoothness of motion across video frames. Some papers to get you started: - [Computer-vision based method for quantifying rising from chair in Parkinson's disease patients](https://www.sciencedirect.com/science/article/pii/S2666521221000223) - [Predicting conversion to wet age-related macular degeneration using deep learning](https://www.nature.com/articles/s41591-020-0867-7) - [Detection of Longitudinal Visual Field Progression in Glaucoma Using Machine Learning](https://www.sciencedirect.com/science/article/abs/pii/S000293941830271X) - [Automated Diagnosis of Plus Disease in Retinopathy of Prematurity Using Deep Convolutional Neural Networks](https://jamanetwork.com/journals/jamaophthalmology/article-abstract/2680579) ### Preoperative surgical planning [Preoperative surgical planning](https://en.wikipedia.org/wiki/Surgical_planning) (often abbreviated to surgical planning) encompasses all visualization, simulation, and construction of blueprints to be used during a subsequent surgical procedure. Surgical planning is most often used for neurosurgery, or oral or cosmetic surgery, but it can be applied in advance of any surgery. Studies have found that digital templating and other preoperative planning techniques can [reduce costs in the operating room](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9091925/) and [reduce operation time](https://pubmed.ncbi.nlm.nih.gov/30786290/), and for elderly populations, it is even [associated with lower 90-day mortality](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2769079) rates. Computer vision is deeply entrenched in the history of preoperative planning: [as early as the 1970’s](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5144463/), CT scans and other forms of diagnostic imaging started to make possible the construction of anatomical models. This imaging data is used to generate 2D projections called templates, 3D digital reconstructions, or even 3D printed models. The models then allow surgeons to explore various approaches and entry pathways prior to surgery, minimizing time of surgery and invasiveness. In recent years, the application of emerging technologies like deep learning and virtual reality in preoperative planning has begun to streamline surgery even further. According to [a recent study involving 193 knee surgery cases](https://www.materialise.com/en/inspiration/articles/artificial-intelligence-knee-surgey-planning), AI-based preoperative planning has the potential to reduce the average number of intraoperative corrections a surgeon needs to make by up to 50%. Among the advances contributing to this reduction, [precision location of anatomical landmarks](https://josr-online.biomedcentral.com/articles/10.1186/s13018-021-02294-9) (keypoint detection) leads to improved accuracy in selecting the appropriate prosthetic size, and deep learning models can be employed to [identify and evaluate potential trajectories](https://thejns.org/focus/view/journals/neurosurg-focus/52/4/article-pE10.xml?tab_body=fulltext) for surgical instruments. Here are a few papers on AI in preoperative planning: - [Value of 3D preoperative planning for primary total hip arthroplasty based on artificial intelligence technology](https://josr-online.biomedcentral.com/articles/10.1186/s13018-021-02294-9) - [A novel surgical planning system using an AI model to optimize planning of pedicle screw trajectories with highest bone mineral density and strongest pull-out force](https://thejns.org/focus/view/journals/neurosurg-focus/52/4/article-pE10.xml?tab_body=fulltext) - [Recent advances in surgical planning & navigation for tumor biopsy and resection](https://qims.amegroups.com/article/view/8005/html) ### Intraoperative surgical guidance _Computer vision can help surgeons perform procedures more precisely. Image courtesy of the National Cancer Institute._ Artificial intelligence and computer vision are also being used inside the operating room to make surgeries smoother and less invasive. One unifying theme for intraoperative intelligence is computer-assisted navigation. These applications take inspiration from traditional robotics, employing [depth estimation](https://paperswithcode.com/task/depth-estimation) and [simultaneous localization and mapping](https://en.wikipedia.org/wiki/Simultaneous_localization_and_mapping) to endoscopy images. In [image guided surgery](https://www.neurosurgery.columbia.edu/patient-care/treatments/image-guided-surgery), a surgeon uses this real time visual and spatial information to navigate with enhanced precision. This technology is at the heart of [minimally invasive surgical](https://www.yalemedicine.org/conditions/minimally-invasive-surgery) (MIS) procedures. Intraoperative images can also be [fused with preoperative images in augmented reality environments](https://karger.com/vis/article/36/6/456/310538/Computer-Vision-in-the-Surgical-Operating-Room). While still far from ubiquitous, we are also beginning to see computer vision-enabled navigation, AI-driven decision making, and robotic maneuverability come together in partially and even fully autonomous robotic surgical systems. Last year, researchers at Johns Hopkins University created the [Smart Tissue Autonomous Robot](https://www.science.org/doi/10.1126/scirobotics.abj2908) (STAR), which automated 83% of the task of suturing the small bowel by combining tissue motion tracking and anatomical landmark tracking with a [motorized suturing tool](https://www.corporis-medical.com/product-page/reusable-laparoscopic-suturing-device-374mm). Here are some papers on applications of computer vision during surgery: - [Computer Vision in the Surgical Operating Room](https://www.karger.com/Article/FullText/511934) - [Autonomous robotic laparoscopic surgery for intestinal anastomosis](https://www.science.org/doi/10.1126/scirobotics.abj2908) For an overview of AI applications in surgery, including preoperative and intraoperative applications, check out these papers: - [Artificial Intelligence in Surgery: Promises and Perils](https://dspace.mit.edu/bitstream/handle/1721.1/129454/nihms951789.pdf?sequence=2&isAllowed=y) - [Application of artificial intelligence in surgery](https://journal.hep.com.cn/fmd/EN/article/downloadArticleFile.do?attachType=PDF&id=27515) ### Assisting people with vision loss _Computer vision can help the visually impaired to navigate, providing an alternative to (or working in conjunction with) guide dogs. Image courtesy of Unsplash._ In many applications of computer vision, AI models are employed to take on tasks in order to free up humans from tedious or hazardous work. Perhaps nowhere is computer vision’s liberating effect on people more pronounced than in assisting people with vision loss. By helping people with blindness or limited vision map their surroundings and navigate indoor and outdoor environments, computer vision makes it easier for them to live on their own and attend work and school. Computer vision techniques for assisting the visually impaired span the gamut from object recognition and [face recognition](https://en.wikipedia.org/wiki/Facial_recognition_system), to [monetary denomination verification](https://arxiv.org/pdf/2204.03738.pdf) and [navigation](https://paperswithcode.com/task/visual-navigation). A few papers to pique your interest: - [Efficient Multi-Object Detection and Smart Navigation Using Artificial Intelligence for Visually Impaired People](https://www.mdpi.com/1099-4300/22/9/941) - [Computer Vision-based Assistance System for the Visually Impaired Using Mobile Edge Artificial Intelligence](https://openaccess.thecvf.com/content/CVPR2021W/MAI/papers/Mahendran_Computer_Vision-Based_Assistance_System_for_the_Visually_Impaired_Using_Mobile_CVPRW_2021_paper.pdf) - [Implementation and Analysis of AI-Based Gesticulation Control for Impaired People](https://www.hindawi.com/journals/wcmc/2022/4656939/) ### Other AI in healthcare These highlighted applications only scratch the surface of how artificial intelligence is transforming healthcare and the delivery of medical care in 2023. Computer vision is also being used to [automate cell counting](https://www.nature.com/articles/s41598-021-01929-5), [make operating rooms smarter](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5320916/), [facilitate medial skill training](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7367541/), and [aid in post-traumatic rehabilitation](https://link.springer.com/article/10.1007/s00530-021-00815-4). Beyond computer vision, artificial intelligence is revolutionizing the way we process, understand, and draw insights from biomedical data. In 2019, following the release of Google’s Bidirectional Encoder Representations (BERT) large language model, a team of researchers from Korea University and Clova AI published [_BioBERT: a pre-trained biomedical language representation model for biomedical text mining_](https://arxiv.org/abs/1901.08746). In the intervening years, large language models like BioBERT, including [Med-BERT](https://www.nature.com/articles/s41746-021-00455-y) and [BEHRT](https://www.nature.com/articles/s41598-020-62922-y) (for Electronic Health Records) and [Clinical BERT](https://github.com/EmilyAlsentzer/clinicalBERT) have been used for [hypothesis generation and knowledge discovery](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1204/reports/custom/report29.pdf#page=9&zoom=100,144,270), [diagnosis](https://aclanthology.org/2021.eacl-main.75.pdf), and [hospital readmission prediction](https://arxiv.org/pdf/1904.03323.pdf). Most recently, [Medical ChatGPT](https://github.com/lucidrains/medical-chatgpt), Microsoft’s [BioGPT](https://github.com/microsoft/BioGPT), and [BioMedLM](https://github.com/stanford-crfm/BioMedLM) have shown that generative pretrained transformer (GPT) models - the same architecture powering ChatGPT - can show impressive results on biomedical tasks. Expect this trend to accelerate in the coming months. ### A double dose of caution While biases inherent in artificially intelligent models can have unforeseen and often unwanted consequences across all industries, in healthcare, if AI models are not deployed with careful consideration, they can exacerbate existing inequities and do more harm than good. For a thorough discussion of these issues, check out the following resources: - [Sources of bias in artificial intelligence that perpetuate healthcare disparities—A global review](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9931338/) - [Addressing fairness in artificial intelligence for medical imaging](https://www.nature.com/articles/s41467-022-32186-3) ## Companies at the cutting edge of computer vision in healthcare ### .lumen _The .lumen glasses were built to empower the blind._ Founded in 2020, [.lumen](https://www.dotlumen.com/) is on a mission to help the world’s visually impaired — a population that is expected to reach 100 million by 2050. Started by [Cornel Amariei](https://ro.linkedin.com/in/cornelamariei), and staffed with more than 50 engineers, scientists, and technologists, Romania’s first deep tech startup is building AI powered “glasses” that give the visually impaired enhanced mobility. This additional freedom could potentially enable millions of people to study, pursue careers, and live unassisted. Working with 600+ blind individuals, .lumen designed the glasses to replicate the main features of having a guide dog like helping you avoid obstacles, reacting to commands, and even gently _pulling_ you in the right direction. The difference: guide dogs can cost [up to $50,000 to train](https://www.guidingeyes.org/guide-dogs-101/), making them inaccessible to most blind individuals. The company’s glasses consist of six separate cameras, which together record almost 10GB of data per minute. Leveraging information from these onboard cameras, the .lumen glasses interpret the surrounding environment via computer vision tasks including object detection, image classification, semantic segmentation, and optical character recognition, all with centimeter-level precision. The glasses are also equipped with speech-to-text and text-to-speech capabilities. While some tasks like reading are enhanced by internet connectivity, all safety-related inference happens on the edge. Visual information is imparted to the person wearing the glasses via sound and touch. To make this happen, the company uses advanced AI and robotics technology to integrate visual, auditory, and haptic systems. All told, the glasses have a sub-100ms latency, In 2021, .lumen received €9.3M in funding from the European Union Innovation Council, and won the [Red Dot: Luminary](https://www.red-dot.org/design-concept/red-dot-luminary) award for the design of their headset. In 2022, the Romanian startup was honored as a [Deep Tech Pioneer](https://hello-tomorrow.org/deep-tech-pioneers/). ### Future Fertility _Image of oocytes. Image courtesy of [Atlas of Human Embryology](https://atlas.eshre.eu/es/14555410245416920)._ In 2021, [2.4% of all births in the United States were conceived via assistive reproductive technologies](https://www.cofertility.com/family-learn/fertility-statistics) (ARTs) like in vitro fertilization (IVF). However, IVF can be quite expensive, with the cost of a [single cycle reaching up to $30,000](https://www.forbes.com/health/family/how-much-does-ivf-cost/). What’s more, 67% of the time, a mother’s first IVF cycle does not result in pregnancy. Canadian biotech startup [Future Fertility](https://futurefertility.com/en/) is using AI and computer vision to usher in the next era in personalized fertility medicine to help doctors, embryologists, and patients make more informed decisions along the fertility journey. It all starts with the [_oocyte_](https://en.wikipedia.org/wiki/Oocyte). Oocytes are female egg cells that are retrieved from the ovaries as part of egg freezing and IVF treatments. Through the IVF process, the oocyte is later fertilized to produce an embryo. As such, oocyte quality is an essential factor in determining the outcome of an IVF cycle. Nevertheless, even the trained eye of an embryologist has difficulty assessing the quality of an oocyte from visual features. This contrasts with embryos, where manual assessment has been possible and reliable using a robust grading system as the standard of care for years. To bridge this gap, Future Fertility has developed a patented, non-invasive software solution that provides personalized oocyte quality assessments by analyzing single 2D oocyte images captured within fertility labs _._ The software can generate reports for both egg freezing patients and IVF patients that assess each oocyte’s likelihood of forming a blastocyst (a day 5 or 6 embryo). Not only does this bring much-needed standardization and objectivity to the oocyte assessment process, but the software also predicts blastocyst outcomes over 20% more accurately than embryologists. The resulting reports enable fertility specialists to better manage their patients’ expectations for pregnancy success and help guide their treatment decisions for future cycles by understanding their patient’s personal oocyte quality versus age-based population health statistics. One of the main challenges Future Fertility overcame when developing their model was constructing a high quality dataset. Their solution based on Deep Neural Networks was trained and tested using 100,000+ oocyte images spanning different countries, outcomes, lab equipment, and varied image quality. To learn more about how Future Fertility framed the problem and built a solution using machine learning, check out [this insightful blog post by their team](https://futurefertility.com/en/resources/decoding-good-ai-in-reproductive-medicine-building-and-assessing-high-quality-machine-learning-models-for-clinical-practice/). In early 2022, the company raised a $6M series A round of funding. ### Hologic _Hologic’s [3DQuorum Technology](https://www.hologic.com/hologic-products/breast-health-solutions/3dquorum-imaging-technology) combines multiple slices into SmartSlices with clinically relevant information._ A Nasdaq-traded corporation focused on women’s health, [Hologic](https://www.hologic.com/) is paving the way in AI-enabled breast cancer detection. With almost 7,000 employees, thousands of patents, and a presence in 36 countries, Hologic has a hand in countless screening, diagnostic, and laboratory technologies, including a suite of 2D and 3D medical imaging solutions. Hologic’s [3DQuorum](https://www.hologic.com/hologic-products/breast-health-solutions/3dquorum-imaging-technology) and [Genius AI Detection](https://www.hologic.com/sites/default/files/2020_12/WP-00178_Rev02_GeniusAI_Detection-white-paper-6979r10p.pdf) technologies employ deep learning to make breast cancer detection faster, safer, and more accurate. Traditionally, two-dimensional X-rays known as mammograms were an integral part of breast cancer screens. Recently, a 3D imaging technique called tomosynthesis has gained traction, with the added dimensionality of images allowing for improved sensitivity _and_ specificity. Conventional tomosynthesis uses slices that are 1 mm in thickness, so in order to review all image slices for a screening a radiologist must inspect 240 images. Hologic’s 3DQuorum Imaging technology analyzes groups of 1-mm slices using computer-aided detection algorithms, and synthesizes new images for each group by combining clinically relevant regions. The resulting images are synthetic 6-mm slices called SmartSlices, which allow radiologists to [save an hour per day](https://www.hologic.com/hologic-products/breast-health-solutions/3dquorum-imaging-technology#:~:text=Expedited%20Clarity%20HD%20High%20Resolution,image%20interpretation%20time%20per%20day.) without reduction in performance. Hologic’s computer-aided detection system, Genius AI Detection, was trained on large quantities of clinical data to locate lesions in breast tomosynthesis images. It uses a Faster-RCNN model to identify potentially relevant regions, U-Net models for segmentation, and Inception models for classification. Genius AI Detection has separate modules for regions of interest containing soft tissue lesions and those containing calcification clusters. For each identified lesion, a score is assigned based on the model’s confidence that the lesion is malignant. In clinical studies, radiologists assisted by Genius AI Detection exhibited improved sensitivity, and achieved higher [area under curve](https://towardsdatascience.com/understanding-auc-roc-curve-68b2303cc9c5) (AUC) than those without. ### Iterative Health _Iterative Health’s [SKOUT](https://iterative.health/products/skout/) polyp detection in action. Image Courtesy of Iterative Health._ Founded in 2017 after CEO [Jonathan Ng](https://www.linkedin.com/in/jonathan-ng-55784b61/)’s visit to Cambodia, [Iterative Health](https://iterative.health/) was built to bring precision medicine to all, regardless of location or socio-economic position. The series B startup, which has raised more than $193M from investors including Insight Ventures and Obvious Ventures, is focused on applying AI and computer vision to gastroenterology (GI). One of their flagship products, [SKOUT](https://iterative.health/products/skout/), applies object detection techniques to images streaming from colonoscopic video feeds in real time to detect polyps, adenomas. With SKOUT, physicians are able to [identify 27% more adenomas per colonoscopy](https://www.gastrojournal.org/article/S0016-5085(22)00519-4/fulltext). In some computer vision applications, small objects can be the most challenging to identify. For Computer aided detection of polyps, however, when devices detect increased numbers of polyps, these polyps are for the most part diminutive (<5mm), and as a result, less likely to indicate health concerns. SKOUT demonstrates enhanced detection capabilities for both small and large polyps. Additionally, SKOUT is designed to be integrated into physicians’ existing workflows. According to a [study published in the journal Gastroenterology](https://www.gastrojournal.org/article/S0016-5085(22)00519-4/fulltext), colonoscopies employing SKOUT do not take significantly longer. In late 2022, Iterative Health [received FDA clearance](https://www.businesswire.com/news/home/20220922005730/en/Iterative-Scopes-Receives-FDA-Clearance-for-AI-Assisted-Polyp-Detection-Device-SKOUT%E2%84%A2) for their artificial intelligence based polyp detection device. ### Pixee Medical _Pixee Medical’s [Knee+](https://www.pixee-medical.com/en/knee/) solution augments the surgeon’s view using computer vision and augmented reality. Image courtesy of Pixee Medical._ Founded in 2017, French medical device manufacturer [Pixee Medical](https://www.pixee-medical.com/en/) is using computer vision, artificial intelligence, and augmented and virtual reality to transform orthopedic surgery. Pixee Medical’s augmented reality (AR) surgical glasses help surgeons perform operations with high precision while minimizing invasiveness. At the heart of Pixee’s technology is a suite of state-of-the-art tracking algorithms, which leverage data from a monocular camera embedded in the glasses to locate objects down to less than 1 mm in three dimensional space, even when surgical equipment partially occludes the tracked objects. These location and tracking algorithms make it possible for surgical teams to identify the positions of anatomical landmarks (keypoint markers on the human body) without invasive probes or costly imaging. In 2021, the company launched [Knee+](https://www.pixee-medical.com/en/knee/), a knee arthroplasty solution that helps surgeons navigate during an operation using the AR glasses. Surgical instruments are tagged with QR codes, which allow for precise localization in 3D. The glasses integrate this information to calculate alignment angles for prosthesis, and then project this information in the form of a hologram onto the scene. In March, 2023, the company was crowned a member of the inaugural [French Tech Health20](https://lafrenchtech.com/fr/la-france-aide-les-startup/french-tech-health20/) program. ### SafelyYou Image courtesy of [Discovery Village senior living](https://www.discoveryvillages.com/senior-living-blog/understanding-the-dangers-of-falling-when-you-age/). San Francisco-based [SafelyYou](https://www.safely-you.com/) is developing artificial intelligence and computer vision-enabled solutions to create safer environments for the [55+ million individuals worldwide who live with dementia](https://www.who.int/news-room/fact-sheets/detail/dementia#:~:text=Key%20facts,nearly%2010%20million%20new%20cases.). The company was started in 2016 by CEO [George Netscher](https://www.linkedin.com/in/george-netscher-a4696b59/), spun out of his doctoral research at the [Berkeley AI Research](https://bair.berkeley.edu/) (BAIR) lab, and [motivated by experience with Alzheimers within his family](https://www.safely-you.com/about-us/). Since then, the company has raised more than $61 million to use AI to detect and prevent falls. SafelyYou operates within assisted living communities, where they apply models to every frame from ongoing camera feeds to detect whether or not someone is on the ground. Their models are tuned for high recall, and their world-leading technology is supported with insights from clinical experts who partner with on-site staff, determining the best interventions to prevent future falls. SafelyYou’s fall detection model detects over 99% of on-the-ground events, and when a fall happens, their systems alert caregivers in real-time. To achieve this level of accuracy, they address challenges including detection in cluttered environments, edge cases where people are not on the ground but are in a vulnerable position, and the overarching reality that falls are “long tail” events: there is way more footage of people not on the ground than there is of people on the ground. To date, they have detected more than 100,000 falls, and by getting residents help immediately and improving outcomes, [SafelyYou _doubles_ the average length of stay](https://info.safely-you.com/length-of-stay-carlton-whitepaper-ds). Building on the resounding success of their fall detection technology, in January of 2023, SafelyYou announced the launch of [SafelyYou Aware](https://www.safely-you.com/news/safelyyou-launches-safelyyou-aware-to-support-and-transform-clinical-care-at-senior-living-communities-with-remote-hourly-nighttime-wellness-checks/), which provides hourly video assessments from SafelyYou team members every night to assess risks to a resident's safety. ### Other companies making strides: - [Aidence](https://www.aidence.com/): recently acquired by [RadNet](https://www.radnet.com/), Aidence delivers AI radiology solutions for detection of lung cancer nodules and other disorders. - [Depuy Synthes](https://www.jnjmedtech.com/en-US/companies/depuy-synthes): subsidiary of Johnson & Johnson. [VELYS Hip Navigation](https://www.jnjmedtech.com/en-US/products/digital-surgery/velys-hip-navigation) platform to improve surgical outcomes. - [Intuitive Surgical](https://www.intuitive.com/en-us): famous for their [da Vinci Surgical System](https://en.wikipedia.org/wiki/Da_Vinci_Surgical_System), Intuitive Surgical uses advanced robotics, computer vision, and AI to assist in surgical procedures, and even enable remote operations. - [Qynapse](https://qynapse.com/): French neuroimaging startup applying segmentation to white matter lesions in MRI scans. ## Healthcare industry datasets and challenges The [National Institutes of Health](https://www.nih.gov/) plays a central role in creating and maintaining publicly available health-related datasets, including machine learning and computer vision datasets such as [DeepLesion](https://paperswithcode.com/dataset/deeplesion), [OASIS](https://www.cms.gov/Medicare/Quality-Initiatives-Patient-Assessment-Instruments/HomeHealthQualityInits/DataSpecifications), and [ChestX-ray8](https://openaccess.thecvf.com/content_cvpr_2017/papers/Wang_ChestX-ray8_Hospital-Scale_Chest_CVPR_2017_paper.pdf). Nevertheless, they are far from the only players in town. Here are some of the most popular public datasets and challenges at the intersection of computer vision, machine learning, and healthcare: ### Cancer - The [International Skin Imaging Collaboration](https://www.isic-archive.com/#!/topWithHeader/wideContentTop/main) (ISIC) has multiple datasets and challenges, including the skin cancer [ISIC Kaggle Challenge](https://www.kaggle.com/datasets/nodoubttome/skin-cancer9-classesisic). - The [Cancer Imaging Archive](https://www.cancerimagingarchive.net/) has [125+ collections](https://www.cancerimagingarchive.net/collections/) spanning various types of cancer and imaging modalities. ### Chest X-Ray (CXR) - [CheXpert](https://stanfordmlgroup.github.io/competitions/chexpert/): dataset from the Stanford ML Group containing 224,316 chest radiographs (frontal and lateral views), as well as an associated competition. - [ChestX-ray14](https://nihcc.app.box.com/v/ChestXray-NIHCC): dataset from the NIH Clinical Center consisting of 112,120 frontal-view X-ray images from 30,805 unique patients. The dataset is also [on Kaggle](https://www.kaggle.com/datasets/nih-chest-xrays/data). - [PadChest](https://bimcv.cipf.es/bimcv-projects/padchest/): dataset of 160,000+ high resolution chest X-rays, along with associated patient and case metadata. Reports contain 174 distinct radiographic label types. - [MIMIC-CXR-JPG](https://mimic.mit.edu/docs/iv/modules/cxr/): chest radiographs with structured labels, from MIT’s [Medical Informatics Mart for Intensive Care](https://mimic.mit.edu/). ### MRI - [IVDM3Seg](https://ivdm3seg.weebly.com/data.html): a collection of 24 multi-modal MRI datasets for intervertebral disc (IVD) localization and segmentation. - [MRNet](https://stanfordmlgroup.github.io/competitions/mrnet/): Knee MRI dataset from the Stanford ML Group comprising 1,370 knee MRI exams. ### Other great sources - [MedPix](https://medpix.nlm.nih.gov/home): From the [National Library of Medicine](https://www.nlm.nih.gov/), MedPix contains integrated textual and image data from 12,000 patient cases. This data is searchable by modality, topic, keyword and more. - [Grand Challenge](https://grand-challenge.org/): a platform for algorithms, challenges, and end-to-end ML solutions in biomedical imaging applications. If you would like to see any of these, or other medical or health related computer vision datasets added to the [FiftyOne Dataset Zoo](https://voxel51.com/docs/fiftyone/user_guide/dataset_zoo/index.html), get in touch and we can work together to make this happen! ## Healthcare industry models and frameworks If you’ve made it this far, then you may be interested in: - [Med-PaLM](https://cloud.google.com/blog/topics/healthcare-life-sciences/sharing-google-med-palm-2-medical-large-language-model): A multimodal large language model from Google that can synthesize information from medical images like X-rays and mammograms. - [MedSAM](https://github.com/bowang-lab/MedSAM): a [Segment Anything Model](https://github.com/facebookresearch/segment-anything) fine-tuned for medical image data, released on Apr 24, 2023 - [MONAI](https://monai.io/): The Medical Open Network for Artificial Intelligence - a set of open-source frameworks related to medical imaging ## Wait, What’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It supports everything from point clouds to DICOM! ## Join the FiftyOne community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 2,000+ [FiftyOne Slack](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 370+ repositories - 65+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [AI](https://voxel51.com/blog/tag/ai) [artificial intelligence](https://voxel51.com/blog/tag/artificial-intelligence) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [computer-aided diagnosis](https://voxel51.com/blog/tag/computer-aided-diagnosis) [disease detection](https://voxel51.com/blog/tag/disease-detection) [healthcare](https://voxel51.com/blog/tag/healthcare) [indiustry spotlight](https://voxel51.com/blog/tag/indiustry-spotlight) [machine learning](https://voxel51.com/blog/tag/machine-learning) [medical imaging](https://voxel51.com/blog/tag/medical-imaging) [medicine](https://voxel51.com/blog/tag/medicine) [segment anything](https://voxel51.com/blog/tag/segment-anything) [surgical guidance](https://voxel51.com/blog/tag/surgical-guidance) ![](https://cdn.sanity.io/images/h6toihm1/production/d58692baec7c64699806d60d25a0d14f534105fa-300x300.png?auto=format&dpr=2&fit=max&q=75&w=42) Jacob Marks Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/ebd118c3d8181dcba0532200d024759e23c7ec1e-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Computer vision in healthcare: 12 breakthrough case studies\\ \\ Industry Solutions\\ \\ • \\ \\ Jul 15, 2025](https://voxel51.com/blog/computer-vision-in-healthcare-12-case-studies) [![](https://cdn.sanity.io/images/h6toihm1/production/08ee8607c02188240496e20d6f5ba968ff9e5c5b-1920x1081.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Visual AI in Healthcare: 2025 Landscape\\ \\ May 12, 2025](https://voxel51.com/blog/visual-ai-in-healthcare-2025-landscape) [![](https://cdn.sanity.io/images/h6toihm1/production/63f2810f51764c0187b30ef9f0640ba49b96d70c-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Why Computer Vision in Agriculture is the Future\\ \\ Industry Solutions, Product & News\\ \\ • \\ \\ Jan 31, 2023](https://voxel51.com/blog/how-computer-vision-is-changing-agriculture-in-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-321-lllmstxt|> ## Computer Vision Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the Computer Vision Meetup — August 24, 2023 Aug 25, 2023 • 6 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#2fdf4852ad54) [Removing Backgrounds Automatically or with a User’s Language](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#f634606ebf5f) [AI at the Edge: Optimizing Deep Learning Models for Real-World Applications](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#46077385e31c) [Drones, Data, and One Direction of Computer Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#3aa8461c74b9) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#1e630f6283cc) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#6aa117791fb1) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#4a854e60abb2) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#2fdf4852ad54) [Removing Backgrounds Automatically or with a User’s Language](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#f634606ebf5f) [AI at the Edge: Optimizing Deep Learning Models for Real-World Applications](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#46077385e31c) [Drones, Data, and One Direction of Computer Vision](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#3aa8461c74b9) [Join the Computer Vision Meetup!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#1e630f6283cc) [What’s Next?](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#6aa117791fb1) [Get Involved!](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023#4a854e60abb2) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We just wrapped up the August 24, 2023 [Computer Vision Meetup](https://www.meetup.com/pro/computer-vision-meetups/), and if you missed it or want to revisit it, here’s a recap! In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. ## First, Thanks for Voting for Your Favorite Charity! In lieu of swag, we gave Meetup attendees the opportunity to help guide a $200 donation to charitable causes. The charity that received the highest number of votes this month was [Education Development Center](https://www.edc.org/) (EDC), an organization on a mission to open doors to education, employment, and healthier lives for millions of people. We are sending this event’s charitable donation of $200 to the Education Development Center on behalf of the computer vision community! ![](https://cdn.sanity.io/images/h6toihm1/production/d20a3e8f5bb9a0d2ef0a52fa7e5f5d04a56a783f-576x284.png?auto=format&dpr=2&fit=max&q=75&w=576) Missed the Meetup? No problem. Here are playbacks and talk abstracts from the event. ## Removing Backgrounds Automatically or with a User’s Language https://www.youtube.com/watch?v=ha0q44pu09w Image matting, also known as removing background, refers to extracting the accurate foregrounds in the image, which benefits many downstream applications such as film production and augmented reality. To solve this ill-posed problem, previous methods require extra user inputs with large amounts of manual effort such as trimap or scribbles. In this session, we will introduce our research works, which allow users to automatically remove the background or even flexibly choose the specific foreground by user’s language. We’ll also show some fancy demos and illustrate some downstream applications. [Jizhizi Li](https://www.linkedin.com/in/jizhizili/) has just finished her Ph.D. study in Artificial Intelligence at the University of Sydney. With several papers published in top-tier conferences and journals including CVPR, IJCV, IJCAI and Multimedia, her research interests include computer vision, image matting, multi-modal learning, and AIGC. **Q&A** - What is the difference between semantic and matting branch? - What type of filter(s) do you use for noise suppression? - How did you optimize the architecture of the encoders? - After training, did you have any observations about the internal feature space that allows discrimination between foreground and background? - How much training data did you use? - Would it make sense to use a similarity measure to interpolate background material across naturally similar types of images, to preclude the use of artificially, and non-realistic background material which may essentially warp the internal feature space to non-natural material? **Resource links** - [Jizhizi Li’s GitHub](https://github.com/JizhiziLi) - [Deep Image Matting: A Comprehensive Survey](https://github.com/jizhiziLi/matting-survey) - [\[CVPR\] Referring Image Matting (CORE A\*, CCF A)](https://github.com/JizhiziLi/RIM) - [\[IJCV\] Rethinking Portrait Matting with Privacy Preserving (CORE A\*, CCF A, IF 13.369)](https://github.com/vitae-transformer/vitae-transformer-matting) - [\[IJCV\] Bridging Composite and Real: Towards End-to-End Deep Image Matting (CORE A\*, CCF A, IF 13.369)](https://github.com/JizhiziLi/GFM) - Additional resources and links can be found on [Jizhizi’s website](https://jizhizili.github.io/homepage/#home) ## AI at the Edge: Optimizing Deep Learning Models for Real-World Applications https://www.youtube.com/watch?v=Wz-ciDGVZZ8 As AI technology continues to advance, there is a growing demand for deep learning models to tackle more complex tasks, particularly on edge devices. However, real-time performance and hardware constraints can present significant challenges in deploying these models on such devices. At SightX, we have been exploring ways to optimize deep learning models for top performance on edge devices while minimizing degradation. In this lecture, we will share our insights and techniques for deploying AI on edge devices, specifically focusing on hardware-aware optimization of deep learning models. We’ll review practical ways to effectively deploy deep learning models in real-time scenarios. [Raz Petel](https://www.linkedin.com/in/raz-petel-ai/), SightX’s Head of AI, has been tackling Computer Vision challenges with Deep Learning since 2015, aiming to enhance their efficiency, speed, compactness, and resilience. **Q&A** - Can you mention any of the families of embedded processors that you're working with? - What sort of feature extraction do you use, and do those front ends change with application? - Have you done any comparison between efficiency obtained by pruning vs. reducing the dynamic range in terms of resolution (ie, floating point dynamic range, or int sizes if you're using integers). - Have you seen performance improve as you remove features incrementally? - Does the pruning cause some negative impact to the accuracy of the model? - Could you use an auxiliary network, like an SVM, to use recursive feature elimination (based on feature coefficient magnitudes) to rank order features for iterative removal? - Are wrapper methods too computationally expensive to be practical (as opposed to filter methods)? - Which one is better, pruning or quantization? ## Drones, Data, and One Direction of Computer Vision https://www.youtube.com/watch?v=6wAEShvce80 In this lightning talk, Dan covered the current state of drone applications, some of the challenges computer vision applications face when working with drone data, plus some solutions and emerging trends like Generative AI that these applications can make use of. [Dan Gural](https://www.linkedin.com/in/daniel-gural/) is a machine learning engineer and is part of the developer relations team at Voxel51. **Q&A** _What are the best use cases for FiftyOne? For example, object detection. (Post-acquisition normalization.)_ The best use cases for drone footage in FiftyOne is to curate and identify potential areas for preprocessing. Data can be curated by removing samples unlike those expected at deployment. Data can also be preprocessed to access issues such as seen [here](https://github.com/jacobmarks/image-quality-issues) with FiftyOne plugins _Are there homography libraries in FiftyOne? (It could help with moving cameras like in drone applications.)_ Homography and other preprocessing applications are not intrinsically offered by FiftyOne, but the finished images of homography no matter the shape or resolution can be supported. **Resource links** - Learn more about the open source [FiftyOne computer vision toolset](https://docs.voxel51.com/) - Explore datasets in your browser with [Try FiftyOne](https://try.fiftyone.ai/) (no install required!) - Explore [FiftyOne Plugins](https://github.com/voxel51/fiftyone-plugins) ## Join the Computer Vision Meetup! Computer Vision Meetup membership has grown to more than [5,000 members](https://www.meetup.com/pro/computer-vision-meetups/) in just one year! The goal of the Meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of computer vision and complementary technologies. Join one of the 13 Meetup locations closest to your timezone. - [Ann Arbor](https://www.meetup.com/ann-arbor-computer-vision-meetup/) - [Austin](https://www.meetup.com/austin-computer-vision-meetup/) - [Bangalore](https://www.meetup.com/bangalore-computer-vision-meetup-group/) - [Boston](https://www.meetup.com/boston-computer-vision-meetup/) - [Chicago](https://www.meetup.com/chicago-computer-vision-meetup/) - [London](https://www.meetup.com/london-computer-vision-meetup/) - [New York](https://www.meetup.com/new-york-computer-vision-meetup/) - [Peninsula](https://www.meetup.com/peninsula-computer-vision-meetup/) - [San Francisco](https://www.meetup.com/san-francisco-computer-vision-meetup/) - [Seattle](https://www.meetup.com/seattle-computer-vision-meetup/) - [Silicon Valley](https://www.meetup.com/silicon-valley-computer-vision-meetup/) - [Singapore](https://www.meetup.com/singapore-computer-vision-meetup/) - [Toronto](https://www.meetup.com/toronto-computer-vision-meetup/) We have exciting speakers already signed up over the next few months! Become a member of the [Computer Vision Meetup closest to you](https://www.meetup.com/pro/computer-vision-meetups/), then register for the Zoom. ## What’s Next? ![](https://cdn.sanity.io/images/h6toihm1/production/de0dfab7985c4eb37e7f0473c433ecc2e991ee97-1024x576.png?auto=format&dpr=2&fit=max&q=75&w=1024) Up next on Sept 7 at 10 AM Pacific we have a great line up speakers including: - **Monitoring Large Language Models (LLMs) in Production –** Sage Elliot, Technical Evangelist – Machine Learning & MLOps at WhyLabs - **Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos –** Minye Wu, Postdoctoral researcher, KU Leuven - **Egoschmema: A Dataset for Truly Long-Form Video Understanding –** Karttikeya Mangalam, PhD student at UC Berkeley Register for the Zoom [here](https://voxel51.com/computer-vision-events/ai-ml-data-science-meetup-sept-7/). You can find a complete schedule of upcoming Meetups on [the Voxel51 Events page](https://voxel51.com/computer-vision-events/). ## **Get Involved!** There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping me over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. _The Computer Vision Meetup network is sponsored by [Vox](https://voxel51.com/) [e](https://voxel51.com/) [l51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [computer vision meetup](https://voxel51.com/blog/tag/computer-vision-meetup) [meetup](https://voxel51.com/blog/tag/meetup) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/af656fd506d11b555a019950b688830000b62f30-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Recapping the Vector Search-Themed Computer Vision Meetup — July 13, 2023\\ \\ Event Recaps, Vector Search\\ \\ • \\ \\ Jul 17, 2023](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/8e16b4cba8a09b0b3b345de76b6bde139e8814fb-1636x920.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup — August 10, 2023\\ \\ Event Recaps\\ \\ • \\ \\ Aug 14, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/991b89d515d9f7fe4eea26c14396eb51116b4a8b-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Recapping the Computer Vision Meetup – April 27, 2023\\ \\ Event Recaps\\ \\ • \\ \\ Apr 28, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-april-27-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-322-lllmstxt|> ## FiftyOne CLI Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) Exploring the CLI – FiftyOne Tips and Tricks – Aug 25th, 2023 Aug 26, 2023 • 3 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#356c2050c0d0) [Navigating the FiftyOne CLI](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#b999b830f0ac) [View Your Datasets](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#fff9a3a15d6e) [Pruning Old Datasets](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#3064abb7647f) [Looking Closer at Your Datsets](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#7dbae2ebdf68) [Draw on Labels with CLI](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#79606b2d6966) [FiftyOne Plugin CLI](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#ff445e781275) [Join the FiftyOne Community!](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#cf8b76a13aed) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#356c2050c0d0) [Navigating the FiftyOne CLI](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#b999b830f0ac) [View Your Datasets](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#fff9a3a15d6e) [Pruning Old Datasets](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#3064abb7647f) [Looking Closer at Your Datsets](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#7dbae2ebdf68) [Draw on Labels with CLI](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#79606b2d6966) [FiftyOne Plugin CLI](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#ff445e781275) [Join the FiftyOne Community!](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023#cf8b76a13aed) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we cover interesting workflows and features of FiftyOne! This week, we’ll be exploring FiftyOne’s [Command Line Interface](https://docs.voxel51.com/cli/index.html). ## **Wait, What’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! Also feel free to follow along in our [notebook](https://github.com/voxel51/fiftyone-examples/blob/cli_tnt/examples/Tips_and_Tricks_CLI.ipynb) or on [YouTube](https://www.youtube.com/watch?v=hVHPWN4vOn8)! ## **Navigating the FiftyOne CLI** Leveraging CLI commands can be a powerful way to streamline your workflow. We will do a quick dive on some powerful options FiftyOne provides that can save you time on your next computer vision project. Let's take a look. To start, let’s first turn on auto completion in the terminal. To complete any FiftyOne command just tab to see all your options. For bash, enter the command below. If you use a different terminal type, look [here](https://docs.voxel51.com/cli/index.html#tab-completion). ```bash 1eval "$(register-python-argcomplete fiftyone)" ``` ![](https://cdn.sanity.io/images/h6toihm1/production/878ad4ffe165325bdff6478982b74c1ce9f9aa16-560x155.png?auto=format&dpr=2&fit=max&q=75&w=560) ## **View Your Datasets** Sometimes, you just want to find out which datasets you have on the fly. The CLI offers the `fiftyone datasets list` command to show all your currently available datasets. It is easy for you to now know which can be loaded into the App, double check any spellings, or check if it is time to clean up some of those old datasets! ```bash 1fiftyone datasets list ``` ```raw 1open-images-v6-validation-200 2quickstart 3tips+tricks ``` ## **Pruning Old Datasets** We’ve all been there. That giant dataset you meant to delete a year ago is still sitting on your machine, hiding behind hundreds of other files. With the CLI, we can sort our datasets by their creation date and delete the ones we don't need anymore. Let’s see how: ```bash 1fiftyone datasets info --sort-by created_at ``` ```raw 1name                           created_at           last_loaded_at       version    persistent    media_type    tags    num_samples 2 3-----------------------------  -------------------  -------------------  ---------  ------------  ------------  ------  ------------- 4 5tips_and_tricks                2023-08-24 18:47:00  2023-08-24 18:55:23  0.21.6     ✓             image                 200 6 7open-images-v6-validation-200  2023-06-29 19:17:58  2023-08-24 17:44:21  0.21.6     ✓             image                 200 8 9evaluate-detections-tutorial   2023-06-28 20:27:39  2023-08-24 17:55:04  0.21.6     ✓             image                 5005 ``` In our example, the `evaluate-detections-tutorial` dataset is past its time and is ready to be deleted. Let's remove it with the following: ```bash 1fiftyone datasets delete evaluate-detections-tutorial ``` ## **Looking Closer at Your Datsets** Another common use case you can find yourself in is quickly needing to grab some basic information about your datasets. No need to boot up the App or open up the Python kernel, quick information can be grabbed straight from the CLI. Here are two quick examples: _What is the size of my dataset on disk?_ ```bash 1fiftyone datasets stats quickstart ``` ```raw 1key            value 2 3-------------  ------- 4 5samples_count  200 6 7samples_bytes  1270762 8 9samples_size   1.2MB 10 11total_bytes    1270762 12 13total_size     1.2MB ``` _Did I add my predictions to my dataset?_ ```bash 1fiftyone datasets info quickstart ``` ```raw 1Name:        quickstart 2Media type:  image 3Num samples: 200 4Persistent:  True 5Tags:        [] 6Sample fields: 7    id:           fiftyone.core.fields.ObjectIdField 8    filepath:     fiftyone.core.fields.StringField 9    tags:         fiftyone.core.fields.ListField(fiftyone.core.fields.StringField) 10    metadata:     fiftyone.core.fields.EmbeddedDocumentField(fiftyone.core.metadata.ImageMetadata) 11    ground_truth: fiftyone.core.fields.EmbeddedDocumentField(fiftyone.core.labels.Detections) 12    uniqueness:   fiftyone.core.fields.FloatField 13    predictions:  fiftyone.core.fields.EmbeddedDocumentField(fiftyone.core.labels.Detections) ``` ## **Draw on Labels with CLI** FiftyOne makes it incredibly easy to visualize and inspect predictions that are added to the dataset. If you have a dataset where you are really proud of your predictions, or maybe a prediction that needs to be sent to another data science team, samples can be exported with their drawn on labels. It is easy and seamless to perform this with the CLI, all you need is the name of the dataset and label field! ```bash 1fiftyone datasets draw -d drawn_labels -f predictions quickstart ``` ```raw 1100% |█████████████████| 200/200 [20.9s elapsed, 0s remaining, 8.4 samples/s] 2Rendered media written to 'drawn_labels' ``` ![](https://cdn.sanity.io/images/h6toihm1/production/878ad4ffe165325bdff6478982b74c1ce9f9aa16-560x155.png?auto=format&dpr=2&fit=max&q=75&w=560) ## **FiftyOne Plugin CLI** Plugins are a hot topic in the FiftyOne universe. With the ability to extend the FiftyOne App’s capabilities with easy to write Python code, the possibilities with FIftyOne Plugins are endless! Check out the [AI Art Gallery](https://voxel51.com/blog/build-your-own-ai-art-gallery/), [VoxelGPT](https://github.com/voxel51/voxelgpt), or the [Plugins Repo](https://github.com/voxel51/fiftyone-plugins)! Let’s explore the Plugins CLI: ![](https://cdn.sanity.io/images/h6toihm1/production/2112efbb160cc5810782a3c6ff6fb7f153fad3e6-1600x848.gif?auto=format&dpr=2&fit=max&q=75&w=1600) **To download a plugin:** ```bash 1fiftyone plugins download https://github.com/path/to/repo ``` **To explore what plugins are installed and operators they contain, look for the following:** ```generic 1fiftyone plugins list ``` ```raw 1plugin                version    enabled    directory 2--------------------  ---------  ---------  --------------------------------------------------- 3@voxel51/hello-world  1.0.0      ✓          /home/dan/fiftyone/__plugins__/@voxel51/hello-world 4@voxel51/voxelgpt     1.0.0      ✓          /home/dan/fiftyone/__plugins__/@voxel51/voxelgpt 5@voxel51/io           1.0.0      ✓          /home/dan/fiftyone/__plugins__/@voxel51/io ``` ```bash 1fiftyone operators list ``` ```raw 1uri                                               enabled    builtin    on_startup    unlisted    dynamic 2------------------------------------------------  ---------  ---------  ------------  ----------  --------- 3@voxel51/hello-world/count                        ✓ 4@voxel51/voxelgpt/ask_voxelgpt                    ✓ 5@voxel51/voxelgpt/ask_voxelgpt_panel              ✓                                   ✓ 6@voxel51/voxelgpt/open_voxelgpt_panel             ✓                                   ✓ 7@voxel51/voxelgpt/open_voxelgpt_panel_on_startup  ✓                     ✓             ✓ 8@voxel51/voxelgpt/vote_for_query                  ✓                                   ✓ 9@voxel51/io/add_samples                           ✓                                               ✓ 10@voxel51/io/export_samples                        ✓                                               ✓ 11@voxel51/operators/clone_selected_samples         ✓          ✓                                    ✓ 12@voxel51/operators/clone_sample_field             ✓          ✓                                    ✓ 13@voxel51/operators/rename_sample_field            ✓          ✓                                    ✓ 14@voxel51/operators/delete_selected_samples        ✓          ✓                                    ✓ 15@voxel51/operators/delete_sample_field            ✓          ✓                                    ✓ 16@voxel51/operators/print_stdout                   ✓          ✓                        ✓ ``` **To disable or enable a plugin:** ```bash 1fiftyone plugins disable @voxel51/hello-world 2 3fiftyone plugins enable @voxel51/hello-world ``` Keep an eye out for more exciting plugin content as we kick off our 10 Weeks of Plugins! Hopefully, these tips will help you speed up your workflows and create a better computer vision experience on your next project! Good Luck! To learn more about the FiftyOne CLI head over to our [CLI Docs](https://docs.voxel51.com/cli/index.html) for more information! ## **Join the FiftyOne Community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 2,000+ [FiftyOne Slack](https://slack.voxel51.com/) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 370+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [CLI](https://voxel51.com/blog/tag/cli) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [open source](https://voxel51.com/blog/tag/open-source) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/79d00d175a8098516cb2f4a7711131cbf322d01a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Finding and Correcting Mistakes – FiftyOne Tips and Tricks – Aug 18, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Aug 18, 2023](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/395ba1a1dacb511782456902b224aa4fa8552dc0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Understanding Grouped Datasets – FiftyOne Tips and Tricks – Sep 1, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 1, 2023](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/aa8b07201d3c2581a3e78336d8737022fcfca6b1-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Dynamic Groups – FiftyOne Tips and Tricks – Sep 8, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 9, 2023](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-323-lllmstxt|> ## Ask Images Anything [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Plugins](https://voxel51.com/blog/category/plugins), [Tutorials](https://voxel51.com/blog/category/tutorials) Ask Your Images Anything Sep 1, 2023 • 5 min read Article content In this article [Run Visual Question Answering Models on Your Images Without Code](https://voxel51.com/blog/ask-your-images-anything#bf031dc63835) [Visual Question Answering 🖼️❓🗨️](https://voxel51.com/blog/ask-your-images-anything#ef0d191d8658) [Plugin Overview & Functionality](https://voxel51.com/blog/ask-your-images-anything#74260785958e) [Installing the Plugin](https://voxel51.com/blog/ask-your-images-anything#e387459c0efc) [Lessons Learned](https://voxel51.com/blog/ask-your-images-anything#4a8642525260) [Conclusion](https://voxel51.com/blog/ask-your-images-anything#65a2037f9159) [Week 2 Community Plugins](https://voxel51.com/blog/ask-your-images-anything#d2380ba4ed40) In this article [Run Visual Question Answering Models on Your Images Without Code](https://voxel51.com/blog/ask-your-images-anything#bf031dc63835) [Visual Question Answering 🖼️❓🗨️](https://voxel51.com/blog/ask-your-images-anything#ef0d191d8658) [Plugin Overview & Functionality](https://voxel51.com/blog/ask-your-images-anything#74260785958e) [Installing the Plugin](https://voxel51.com/blog/ask-your-images-anything#e387459c0efc) [Lessons Learned](https://voxel51.com/blog/ask-your-images-anything#4a8642525260) [Conclusion](https://voxel51.com/blog/ask-your-images-anything#65a2037f9159) [Week 2 Community Plugins](https://voxel51.com/blog/ask-your-images-anything#d2380ba4ed40) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Run Visual Question Answering Models on Your Images Without Code Welcome to week two of _Ten Weeks of Plugins_. During these ten weeks, we will be building a FiftyOne Plugin (or multiple!) each week and sharing the lessons learned! If you’re new to them, FiftyOne Plugins provide a flexible mechanism for anyone to extend the functionality of their FiftyOne App. You may find the following resources helpful: - [FiftyOne Plugins Repo](https://github.com/voxel51/fiftyone-plugins) - [FiftyOne Plugin Docs](https://docs.voxel51.com/plugins/index.html#downloading-plugins) - Plugins Channel in the [FiftyOne Community Slack](https://slack.voxel51.com/) - Week 1 Plugins: [AI Art Gallery](https://github.com/jacobmarks/ai-art-gallery) & [Twilio Automation](https://github.com/jacobmarks/twilio-automation-plugin) Ok, let’s dive into this week’s FiftyOne Plugin! ## Visual Question Answering 🖼️❓🗨️ Imagine a world where your dataset could talk to you. A world where you could ask any question — specific or open-ended — and get a meaningful answer. How much more dynamic would your data exploration be? Welcome to Visual Question Answering (VQA), a machine learning task squarely situated at the intersection of computer vision and natural language processing. In recent years, transformer models like [Salesforce’s BLIPv2](https://arxiv.org/abs/2301.12597) have taken this open-ended data exploration to a new level. This plugin brings the power of VQA models directly to your image dataset. Now you can start asking your images those burning questions — all without a single line of code! ## Plugin Overview & Functionality For the second week of _10 Weeks of Plugins_, I built a Visual Question Answering (VQA) Plugin. This plugin allows you to ask open-ended questions to your images — effectively chatting with your data — within the FiftyOne App. Out of the box, this plugin supports two models (and two types of usage): 1. A Vision-Language Transformer (fine-tuned on the [VQAv2 dataset](https://visualqa.org/)), which is the [default VQA model](https://huggingface.co/dandelin/vilt-b32-finetuned-vqa) in the [Visual Question Answering pipeline](https://huggingface.co/tasks/visual-question-answering) from [Hugging Face’s Transformers library](https://huggingface.co/docs/transformers/index). This model is run locally. 2. [BLIP2](https://replicate.com/andreasjansson/blip-2) from Salesforce, which is accessed via a [Replicate](https://replicate.com/) inference endpoint. After you install the plugin, when you open the operators list (pressing “ `` ` ``” in the FiftyOne App) and click into the `answer_visual_question` operator, you can choose which of these models to use. Enter your question in the `question` box, and the answer will be displayed in the operator’s output: No data is added to the underlying dataset. ## Installing the Plugin If you haven’t already done so, install FiftyOne: ```bash 1pip install fiftyone ``` Then you can download this plugin from the command line with: ```bash 1fiftyone plugins download https://github.com/jacobmarks/vqa-plugin ``` Refresh the FiftyOne App, and you should see the `answer_visual_question` operator in your operators list when you press the “ `` ` ``” key. To use the Vision Language transformer (ViLT), install Hugging Face’s transformers library. ```bash 1pip install transformers ``` To use BLIPv2, set up an account with [Replicate](https://replicate.com/), install the Replicate Python library: ```bash 1pip install replicate ``` And add your Replicate API Token to your environment variables: ```bash 1export REPLICATE_API_TOKEN=... ``` You do not need both to use the plugin — the operator checks your environment variables and only shows as options models accessible via the corresponding APIs. If you want to use a different VQA model (or fine tune your own version of one of these!), locally or via API, it should be easy to extend this code. ## Lessons Learned The Visual Question Answering plugin is a Python Plugin consisting of four files: - `__init__.py`: defining the operator - `fiftyone.yml`: making the plugin _register_ for download and installation - `README.md`: describing the plugin - `requirements.txt`: listing the requirements. Both `transformers` and `replicate` are commented out by default because neither is strictly required. ### Using Selected Samples Visual question answering models like BLIPv2 typically answer questions about one image at a time. As a result, it only makes sense for the `answer_visual_question` operator to likewise act on a single image. But how does the operator _know_ which image to answer a question about? Just like the FiftyOne App, whose `session` has a `selected` attribute (see [Selecting sample](https://docs.voxel51.com/user_guide/app.html#selecting-samples)), the plugin’s context, `ctx`, has a `selected` attribute. In direct analogy with the session, `ctx.selected` is a list of sample IDs that are currently selected in the FiftyOne App. The VQA plugin looks at the number of selected samples in the `resolve_input()` method: ```python 1num_selected = len(ctx.selected) ``` And only allows the user to enter a question if exactly one sample is selected. **💡Note**: to use `ctx.selected`, when you expect the selected sample(s) to be changing, you _must_ pass `dynamic=True` into the operator’s configuration. In this case, the operator config was: ```python 1@property 2    def config(self): 3        return foo.OperatorConfig( 4            name="answer_visual_question", 5            label="VQA: Answer question about selected image", 6            dynamic=True, 7        ) ``` ### Returning Informative Outputs The VQA plugin doesn’t write anything onto the samples themselves, but we still need a way to see the results of the model’s run: the “answer”. In this plugin, I return the model’s answer as output, using the `resolve_output()` method. Output in a Python plugin works in much the same way as input. In `resolve_input()`, we create an inputs object `inputs = types.Object()`, add elements to this, e.g. `inputs.str("question", label="Question", required=True)`, and then return these inputs via `types.Property(inputs, view=...)`. In `resolve_output()`, we create an output object `outputs = types.Output()` object, add elements to this, e.g. `outputs.str("question", label="Question")`,and return these outputs via `types.Property(outputs, view=...)`. The main difference between inputs and outputs is that in `resolve_input()`, the input values come from the user. How are variables communicated to `resolve_output()`? You can return them as a dictionary from `execute()`, and then use the values, referencing them by key. In this plugin, I pass the question and answer from `execute()`: ```python 1return {"question": question, "answer": answer} ``` Then in `resolve_output()`, I access these values: ```python 1outputs.str("question", label="Question") 2outputs.str("answer", label="Answer") ``` This works for a variety of data types, not just strings! ## Conclusion If you want to chat with your entire dataset, [VoxelGPT](https://github.com/voxel51/voxelgpt) is a great option. VoxelGPT is another example of a FiftyOne Plugin, which we launched earlier this year. It translates your natural language prompts into actions that organize and explore your data. On the other hand, if you want to ask open-ended questions about specific images in your dataset — without departing from your existing workflows — then this Visual Question Answering plugin is for you! Stay tuned over the remaining weeks in the _Ten Weeks of FiftyOne Plugins_ while we continue to pump out a killer lineup of plugins! You can track our journey in our [ten-weeks-of-plugins repo](https://github.com/jacobmarks/ten-weeks-of-plugins) — and I encourage you to fork the repo and join me on this journey! ## Week 2 Community Plugins 🚀Check out this awesome [line2d](https://github.com/wayofsamu/line2d) plugin 📉by [wayofsamu](https://github.com/wayofsamu) for visualizing `(x,y)` points as a line chart! [BLIP](https://voxel51.com/blog/tag/blip) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [plugins](https://voxel51.com/blog/tag/plugins) [transformers](https://voxel51.com/blog/tag/transformers) [Visual Question Answering](https://voxel51.com/blog/tag/visual-question-answering) [VQA](https://voxel51.com/blog/tag/vqa) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/fedd0c008df994c1834839ea35bc65b34e47fb5d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Double Trouble: Eliminate Image Duplicates with FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 14, 2023](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/bad75ba72dfae8cdefc3d9afe33a1ea9a9c4ec36-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Optical Character Recognition with PyTesseract\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 21, 2023](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract) [![](https://cdn.sanity.io/images/h6toihm1/production/3efef5551e07ae9c6a1d190a5bd256b9723c78c7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Zero-Shot Prediction Plugin for FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 28, 2023](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-324-lllmstxt|> ## Grouped Datasets Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) Understanding Grouped Datasets – FiftyOne Tips and Tricks – Sep 1, 2023 Sep 1, 2023 • 5 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#9467c2642417) [What Is a Grouped Dataset?](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#ab515878eb75) [Kickstarting Your First Grouped Dataset](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#7e130f2604cf) [Working with Grouped Datasets](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#9021c60349f0) [Conclusion](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#d7e2eff1e95a) [Join the FiftyOne Community!](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#716b09fc85fb) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#9467c2642417) [What Is a Grouped Dataset?](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#ab515878eb75) [Kickstarting Your First Grouped Dataset](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#7e130f2604cf) [Working with Grouped Datasets](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#9021c60349f0) [Conclusion](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#d7e2eff1e95a) [Join the FiftyOne Community!](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023#716b09fc85fb) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we cover interesting workflows and features of FiftyOne! This week will be Part one of a two part series exploring FiftyOne’s [Grouped Datasets](https://docs.voxel51.com/user_guide/groups.html#grouped-aggregations). We aim to cover the basics like creating your first grouped dataset and explaining how you can work with and create powerful views on your new dataset. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! Also feel free to follow along in our [notebook](https://github.com/voxel51/fiftyone-examples/blob/tnt_groups/examples/Grouped%20Datasets.ipynb) or on [YouTube](https://www.youtube.com/watch?v=MgWiDCrFz7g)! ## What Is a Grouped Dataset? Before diving into the topic today, let's first understand what a group dataset is and why we would want to use one. A [grouped dataset](https://docs.voxel51.com/user_guide/groups.html) is a collection of multiple slices of samples of possibly different modalities (image, video, or point cloud) that are organized into groups. Another way to look at it is multiview datasets are also representative of grouped datasets. Samples in the same group are _related_.For example _,_ grouped datasets can be used to represent multiview scenes, where data for multiple perspectives of the same scene can be stored, visualized, and queried in ways that respect the relationships between the slices of data. ![](https://cdn.sanity.io/images/h6toihm1/production/abad3ea7f3410ff83d15ccab154f7154384deea5-1324x710.gif?auto=format&dpr=2&fit=max&q=75&w=1324) ## Kickstarting Your First Grouped Dataset ```python 1import fiftyone as fo 2 3dataset = fo.Dataset("first-group-dataset") 4dataset.add_group_field("group", default="center") ``` To get started, create a dataset and add a group field. All grouped datasets must contain a group field, where samples of our chosen media will be placed. The optional parameter for \`default\` refers to the default slice of the group that will be returned when interacting with the dataset via Python, and the slice that will be shown when you first launch a session of the FiftyOne App. This can all be changed with ease and will be detailed later. For now let's add some images to our dataset: ```python 1import fiftyone.utils.random as four 2import fiftyone.zoo as foz 3 4groups = ["left", "center", "right"] 5 6d = foz.load_zoo_dataset("quickstart") 7 8four.random_split(d, {g: 1 / len(groups) for g in groups}) 9 10filepaths = [d.match_tags(g).values("filepath") for g in groups] 11filepaths = [dict(zip(groups, fps)) for fps in zip(*filepaths)] ``` Preparing the data is easy for your grouped dataset. Define your groups and create a dictionary of the filepath of the sample as well and the group it is in. With our data ready, it is time to throw it into our grouped dataset. ```python 1samples = [] 2for fps in filepaths: 3 group = fo.Group() 4 for name, filepath in fps.items(): 5 sample = fo.Sample(filepath=filepath, 6group=group.element(name) 7) 8 samples.append(sample) 9 10dataset.add_samples(samples) 11print(dataset) ``` ```raw 1Name:        first-group-dataset 2Media type:  group 3Group slice: center 4Num groups:  145 5Persistent:  False 6Tags:        [] 7Sample fields: 8    id:       fiftyone.core.fields.ObjectIdField 9    filepath: fiftyone.core.fields.StringField 10    tags:     fiftyone.core.fields.ListField(fiftyone.core.fields.StringField) 11    metadata: fiftyone.core.fields.EmbeddedDocumentField(fiftyone.core.metadata.Metadata) 12    group:    fiftyone.core.fields.EmbeddedDocumentField(fiftyone.core.groups.Group) ``` All it takes to add your data into your dataset is to iterate through your groups and add all the images or data into their new samples. Once all the samples have been created, we can add all of them at once with `dataset.add_samples(samples)`. Congrats! That's all it takes to create to take a grouped dataset. We can visualize our first dataset with: ```python 1session = fo.launch_app(dataset) ``` ![](https://cdn.sanity.io/images/h6toihm1/production/17258ce2cbf8716b89f2c73411ed8a27049e8999-600x338.gif?auto=format&dpr=2&fit=max&q=75&w=600) ## Working with Grouped Datasets Great, so now that we have our grouped dataset, what can we do with it? If this is your first time using FiftyOne or maybe you need a refresher on creating views and working with FiftyOne datasets, I recommend brushing up with [Views Guide](https://docs.voxel51.com/user_guide/using_views.html) or some previous [Tips and Tricks](https://voxel51.com/blog/) blogs. If you are here for grouped datasets and grouped datasets only, no worries! For the most part, the Python syntax for interacting with grouped datasets is identical to that of non-grouped datasets. We can start by getting some basic information about our dataset and use that access or filter to our needs. ### What Are the Groups in My Dataset? ```python 1print(dataset.group_slices) 2print(dataset.group_media_types) ``` ```raw 1['left', 'center', 'right'] 2{'left': 'image', 'center': 'image', 'right': 'image'} ``` Here we can see what our group slices are and what kind of media inside of them. Remember that only one slice is active at a time. By default we set it to \`center\` so all functions we run to grab samples or stats will return the center slice. ```python 1sample = dataset.shuffle().first() 2 3print(sample) ``` ```raw 1, 14 15}> ``` We can see it is the center slice under the \`group\` field on the bottom. If we wanted to change the active slice, all we need to do is: ```python 1dataset.group_slice = "left" 2sample = dataset.shuffle().first() 3 4print(sample) ``` ```raw 1, 14 15}> ``` Changing the active group slice also changes it in your App as well! ![](https://cdn.sanity.io/images/h6toihm1/production/878ad4ffe165325bdff6478982b74c1ce9f9aa16-560x155.png?auto=format&dpr=2&fit=max&q=75&w=560) The next natural question in your head is, “What if I want the entire group and not just one sample?” No problem! Just grab the group id and pull like this: ```python 1sample = dataset.shuffle().first() 2group_id = sample.group.id 3group = dataset.get_group(group_id) 4 5print(group) ``` ```raw 1{'left': , 14 15}>, 'center': , 28 29}>, 'right': , 42 43}>} ``` Now we can access each piece of media in the group for the sample we are looking for. ### Iterating Through Your Grouped Dataset There are two suggested ways to iterate through your group dataset: Iterating through your active slice or iterating through each group. Depending on your use case, choose which one is best for you. To iterate through just your active slice you can use: ```python 1print(dataset.group_slice) 2 3# center 4 5for sample in dataset: 6    pass ``` Remember, you can always change the active slice with `dataset.group_slice = slice`! To iterate over your groups, you can use the `iter_groups()` function: ```python 1for group in dataset.iter_groups(): 2    pass ``` ### Creating Views in Your Grouped Dataset One of the best parts about creating a grouped dataset is you have the entire [dataset view language](https://docs.voxel51.com/user_guide/using_views.html#using-views) at your disposal to sort, slice, and search through your dataset. Iterating through, grabbing samples, or any other basic property of grouped datasets carries over when you make a new view. There are tons of possibilities of what views or subsets of your dataset you can make, so I will highlight just a few of the great possibilities: **Filter based on class**: ```python 1from fiftyone import ViewField as F 2 3dataset = foz.load_zoo_dataset("quickstart-groups") 4 5print(dataset.group_slice) 6# left 7 8# Filters based on the content in the 'left' slice 9view = ( 10 dataset 11 .match_tags("train") 12 .filter_labels("ground_truth", F("label") == "Pedestrian") 13) ``` We can even filter on multiple groups at once, using the computed metadata of the samples! ```generic 1from fiftyone import ViewField as F 2 3dataset.compute_metadata() 4 5# Match groups whose `left` image has a height of at least 640 pixels and 6# whose `right` image has a height of at most 480 pixels 7 8left_cond = F("groups.left.metadata.height") >= 640 9right_cond = F("groups.right.metadata.height") <= 480 10view = dataset.match(left_cond & right_cond) 11 12 13print(view) ``` **Create views of joined group slices:** ```python 1print(dataset.count()) # 200 2print(dataset.count("ground_truth.detections")) # 1438 3 4view3 = dataset.select_group_slices(["left", "right"]) 5 6print(view3.count()) # 400 7print(view3.count("ground_truth.detections")) # 2876 ``` If you want to create a view of just two of your groups, you can easily select multiple groups to create a new view that can become a new dataset or just the next step of your filtering or sorting process. Likewise, we can exclude individual groups from our view as well! ```generic 1# Exclude two groups at random 2view = dataset.take(2) 3 4group_ids = view.values("group.id") 5other_groups = dataset.exclude_groups(group_ids) 6assert len(set(group_ids) & set(other_groups.values("group.id"))) == 0 ``` **Aggregations**: ```python 1# Expression that computes the area of a bounding box, in pixels 2bbox_width = F("bounding_box")[2] * F("$metadata.width") 3bbox_height = F("bounding_box")[3] * F("$metadata.height") 4bbox_area = bbox_width * bbox_height 5 6print(dataset.group_slice) 7# left 8 9print(dataset.count("ground_truth.detections")) 10# 1379 11 12print(dataset.mean("ground_truth.detections[]", expr=bbox_area)) 13# 9291.53 14 ``` We can still grab the statistics we want from our dataset and apply these to be used in views: Putting it all together we can create complex views to suit exactly what you need! ```generic 1dataset = foz.load_zoo_dataset("quickstart-groups") 2 3dataset.compute_metadata() 4 5bbox_width = F("bounding_box")[2] 6bbox_height = F("bounding_box")[3] 7bbox_area = bbox_width * bbox_height 8 9view = dataset.filter_labels("ground_truth", (0.05 <= bbox_area) & (bbox_area < 0.5)) 10print(view) ``` ## Conclusion I hope this quick walkthrough has allowed you to understand grouped datasets more! There is really so much that can be accomplished and possibilities are endless. Next week we will cover dynamic grouped datasets and dive into just how much you can customize your FiftyOne datasets! Stay tuned and for more conversations or for help on grouped datasets, hop into our community Slack channel where everyone is eager to help with your FiftyOne experience! ## Join the FiftyOne Community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 2,000+ [FiftyOne Slack](https://slack.voxel51.com/) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 370+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FAQ](https://voxel51.com/blog/tag/faq) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [grouped datasets](https://voxel51.com/blog/tag/grouped-datasets) [groups](https://voxel51.com/blog/tag/groups) [open source](https://voxel51.com/blog/tag/open-source) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/79d00d175a8098516cb2f4a7711131cbf322d01a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Finding and Correcting Mistakes – FiftyOne Tips and Tricks – Aug 18, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Aug 18, 2023](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/0ecb0645c4938217bcade4d3d80cf59f7b05329b-1200x677.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision View Stages Tips and Tricks – Jan 20, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 21, 2023](https://voxel51.com/blog/fiftyone-computer-vision-view-stages-tips-and-tricks-jan-20-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-325-lllmstxt|> ## Drone Data Training Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Datasets](https://voxel51.com/blog/category/datasets) Mastering Drone Data Training Sep 26, 2023 • 6 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/computer-vision-mastering-drone-data-training#31f4dd8665e8) [Getting Started With Your Data](https://voxel51.com/blog/computer-vision-mastering-drone-data-training#131fd2ba9ad6) [Training With YOLOv5](https://voxel51.com/blog/computer-vision-mastering-drone-data-training#7ad346c6aa36) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/computer-vision-mastering-drone-data-training#31f4dd8665e8) [Getting Started With Your Data](https://voxel51.com/blog/computer-vision-mastering-drone-data-training#131fd2ba9ad6) [Training With YOLOv5](https://voxel51.com/blog/computer-vision-mastering-drone-data-training#7ad346c6aa36) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### _A Comprehensive Guide Using FiftyOne and Ultralytics YOLOv5_ Drones are revolutionizing industries from agriculture to surveillance, and the ability to accurately detect and analyze objects from aerial imagery is becoming increasingly invaluable. Embarking on the journey of harnessing the potential of drone data through advanced object detection techniques is an exciting endeavor. This tutorial dives deep into the powerful synergy between FiftyOne and Ultralytics YOLOv5, two cutting-edge tools that, when combined, offer a robust solution for training drone data. Whether you're a seasoned machine learning practitioner or a drone enthusiast venturing into the realm of AI, this guide will equip you with the knowledge and tools to navigate the complexities of drone data training and object detection. ## Wait, What’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Settings QualityAuto SpeedNormal Picture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) ## Getting Started With Your Data Drone data is very sensitive in computer vision and it is extremely important to keep all factors in mind. One small misrepresentation can severely hurt model accuracy. When curating or collecting data, it is important to take into account the angle, height, lens type, and more as to keep it consistent with how you plan on using your drone for its computer vision task. That task may be search and rescue, surveying, agriculture, or traffic detection. All of these tasks have different flight requirements and will be “seeing” the work it is doing differently. ![](https://cdn.sanity.io/images/h6toihm1/production/878ad4ffe165325bdff6478982b74c1ce9f9aa16-560x155.png?auto=format&dpr=2&fit=max&q=75&w=560) An easy to depict issue that can arise is the changing height of the terrain below the drone, which can lead to troublesome data. If there is an expected height above target for the use case, that should also be consistent on data collection. If the drone should be 50 feet above the ground but you are flying above a hill, the drone should also raise its altitude. Neglecting to do so will not only change the appearance of the image, but you can be missing coverage in your scan. None of these issues are going to totally impede training your model but could impact performance. After all the data has been collected, the only actions you can make are additional annotations or data curation. Setting some guidelines for what data is acceptable and what is not for the use case is a great start. ![](https://cdn.sanity.io/images/h6toihm1/production/83082b000c0d514f7bde08c1d022c3bdb9d0f7fa-569x215.png?auto=format&dpr=2&fit=max&q=75&w=569) For this walkthrough on using FiftyOne and Ultralytics to train an ML model on drone data, I will be using the Kaggle [Roundabout Aerial Images](https://www.kaggle.com/datasets/javiersanchezsoriano/roundabout-aerial-images-for-vehicle-detection) dataset, but you can apply the same workflow to any annotated detection dataset though, so feel free to use your own. ## Training With YOLOv5 One of the most painful experiences one goes through with training detection models is changing data types from one to another. Thankfully the days of parsing through several GB large COCO json files or scrambling together VOC xml files is over. FiftyOne allows you to natively convert your dataset to any type quickly and easily. We can start from data all the way to training in a few steps. ### Step 1: Load Your Data Let's start by loading our VOC dataset. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3import fiftyone.utils.random as four 4import fiftyone.utils.yolo as fouy 5from fiftyone import Dataset 6from fiftyone.types import VOCDetectionDataset 7 8# Path to the dataset directory 9dataset_dir = "./original/original" 10 11# Create a VOCDetectionDataset 12dataset = Dataset.from_dir( 13 dataset_dir, 14 dataset_type=VOCDetectionDataset, 15 label_field="ground_truth", 16 name = "drone_original", 17 overwrite=True 18) 19dataset.persistent = True ``` Now that all of our data is loaded, it is time for us to convert it into into the YOLO format. ### Step 2: Convert Into YOLO Format To do this we will shuffle our data; we are using FiftyOne `random_split`. With it, we can easily split our data into a 85/15 split and [export](https://docs.voxel51.com/user_guide/export_datasets.html#yolov5dataset) these datasets to the YOLO format. The operation will create a yaml file that points to the proper directories and will assist us in training. With these few steps, our data is already prepared to go train with Ultralytics! ```python 1four.random_split(dataset, {"val": 0.15, "train": 0.85}) 2val_view = dataset.match_tags("val") 3train_view = dataset.match_tags("train") 4 5val_view.export( 6 export_dir="yolo_drone/", 7 split="val", 8 dataset_type=fo.types.YOLOv5Dataset, 9) 10train_view.export( 11 export_dir="yolo_drone/", 12 split="train", 13 dataset_type=fo.types.YOLOv5Dataset, 14) ``` ### Step 3: Training With Ultralytics With our data all ready to go, next up is to train a model. Providing great tools as well as an extremely powerful model, Ultralytics is a great option when training your object detection models. To download, clone their [YOLOv5 repository](https://github.com/ultralytics/yolov5) by running the commands in your terminal: ```bash 1git clone https://github.com/ultralytics/yolov5  # clone 2cd yolov5 3pip install -r requirements.txt  # install ``` Ultralytics offers many different variations of training such as Multi-GPU, hyperparameter search, as well as pruning. For the sake of this demo, we will just be doing the default training. It is recommended that to achieve the highest accuracy you should tweak the training to fit your use case and data the best. For tutorials or tips on how to take advantage of YOLOv5 features, hop on over to their [docs](https://docs.ultralytics.com/yolov5/). For the default training, follow along with: ```bash 1python3  /path/to/yolov5/train.py --data ./yolo_drone/dataset.yaml --weights yolov5s.pt --img 640 ``` As is the case with most training runs, this can take quite a while on most machines and requires a graphics card. After the model has finished training and you are satisfied with the results, we can hop back into FiftyOne for some insights on how well the model trained. ### Step 4: Post Training Our model has been trained and our weights have been saved. Now, in order to deploy or run inference on our trained model, we need to export it to the format of our choosing. I chose TorchScript due to its fast compiled nature and ease to work with. First locate your weights at `/path/to/yolov5/runs/train/exp#/weights/best.pt` then to export to Torchscript, run the following command: ```bash 1python /path/to/yolov5/export.py --weights YOUR_WEIGHTS.pt --include torchscript ``` This will save the TorchScript file at the same location. The Ultralytics training run should have provided us with several evaluation marks such as loss or mAP. However, it is beneficial to take a deeper look to find exactly which images our model struggled with to develop a strategy to curtail this in future experiments. Let’s hop into FiftyOne for the answers! ### **Prepping the Model for Inference** ```python 1import torch 2import torchvision 3 4#Load the Model 5model = torch.load("/path/to/yolov5/runs/train/exp1/weights/best.torchscript") 6 7#Define the Postprocessing Function 8def non_max_suppression( 9 prediction, 10 conf_thres=0.25, 11 iou_thres=0.45, 12 classes=None, 13 agnostic=False, 14 multi_label=False, 15 labels=(), 16 max_det=300, 17 nm=0, # number of masks 18): 19 """Non-Maximum Suppression (NMS) on inference results to reject overlapping detections 20 21 Returns: 22 list of detections, on (n,6) tensor per image [xyxy, conf, cls] 23 """ 24 25 # Checks 26 assert 0 <= conf_thres <= 1, f'Invalid Confidence threshold {conf_thres}, valid values are between 0.0 and 1.0' 27 assert 0 <= iou_thres <= 1, f'Invalid IoU {iou_thres}, valid values are between 0.0 and 1.0' 28 if isinstance(prediction, (list, tuple)): # YOLOv5 model in validation model, output = (inference_out, loss_out) 29 prediction = prediction[0] # select only inference output 30 31 device = prediction.device 32 mps = 'mps' in device.type # Apple MPS 33 if mps: # MPS not fully supported yet, convert tensors to CPU before NMS 34 prediction = prediction.cpu() 35 bs = prediction.shape[0] # batch size 36 nc = prediction.shape[2] - nm - 5 # number of classes 37 xc = prediction[..., 4] > conf_thres # candidates 38 39 # Settings 40 # min_wh = 2 # (pixels) minimum box width and height 41 max_wh = 7680 # (pixels) maximum box width and height 42 max_nms = 30000 # maximum number of boxes into torchvision.ops.nms() 43 time_limit = 0.5 + 0.05 * bs # seconds to quit after 44 redundant = True # require redundant detections 45 multi_label &= nc > 1 # multiple labels per box (adds 0.5ms/img) 46 merge = False # use merge-NMS 47 48 mi = 5 + nc # mask start index 49 output = [torch.zeros((0, 6 + nm), device=prediction.device)] * bs 50 for xi, x in enumerate(prediction): # image index, image inference 51 # Apply constraints 52 # x[((x[..., 2:4] < min_wh) | (x[..., 2:4] > max_wh)).any(1), 4] = 0 # width-height 53 x = x[xc[xi]] # confidence 54 55 # Cat apriori labels if autolabelling 56 if labels and len(labels[xi]): 57 lb = labels[xi] 58 v = torch.zeros((len(lb), nc + nm + 5), device=x.device) 59 v[:, :4] = lb[:, 1:5] # box 60 v[:, 4] = 1.0 # conf 61 v[range(len(lb)), lb[:, 0].long() + 5] = 1.0 # cls 62 x = torch.cat((x, v), 0) 63 64 # If none remain process next image 65 if not x.shape[0]: 66 continue 67 68 # Compute conf 69 x[:, 5:] *= x[:, 4:5] # conf = obj_conf * cls_conf 70 71 # Box/Mask 72 box = xywh2xyxy(x[:, :4]) # center_x, center_y, width, height) to (x1, y1, x2, y2) 73 mask = x[:, mi:] # zero columns if no masks 74 75 # Detections matrix nx6 (xyxy, conf, cls) 76 if multi_label: 77 i, j = (x[:, 5:mi] > conf_thres).nonzero(as_tuple=False).T 78 x = torch.cat((box[i], x[i, 5 + j, None], j 79 mask[i]), 1) 80 else: # best class only 81 conf, j = x[:, 5:mi].max(1, keepdim=True) 82 x = torch.cat((box, conf, j.float(), mask), 1)[conf.view(-1) > conf_thres] 83 84 # Filter by class 85 if classes is not None: 86 x = x[(x[:, 5:6] == torch.tensor(classes, device=x.device)).any(1)] 87 88 # Apply finite constraint 89 # if not torch.isfinite(x).all(): 90 # x = x[torch.isfinite(x).all(1)] 91 92 # Check shape 93 n = x.shape[0] # number of boxes 94 if not n: # no boxes 95 continue 96 x = x[x[:, 4].argsort(descending=True)[:max_nms]] # sort by confidence and remove excess boxes 97 98 # Batched NMS 99 c = x[:, 5:6] * (0 if agnostic else max_wh) # classes 100 boxes, scores = x[:, :4] + c, x[:, 4] # boxes (offset by class), scores 101 i = torchvision.ops.nms(boxes, scores, iou_thres) # NMS 102 i = i[:max_det] # limit detections 103 if merge and (1 < n < 3E3): # Merge NMS (boxes merged using weighted mean) 104 # update boxes as boxes(i,4) = weights(i,n) * boxes(n,4) 105 iou = box_iou(boxes[i], boxes) > iou_thres # iou matrix 106 weights = iou * scores[None] # box weights 107 x[i, :4] = torch.mm(weights, x[:, :4]).float() / weights.sum(1, keepdim=True) # merged boxes 108 if redundant: 109 i = i[iou.sum(1) > 1] # require redundancy 110 111 output[xi] = x[i] 112 if mps: 113 output[xi] = output[xi].to(device) 114 115 return output 116 117def xywh2xyxy(x): 118 # Convert nx4 boxes from [x, y, w, h] to [x1, y1, x2, y2] where xy1=top-left, xy2=bottom-right 119 y = x.clone() if isinstance(x, torch.Tensor) else np.copy(x) 120 y[..., 0] = x[..., 0] - x[..., 2] / 2 # top left x 121 y[..., 1] = x[..., 1] - x[..., 3] / 2 # top left y 122 y[..., 2] = x[..., 0] + x[..., 2] / 2 # bottom right x 123 y[..., 3] = x[..., 1] + x[..., 3] / 2 # bottom right y 124 return y 125 126def format_detections(preds): 127 detections = [] 128 for x in preds: 129 label = x[5].cpu().detach().numpy() 130 score = x[4].cpu().detach().numpy() 131 box = x[:4].cpu().detach().numpy() 132 x1, y1, x2, y2 = box 133 rel_box = [x1 / w, y1 / h, (x2 - x1) / w, (y2 - y1) / h] 134 detections.append( 135 fo.Detection( 136 label=classes[int(label)], 137 bounding_box=rel_box, 138 confidence=score 139 ) 140 ) 141 return detections ``` With our model loaded and ready to go, let's take a slice of our dataset and introspect a bit. We start with creating a prediction view of 100 samples: ```python 1predictions_view = dataset.take(100, seed=51) ``` We follow up by creating an inference loop that will inference on the 100 images and add the detections to the samples. This will allow us to compare the detections in the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html) in the future. ```python 1from PIL import Image 2import cv2 3from torchvision.transforms import functional as func 4 5device = "cuda" 6classes = ["vehicle", "cycle", "truck", "bus", "van"] 7model = model.to(device) 8 9 10 11with fo.ProgressBar() as pb: 12 for sample in pb(predictions_view): 13 # Load image 14 image = cv2.imread(sample.filepath) 15 image = cv2.resize(image, (640, 640)) 16 image = func.to_tensor(image).to(device) 17 c, h, w = image.shape 18 #pint(image.shape) 19 20 # Perform inference 21 preds = model(image.unsqueeze(0)) 22 out = non_max_suppression(preds) 23 detections = format_detections(out[0]) 24 # Save predictions to dataset 25 sample["yolov5"] = fo.Detections(detections=detections) 26 sample.save() 27 28session.view = predictions_view 29 ``` With the App restarted and our new view in place, we can extract fresh insights from our data. Immediately, we're able to observe our detections superimposed on the ground truths, facilitating a performance evaluation. Additionally, the option to hide the ground truths provides a clear picture of missed detections. Another valuable suggestion is to experiment with the label confidence sliders, offering a glimpse of our high and low-confidence predictions. ![](https://cdn.sanity.io/images/h6toihm1/production/4d0a6951b69d77580ec5925f6a1ddb8ed58f44c1-1920x1080.gif?auto=format&dpr=2&fit=max&q=75&w=1600) Upon analyzing the data, I can draw conclusions that were previously inaccessible without a thorough examination: 1. The model performs strongly with high confidence on cars going around the roundabout 2. The model struggles at closely bunched cars, such as parked cars 3. The model mistakes rectangular objects as vehicles potentially ### **Investigating the Embeddings** With the [FiftyOne Brain](https://docs.voxel51.com/user_guide/brain.html), we can take an even deeper look at our data and predictions by using embeddings. By executing a couple commands beforehand, we can open a power visualization tool that shows the groups within your data and how your model is performing on them. Take a look below: ```python 1import fiftyone.brain as fob 2 3#Grab the mAP and IoUs 4results = predictions_view.evaluate_detections( 5 "yolov5", 6 gt_field="ground_truth", 7 eval_key="eval", 8 compute_mAP=True, 9) 10 11#Compute ground_truth embeddings 12results = fob.compute_visualization( 13 predictions_view, patches_field="ground_truth", brain_key="gt_viz" 14) 15 16#Compute yolov5 embeddings 17results = fob.compute_visualization( 18 predictions_view, patches_field="yolov5", brain_key="yolo_viz" 19) 20 21session.view = predictions_view ``` ![](https://cdn.sanity.io/images/h6toihm1/production/7c17b871846b84a00fbb9b240bcaf39e9accaeeb-1920x1080.gif?auto=format&dpr=2&fit=max&q=75&w=1600) We can grab awesome insights of where maybe we need to add to our dataset. We can potentially try to add more parked car views, different street views other than roundabouts, or even think about decreasing the height of the drone to increase the information we get about each car. Looking through our highest and lowest confidence predictions compared to ground truths allows us to learn more about data than a simple training run or evaluation could ever do. If you'd like learn more, check out this talk I gave about many of the topics in this blog in August at the Computer Vision Meetup. In summary, this tutorial has explored the exciting realm of harnessing the potential of drone data, spotlighting the synergy between FiftyOne and Ultralytics YOLOv5. By meticulously curating drone data and leveraging these powerful tools, this guide equips both AI enthusiasts and machine learning practitioners with the expertise to navigate the complexities of object detection training. The provided steps, from data preparation and model training to post-training analysis, underscore the importance of tailored approaches for specific use cases of working with drone data. By offering insights into model performance and utilizing embeddings for deeper analysis, this tutorial empowers users to drive accuracy and make informed decisions in the dynamic intersection of drone technology and artificial intelligence. [drones](https://voxel51.com/blog/tag/drones) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [open source](https://voxel51.com/blog/tag/open-source) ![](https://cdn.sanity.io/images/h6toihm1/production/3b39056326e925c10b46da1324bc3c5840a1629c-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Dan Gural Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/7ac3933208d45ce7ee8d08eb935a43eadfa2475e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Teaching Androids to Dream of Sheep\\ \\ Computer Vision, Datasets, Product & News, Vector Search\\ \\ • \\ \\ Aug 7, 2023](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep) [![](https://cdn.sanity.io/images/h6toihm1/production/c9967da6d043cb267a4432030ad161d3443f5ba2-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Segment Anything in a CT Scan with NVIDIA VISTA-3D\\ \\ Computer Vision, Plugins\\ \\ • \\ \\ Jul 2, 2024](https://voxel51.com/blog/segment-anything-in-a-ct-scan-with-nvidia-vista-3d) [![](https://cdn.sanity.io/images/h6toihm1/production/aa202a26af141cab6f1c7785026a9edf71ef8101-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How to Make the Best Self-Driving Dataset\\ \\ Computer Vision\\ \\ • \\ \\ Jan 15, 2025](https://voxel51.com/blog/how-to-make-the-best-self-driving-dataset) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-326-lllmstxt|> ## Custom Computer Vision Applications [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Plugins](https://voxel51.com/blog/category/plugins), [Tutorials](https://voxel51.com/blog/category/tutorials) Build Custom Computer Vision Applications Sep 7, 2023 • 6 min read Article content In this article [Embed Custom UIs Directly into Data-Centric Workflows](https://voxel51.com/blog/build-custom-computer-vision-applications#7e3e3cbabf4f) [Custom YouTube Player Panel 📺🔧🔘](https://voxel51.com/blog/build-custom-computer-vision-applications#de89a6ca1d91) [Plugin Overview & Functionality](https://voxel51.com/blog/build-custom-computer-vision-applications#2b8329cabe68) [Installing the Plugin](https://voxel51.com/blog/build-custom-computer-vision-applications#f8943aafae29) [Lessons Learned](https://voxel51.com/blog/build-custom-computer-vision-applications#6171d62212b7) [Conclusion](https://voxel51.com/blog/build-custom-computer-vision-applications#d474398c13a5) In this article [Embed Custom UIs Directly into Data-Centric Workflows](https://voxel51.com/blog/build-custom-computer-vision-applications#7e3e3cbabf4f) [Custom YouTube Player Panel 📺🔧🔘](https://voxel51.com/blog/build-custom-computer-vision-applications#de89a6ca1d91) [Plugin Overview & Functionality](https://voxel51.com/blog/build-custom-computer-vision-applications#2b8329cabe68) [Installing the Plugin](https://voxel51.com/blog/build-custom-computer-vision-applications#f8943aafae29) [Lessons Learned](https://voxel51.com/blog/build-custom-computer-vision-applications#6171d62212b7) [Conclusion](https://voxel51.com/blog/build-custom-computer-vision-applications#d474398c13a5) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Embed Custom UIs Directly into Data-Centric Workflows Welcome to week three of _Ten Weeks of Plugins_. During these ten weeks, we will be building a FiftyOne plugin (or multiple!) each week and sharing the lessons learned! If you’re new to them, FiftyOne Plugins provide a flexible mechanism for anyone to extend the functionality of their FiftyOne App. You may find the following resources helpful: - [FiftyOne Plugins Repo](https://github.com/voxel51/fiftyone-plugins) - [FiftyOne Plugin Docs](https://docs.voxel51.com/plugins/index.html#downloading-plugins) - Plugins Channel in the [FiftyOne Community Slack](https://slack.voxel51.com/) Here’s what we’ve built so far: - Week 1: [AI Art Gallery](https://github.com/jacobmarks/ai-art-gallery) & [Twilio Automation](https://github.com/jacobmarks/twilio-automation-plugin) - Week 2: [Visual Question Answering](https://github.com/jacobmarks/vqa-plugin) Ok, let’s dive into this week’s FiftyOne Plugin! ## Custom YouTube Player Panel 📺🔧🔘 Over the first two weeks of this 10 Weeks of Plugins journey, all three of the plugins that I built were pure [Python plugins](https://docs.voxel51.com/api/fiftyone.operators.types.html#module-fiftyone.operators.types) plugins that allow you to create _operators_, which execute custom code behind the scenes. These plugins enable a plethora of workflows, and as primarily a Python programmer myself, I find the process for creating Python plugins to be quite intuitive, especially after creating a few. Sometimes, however, it’s nice to be able to create a custom user interface. Take [VoxelGPT](https://github.com/voxel51/voxelgpt) for instance: the chatbot-like interface (plus easy statefulness!) is really only possible when given its own devoted space within the FiftyOne App. In FiftyOne, creating custom interfaces like this is possible via JavaScript Plugins, which give you blank canvases on which to design workflows and experiences. In FiftyOne, these “canvases” are called `Panels`, and when you write a JavaScript plugin, you can build your own panels from scratch. Using JavaScript (and React Material UI components) you can essentially design a custom UI for yourself and/or others, which you can pull up at the click of a button. You’ll see what I mean shortly! For the third week of _10 Weeks of Plugins_, I built a YouTube Player Panel Plugin. For this plugin, I built a relatively simple UI consisting of a YouTube video player and a few selectors to make it interactive. While this plugin itself is not terribly complex, and my UI is not terribly artistic, it showcases the basic elements required to build a panel to your liking. The best part: I used ChatGPT to write the JavaScript for me! ## Plugin Overview & Functionality The YouTube Player Panel Plugin is a joint Python/JavaScript plugin with a single operator: `open_youtube_player_panel`, which opens a custom JavaScript component named `YouTubePlayerPanel`. You can choose to play a video from a preset list of YouTube videos from the [Voxel51 YouTube channel](https://www.youtube.com/@voxel51), or you can play a YouTube video from URL. There are three ways to execute the `open_youtube_player_panel` operator and open the panel: 1. Press the YouTube player button in the [Sample Actions Menu](http://samples_grid_actions/): \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop 1. Click on the `+` icon next to the `Samples` tab and select `YouTube Player` from the dropdown menu: \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop 1. Press “ `` ` ``” to pull up your list of operators, and select `open_youtube_player_panel`: \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop The Python portion of the plugin is exceedingly simple: - Storing an SVG icon for YouTube in an `assets` folder within the plugin’s directory, I specified this as the icon to use to represent the operator in two ways: first, in the operator’s `config()` method, so that this would be the icon in the operators list, and second, in its `resolve_placement()` method, so that the button in the action menu is also the YouTube icon. - In the `execute()` method, the operator triggers an `open_panel` action, and the params passed into `trigger()` instruct that the panel to be opened is named `YouTubePlayerPanel`. The JavaScript portion is a bit more involved. It consists of: - `YouTubeIcon`: the SVG icon that appears next to the label `YouTube Player` on the tab in the panel. - `YoutubeEmbed`: an [iframe](https://www.techtarget.com/whatis/definition/IFrame-Inline-Frame#:~:text=An%20inline%20frame%20(iframe)%20is,webpage%20within%20the%20parent%20page.) component (basically an embedded live view of HTML loaded on another page — in this case YouTube). This component takes as an argument an `embedId`, which is the unique identifier for YouTube videos. - `YouTubePlayerPanel`: the custom panel containing the YouTube player iframe, a list of preset videos, a text field for inputting a different video link, and the logic for updating the display based on values of the associated variables. - `registerComponent`: this registers the `YouTubePlayerPanel` we defined as a `Panel` with the name “YouTubePlayerPanel”, so that we can reference it (and open it) from our Python operator. ## Installing the Plugin If you haven’t already done so, install FiftyOne: ```bash 1pip install fiftyone ``` Then you can download this plugin from the command line with: ```bash 1fiftyone plugins download https://github.com/jacobmarks/vqa-plugin ``` Refresh the FiftyOne App, and you should see the YouTube button show up in the actions menu. ## Lessons Learned ### Icon Rendering FiftyOne plugins are very flexible when it comes to icon rendering. If you desire, you can have distinct icons for an operator in the action menu, the operator list, and at the top of its panel. Even though this particular plugin only uses one icon, I try to stick to the convention of putting all icons in an `assets` folder, and then referencing the svg file via absolute path within the plugin’s directory. For example: ```python 1icon="/assets/youtube.svg" ``` For the icon in the panel itself, I had to create a separate icon component in JavaScript. The paths in the svg were the same, but I had to use different `viewBox` coordinates to make all of the icons line up nicely in their respective contexts. ### Trigger Composes Operators In the [AI Art Gallery](https://github.com/jacobmarks/ai-art-gallery) and [Twilio Automation](https://github.com/jacobmarks/twilio-automation-plugin) plugins, I had used the `ctx.trigger()` method to perform operations like reloading samples ( `ctx.trigger(“reload_samples”)`), and reloading the dataset ( `ctx.trigger(“reload_dataset”)`). I was even aware from [VoxelGPT](https://github.com/voxel51/voxelgpt) that you could use `ctx.trigger()` to set the session’s view. But it wasn’t until this YouTube Player Plugin that it clicked for me that `ctx.trigger()` [allows you to compose operators](https://docs.voxel51.com/plugins/index.html#operator-composition). You can trigger a Python operator from JavaScript, or a JavaScript operator from Python. You can even trigger operators defined in other plugins! In this plugin the Python operator’s `execute()` method looks like: ```python 1def execute(self, ctx): 2 ctx.trigger( 3 "open_panel", 4 params=dict( 5 name="YouTubePlayerPanel", isActive=True, layout="horizontal" 6 ), 7 ) ``` What this means is that when executed, the `open_youtube_player_panel` operator is triggering another operator, `open_panel`, with the specific input parameters. In this case, the `open_panel` operator is opening the panel named `YouTubePlayerPanel`, which we defined in our `YouTubePlayerPlugin.tsx` file, and is splitting the canvas of the FiftyOne App in two horizontally (as opposed to vertically). You can do so much with these operator compositions. I’m excited to explore this further in coming weeks! ### No JavaScript?; No Problem! My biggest takeaway from building this plugin is that you don’t actually need to know any JavaScript, React, or Typescript to build a “JavaScript” plugin in FiftyOne. The greatest strength of FiftyOne’s JavaScript plugins is their straightforward interface. Because FiftyOne plugin components map fairly directly onto React Material UI components, I found that the [FiftyOne Plugin docs](https://docs.voxel51.com/plugins/index.html) plus the [Material UI docs](https://mui.com/material-ui/getting-started/) plus ChatGPT was more than enough to turn my vision into reality. If you’d like to dive deeper into this, let me know and I’ll put together some resources :) ## Conclusion FiftyOne plugins are incredibly flexible. They give you the blank canvas you need to create tailor-made data-centric machine learning applications. Playing YouTube videos in the FiftyOne App is cool, but the real value is the ability to educate, to showcase workflows, and to create a compelling story — all with data at the forefront! Stay tuned for the remainder of these ten weeks as we continue to pump out a killer lineup of plugins! You can track our journey in our [ten-weeks-of-plugins repo](https://github.com/jacobmarks/ten-weeks-of-plugins) — and I encourage you to fork the repo and join me on this journey! [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [javascript](https://voxel51.com/blog/tag/javascript) [material UI](https://voxel51.com/blog/tag/material-ui) [open source](https://voxel51.com/blog/tag/open-source) [plugins](https://voxel51.com/blog/tag/plugins) [react](https://voxel51.com/blog/tag/react) [youtube](https://voxel51.com/blog/tag/youtube) ![](https://cdn.sanity.io/images/h6toihm1/production/d58692baec7c64699806d60d25a0d14f534105fa-300x300.png?auto=format&dpr=2&fit=max&q=75&w=42) Jacob Marks Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/fedd0c008df994c1834839ea35bc65b34e47fb5d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Double Trouble: Eliminate Image Duplicates with FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 14, 2023](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/bad75ba72dfae8cdefc3d9afe33a1ea9a9c4ec36-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Optical Character Recognition with PyTesseract\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 21, 2023](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract) [![](https://cdn.sanity.io/images/h6toihm1/production/3efef5551e07ae9c6a1d190a5bd256b9723c78c7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Zero-Shot Prediction Plugin for FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 28, 2023](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-327-lllmstxt|> ## FiftyOne Community Update [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) FiftyOne Computer Vision Community Update – September 2023 Sep 8, 2023 • 5 min read Article content In this article [Community Spotlight](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#8e76f3847518) [Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#32f55ec6de4a) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#fe4a44b23b76) [Community Integrations](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#eb091b03892a) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#27096776e450) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#f686c6d5ab2e) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#a054f333d9ef) [Upcoming Computer Vision Events](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#32e2bb4abbc3) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#de118c3cb798) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#9deaac66d609) [What’s Next?](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#4b665715147e) In this article [Community Spotlight](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#8e76f3847518) [Community Rewards](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#32f55ec6de4a) [Product Releases](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#fe4a44b23b76) [Community Integrations](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#eb091b03892a) [FiftyOne on GitHub](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#27096776e450) [FiftyOne Community Slack](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#f686c6d5ab2e) [Computer Vision Meetups](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#a054f333d9ef) [Upcoming Computer Vision Events](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#32e2bb4abbc3) [New Docs, Blogs, Videos, and Tutorials](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#de118c3cb798) [Voxel51’s Commitment to Open Source and Community](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#9deaac66d609) [What’s Next?](https://voxel51.com/blog/fiftyone-computer-vision-community-update-sep-2023#4b665715147e) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to the monthly blog series where we bring you up to speed on recent happenings in the FiftyOne community and celebrate noteworthy milestones. 🙌 🚀 ## Community Spotlight We love hearing how FiftyOne helps you solve challenges and reach new heights! Curious what sorts of use cases are possible with [FiftyOne](https://voxel51.com/fiftyone/)? Here’s a highlight from the open source FiftyOne community. ![](https://cdn.sanity.io/images/h6toihm1/production/2c4107b73ce12103a76143113ce43af96f9bf5ef-1024x628.png?auto=format&dpr=2&fit=max&q=75&w=1024) [Aidence](https://www.aidence.com/) turns complex data science into practical and intuitive solutions that make physicians’ work easier. The company made it their mission to provide AI that empowers healthcare and pharmaceutical professionals to deliver faster, more precise diagnostics and treatments. > “I use FiftyOne on a daily basis to improve the quality of our data and visually inspect model predictions. I much prefer the interactive experience of FiftyOne over static notebooks with matplotlib figures.“ > > — Marijn Lems, Machine Learning Engineer, Aidence Learn more about computer vision use cases in healthcare in the latest industry spotlight: [How Computer Vision Is Changing Healthcare in 2023](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare-in-2023/) ## Community Rewards ![](https://cdn.sanity.io/images/h6toihm1/production/e3ed352e497d0e500ff4a1484b8422b3c9bef5cb-600x600.png?auto=format&dpr=2&fit=max&q=75&w=600) Is your organization using FiftyOne to solve interesting computer vision problems? [Share your success story](https://voxel51.com/fiftyone-computer-vision-success-story-submission/) and claim a box of community rewards as a thank you! ## Product Releases In August there were a few point releases to both the open source and Teams versions of FiftyOne. Here are the links to the releases learn more: ### Open Source FiftyOne - [FiftyOne 0.21.6](https://docs.voxel51.com/release-notes.html#fiftyone-0-21-6) \- August 8 - [FiftyOne 0.21.5](https://docs.voxel51.com/release-notes.html#fiftyone-0-21-5) \- August 7 ### FiftyOne Teams - [FiftyOne Teams 1.3.6](https://docs.voxel51.com/release-notes.html#fiftyone-teams-1-3-6) \- August 8 - [FiftyOne Teams 1.3.5](https://docs.voxel51.com/release-notes.html#fiftyone-teams-1-3-5) \- August 7 ## Community Integrations August brought us three new exciting integrations with computer vision models. ### Segment Anything Several [Segment Anything](https://segment-anything.com/) models have been added to the [FiftyOne Model Zoo](https://docs.voxel51.com/user_guide/model_zoo/). (SAM is a promptable segmentation system with zero-shot generalization to unfamiliar objects and images, without the need for additional training.) - [Segment-anything-vitb-torch](https://docs.voxel51.com/user_guide/model_zoo/models.html#segment-anything-vitb-torch): ViT-B/16 backbone trained on SA-1B - [Segment-anything-vith-torch](https://docs.voxel51.com/user_guide/model_zoo/models.html#segment-anything-vith-torch): ViT-H/16 backbone trained on SA-1B. - [Segment-anything-vitl-torch](https://docs.voxel51.com/user_guide/model_zoo/models.html#segment-anything-vitl-torch): ViT-L/16 backbone trained on SA-1B. ### DINOv2 Several [DINOv2](https://github.com/facebookresearch/dinov2) models produce high-performance visual features that can be directly employed with classifiers as simple as linear layers on a variety of computer vision tasks; these visual features are robust and perform well across domains without any requirement for fine-tuning. The models were pretrained on a dataset of 142 M images without using any labels or annotations. - [Dinov2-vitb14](https://docs.voxel51.com/user_guide/model_zoo/models.html#dinov2-vitb14): ViT-B/14 distilled. - [Dinov2-vitg14](https://docs.voxel51.com/user_guide/model_zoo/models.html#dinov2-vitg14): ViT-g/14. - [Dinov2-vitl14](https://docs.voxel51.com/user_guide/model_zoo/models.html#dinov2-vitl14): ViT-L/14 distilled. ### PyTorch Hub Integration FiftyOne integrates natively with [PyTorch Hub](https://pytorch.org/hub/), so you can load any Hub model and run inference on your FiftyOne datasets with just a few lines of code! Learn more about the integration in the [FiftyOne Docs](https://pytorch.org/hub/). ## FiftyOne on GitHub GitHub is home to the open source FiftyOne project. Here’s the latest snapshot of what’s happening in the [FiftyOne GitHub repo](https://github.com/voxel51/fiftyone): - Total stars: 4,100+ - Total contributors: 72 - Total used by: 379 - Total forks: 410 - Total issues closed so far: 872 ## FiftyOne Community Slack The FiftyOne Community [Slack channel](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) is where you can join more than 2,000 machine learning engineers and data scientists using FiftyOne to improve the quality of their computer vision data and build better models. Last month alone we had almost 150 new members join. Ask questions, answer questions, or simply follow along with the discussion! To make it easy to catch the highlights, every Friday we recap interesting questions and answers from Slack in [Tips & Tricks blog series](https://voxel51.com/blog/category/tips-tricks/). Recent posts include: - [Exploring the CLI – FiftyOne Tips and Tricks – Aug 25](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023/) - [Finding and Correcting Mistakes – FiftyOne Tips and Tricks – Aug 18](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023/) - [FiftyOne Sample Fields Tips and Tricks – Aug 11](https://voxel51.com/blog/fiftyone-sample-fields-tips-and-tricks-aug-11-2023/) - [FiftyOne Computer Vision Tips and Tricks – Aug 4](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-aug-4-2023/) ## Computer Vision Meetups \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Voxel51 sponsors 13 virtual [Computer Vision Meetups](https://www.meetup.com/pro/computer-vision-meetups/) and 12 [AI, Machine Learning, and Data Science Meetups](https://www.meetup.com/pro/ai-machine-learning-data-science-network/) around the world. (To join, visit the previous Meetup links and scroll down to find the location friendliest to your time zone.) The Computer Vision Meetups are geared towards data scientists, machine learning engineers, and open source enthusiasts who want to expand their knowledge of computer vision and complementary technologies. We put an emphasis on open source software, and speakers who are computer vision practitioners or academics doing research in the field. This month’s Meetups include: ### Sept 14 Computer Vision Meetup \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop - ARMBench: An Object-Centric Benchmark Dataset for Robotic Manipulation - Amazon Robotics team - From Model to the Edge, Putting Your Model into Production - Joy Timmermans, Secury360 - Optimizing Distributed Fine-Tuning Workloads for Stable Diffusion with the Intel Extension for PyTorch on AWS - Eduardo Alvarez, Intel ### Recapping August’s Meetups If you missed any of last month’s Meetups, you can get the executive summaries and links to the video playbacks here: - [Recapping the Computer Vision Meetup — August 24](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-24-2023/) - [Recapping the Computer Vision Meetup — August 10](https://voxel51.com/blog/recapping-the-computer-vision-meetup-august-10-2023/) ## Upcoming Computer Vision Events In addition to Meetups, we invite you to join us for one or more of these upcoming [events](https://voxel51.com/computer-vision-events/): - Sept 20 - ​​ [FiftyOne Community Office Hours](https://voxel51.com/computer-vision-events/) - Sept 20 - [ADAS & Autonomous Vehicle Technology Expo](https://www.autonomousvehicletechnologyexpo.com/california/index.php) - Sept 26 - [The AI Conference](https://aiconference.com/) - Sept 27 - [Getting Started with FiftyOne Workshop](https://voxel51.com/computer-vision-events/) ## New Docs, Blogs, Videos, and Tutorials We want everyone to be successful with FiftyOne, and one of the ways we try to do that is by publishing resources that you might find helpful and handy. Here’s a list of some of the new [documentation](https://docs.voxel51.com/), [blogs](https://voxel51.com/blog/), [videos](https://www.youtube.com/@voxel51/videos), [tutorials](https://docs.voxel51.com/tutorials/index.html), [integrations](https://docs.voxel51.com/integrations/index.html), and [cheat sheets](https://docs.voxel51.com/cheat_sheets/index.html) that you may want to check out. ### Blogs - [Ask Your Images Anything](https://voxel51.com/blog/ask-your-images-anything/) - [How Computer Vision Is Changing Healthcare in 2023](https://voxel51.com/blog/how-computer-vision-is-changing-healthcare-in-2023/) - [Build Your Own AI Art Gallery](https://voxel51.com/blog/build-your-own-ai-art-gallery/) - [The Spirit of Competition – OpenCV AI Competition 2023](https://voxel51.com/blog/opencv-ai-competition-2023/) - [Spending My First Week With FiftyOne](https://voxel51.com/blog/spending-my-first-week-with-fiftyone/) - [Celebrating Three Years of FiftyOne!](https://voxel51.com/blog/celebrating-three-years-of-fiftyone/) - [Teaching Androids to Dream of Sheep](https://voxel51.com/blog/teaching-androids-to-dream-of-sheep/) ### Videos - [Finding and Correcting Mistakes — FiftyOne Computer Vision Tips and Tricks - Aug 18](https://www.youtube.com/watch?v=WDl80g7_SBw&t=4s) - [How to Install, Use, and Write FiftyOne Plugins](https://www.youtube.com/watch?v=iJJaHudKlLI&t=298s) ## Voxel51’s Commitment to Open Source and Community Open source, transparency, and giving back to the computer vision community is what we are all about! Whether it’s developing the open source [FiftyOne computer vision toolset](https://github.com/voxel51/fiftyone) to help engineers and data scientists build high-quality datasets and models, sponsoring [Meetups](https://www.meetup.com/pro/computer-vision-meetups/) to help members boost their computer vision knowledge, or [giving to charitable causes](https://voxel51.com/charitable-giving/) on behalf of the community, Voxel51 is committed to bringing transparency and clarity to the world’s data. ## What’s Next? - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [Community Update](https://voxel51.com/blog/tag/community-update) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [OSS community](https://voxel51.com/blog/tag/oss-community) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/338b38d41e6072dd11af86f21f5309337c52f36b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – May ‘23\\ \\ Product & News\\ \\ • \\ \\ May 5, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/286bca6ba83c8386a9b92750a9249e2e8aba7d8a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – June ‘23\\ \\ Product & News\\ \\ • \\ \\ Jun 2, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-june-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/869a02098d1898869a250f4a5a23648c479af7da-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – April ‘23\\ \\ Product & News\\ \\ • \\ \\ Apr 6, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-328-lllmstxt|> ## Dynamic Groups in FiftyOne [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) Dynamic Groups – FiftyOne Tips and Tricks – Sep 8, 2023 Sep 9, 2023 • 4 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#515d8715a539) [What Is a Dynamic Group?](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#a32e792f4622) [Grouping Together Video Frames](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#abde898c2b91) [Working with Dynamically Grouped Views](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#10fab7f476d2) [Applying Dynamic Groups to Your Data](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#12a01b6a012c) [Conclusion](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#141ce8f08cbf) [Join the FiftyOne Community!](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#7d6dece3ca5d) [What’s Next?](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#e27fd789ada7) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#515d8715a539) [What Is a Dynamic Group?](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#a32e792f4622) [Grouping Together Video Frames](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#abde898c2b91) [Working with Dynamically Grouped Views](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#10fab7f476d2) [Applying Dynamic Groups to Your Data](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#12a01b6a012c) [Conclusion](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#141ce8f08cbf) [Join the FiftyOne Community!](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#7d6dece3ca5d) [What’s Next?](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023#e27fd789ada7) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we cover interesting workflows and features of FiftyOne! This week is part two of a two-part series exploring FiftyOne’s [grouped datasets](https://docs.voxel51.com/user_guide/groups.html#grouped-aggregations). If you missed last week, you can still catch it [here](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023/). Today we dive into dynamic grouped datasets and how to create powerful new ways to organize and view your data. ## Wait, What’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! Also feel free to follow along in our [notebook](https://github.com/voxel51/fiftyone-examples/blob/dynamic_groups/examples/Dynamic%20Group.ipynb) or on [YouTube](https://www.youtube.com/watch?v=WDl80g7_SBw&feature=youtu.be)! ## What Is a Dynamic Group? Let's take a look at what dynamic grouping is and when we may want to do this. [Dynamic grouping](https://docs.voxel51.com/user_guide/using_views.html#view-groups) is a feature in FiftyOne that allows you to group samples in your dataset by a particular field or expression. In the most basic of examples we can take a standard classification dataset and dynamically group it to create a new view. It is all possible with the `group_by` method. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5dataset = foz.load_zoo_dataset("cifar10", split="test") 6 7# Take 100 samples and group by ground truth label 8view = dataset.take(100, seed=51).group_by("ground_truth.label") 9 10print(view.media_type)  # group 11print(len(view))  # 10 ``` Our new view is a dynamic group that instead of showing all 100 images, we can now organize by label. Opening the FiftyOne App allows us to see our 10 groups and click to see all of our different samples within each label group. ![](https://cdn.sanity.io/images/h6toihm1/production/f9dd58a6651de61473472a84b81690c602a01ca1-600x338.gif?auto=format&dpr=2&fit=max&q=75&w=600) A dynamic group view is a collection of samples that has been grouped based on a specified matching condition. You can group by scene number, label, [view expression](https://docs.voxel51.com/user_guide/using_views.html#), or more generally any field on your samples. We will walk through a couple examples of how this can be done and how to work with the dataset afterwards. ## Grouping Together Video Frames Here's another way to use `group_by`. In this example, we have a set of videos that I need to break down into frames for my model and use case. We can load the videos for our example with the `quickstart-video` and create the frames from it with `to_frames(sample_frames=True)` ```python 1dataset2 = ( 2 foz.load_zoo_dataset("quickstart-video") 3 .to_frames(sample_frames=True) 4 .clone() 5) 6print(dataset2) #1279 samples 7session = fo.launch_app(dataset2) 8 ``` ![](https://cdn.sanity.io/images/h6toihm1/production/fec7b23838b23cb94088f2ba8618011646ca544c-1203x528.png?auto=format&dpr=2&fit=max&q=75&w=1203) We used to be greeted with each individual video in our dataset. Now, instead, we are met with the first several frames from our first video. To see a sample from our next video, we will need to scroll and scroll to see. No worries, as dynamic group views can clean this up for us if you like. We can group and order our data by using `group_by` with the `sample_id` from the original video and `order_by` the frame number. ```python 1view2 = dataset2.group_by("sample_id", order_by="frame_number") 2 3print(len(view2))  # 10 4print(view2.values("frame_number")) 5 6session.view = view2 ``` ![](https://cdn.sanity.io/images/h6toihm1/production/84624eda9f72725d4b82e05594803bf62a551021-797x621.png?auto=format&dpr=2&fit=max&q=75&w=797) This can also be done within the App by using the dynamic group button! Take a look below on how to do the same only using the UI! ![](https://cdn.sanity.io/images/h6toihm1/production/ac376beb6f8254dfee982086ea3f43fd590436b5-600x338.gif?auto=format&dpr=2&fit=max&q=75&w=600) ## Working with Dynamically Grouped Views Now that we have our data grouped, in order to access this dynamic group, we use the `get_dynamic_group` method to grab the group we want. `get_dynamic_group` works just like `get_group` does in [grouped datasets](https://docs.voxel51.com/api/fiftyone.core.dataset.html#fiftyone.core.dataset.Dataset.get_group), only this time on our dynamic group view! This will bring us a view with our selected group and all its samples. For example, we can grab the group from the first video in our dataset. ```python 1sample_id = dataset2.take(1).first().sample_id 2 3video = view2.get_dynamic_group(sample_id) 4 5print(video.values("frame_number")) ``` Iterating through any datasets or views is also an important and useful method to do. When iterating through a dynamic group view, each iteration gives you a group from the dataset. If we look at the CIFAR dataset again, we can see this in action: ```generic 1# Sort the groups by label 2sorted_view = view.sort_by("ground_truth.label") 3 4for sample in sorted_view: 5 print(sample.ground_truth.label) ``` ```raw 1airplane 2automobile 3bird 4cat 5deer 6dog 7frog 8horse 9ship 10truck ``` To flatten or unroll a view that you have created using dynamic groups, you can use the \`flatten\` function to undo any groups and put it back into a flat collection of samples. ```python 1# Unwind the sorted groups back into a flat collection 2flat_sorted_view = sorted_view.flatten() 3 4print(len(flat_sorted_view))  # 1000 5print(flat_sorted_view.values("ground_truth.label")) # ['airplane', 'airplane', 'airplane', ..., 'truck'] ``` ## Applying Dynamic Groups to Your Data When working with FiftyOne, it is important to think of dynamic groups as one of the many tools in your toolbag. Organizing and bringing clarity to how your data is related is highly valuable in an ML workflow. Pairing related samples can not only help bring you insights into your data, it can speed up your workflow by making training easier. Instead of managing multiple input streams from multiple datasets, with groups you can limit it to just one. Dynamic groups can be used alongside sensor data to help group every instance, let's say a frame, with each output of any number of sensors. This can allow you to easily track sequences in data and bring a more temporal feel to FiftyOne. It can also be great for quick exploration of your dataset by utilizing dynamic group views. A great example can be looking at the groups with the highest numbers detections to make sure there are no cluttered or crowded boxes. We can easily group by number of detections. ```python 1dataset3 = foz.load_zoo_dataset("quickstart") 2 3# Group samples by the number of ground truth objects they contain 4expr = F("ground_truth.detections").length() 5view3 = dataset3.group_by(expr) 6 7print(len(view3))  # 26 8print(len(dataset3.distinct(expr)))  # 26 ``` For a great use case of a sophisticated view created from dynamic groups, check out [Spring](https://try.fiftyone.ai/datasets/spring/samples), a dataset hosted on the [try.fiftyone.ai](https://try.fiftyone.ai/) website so you can instantly see it in action! Explore this dataset to see each frame and all its associated sensor data grouped nicely. This type of grouping can be applied across industries from automotive, agriculture, aerospace, and more! ![](https://cdn.sanity.io/images/h6toihm1/production/b1a63ef4dc2d016162e932faa9b955c7eb273ff3-1999x957.png?auto=format&dpr=2&fit=max&q=75&w=1600) ## Conclusion Wrapping up today's Tips and Tricks, I hope you now feel comfortable using groups in FiftyOne whether through a grouped dataset or a dynamic group view. Alongside these features, organizing multiview or multimodal data can be made easy and done on the fly. Accessing, changing, iterating, or reverting dynamic group views is quick and easy to learn! As always, if you are looking to learn more about dynamic groups or have any questions, I highly encourage you to head over to the community Slack channel for more information! ## Join the FiftyOne Community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 2,000+ [FiftyOne Slack](https://slack.voxel51.com/) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 370+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) ## What’s Next? - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [group datasets](https://voxel51.com/blog/tag/group-datasets) [open source](https://voxel51.com/blog/tag/open-source) [video datasets](https://voxel51.com/blog/tag/video-datasets) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/a3a918e30b0553723b9392ea90763379f98480a0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – April 7, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Apr 7, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-april-7-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/79d00d175a8098516cb2f4a7711131cbf322d01a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Finding and Correcting Mistakes – FiftyOne Tips and Tricks – Aug 18, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Aug 18, 2023](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/99e6a836a21e272da945f0e045f4d4812c1d3583-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the CLI – FiftyOne Tips and Tricks – Aug 25th, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Aug 26, 2023](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-329-lllmstxt|> ## AI Meetup Recap [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Event Recaps](https://voxel51.com/blog/category/event-recaps) Recapping the AI, Machine Learning, and Data Science Meetup — Sept 7, 2023 Sep 9, 2023 • 4 min read Article content In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#c5cdfb358624) [Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#449d03fd731b) [EgoSchema: A Dataset for Truly Long-Form Video Understanding](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#5d7e3300fb56) [Monitoring Large Language Models (LLMs) in Production](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#4ce691159b49) [Join the AI, Machine Learning, and Data Science Meetup!](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#050e9fc3a6f2) [What’s Next?](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#6855f26ce355) [Get Involved!](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#80eada9b3115) In this article [First, Thanks for Voting for Your Favorite Charity!](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#c5cdfb358624) [Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#449d03fd731b) [EgoSchema: A Dataset for Truly Long-Form Video Understanding](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#5d7e3300fb56) [Monitoring Large Language Models (LLMs) in Production](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#4ce691159b49) [Join the AI, Machine Learning, and Data Science Meetup!](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#050e9fc3a6f2) [What’s Next?](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#6855f26ce355) [Get Involved!](https://voxel51.com/blog/recapping-the-ai-ml-data-science-meetup-sept-7-2023#80eada9b3115) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) We just wrapped up the Sept 7, 2023 , [AI, Machine Learning, and Data Science Meetup](https://www.meetup.com/pro/ai-machine-learning-data-science-network/) and if you missed it or want to revisit it, here’s a recap! In this blog post you’ll find the playback recordings, highlights from the presentations and Q&A, as well as the upcoming Meetup schedule so that you can join us at a future event. ## **First, Thanks for Voting for Your Favorite Charity!** In lieu of swag, we gave Meetup attendees the opportunity to help guide a $200 donation to charitable causes. The charity that received the highest number of votes this month was [Coalition for Rainforest Nations](https://www.rainforestcoalition.org/), an organization on a mission to save the World’s last great rainforests to achieve environmental and social sustainability. We are sending this event’s charitable donation of $200 to Coalition for Rainforest Nations on behalf of the computer vision community! Missed the Meetup? No problem. Here are playbacks and talk abstracts from the event. ## **Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos** https://youtu.be/DjqMDeLglfY?si=2FnRo0tIWqbGXP\_z The success of the Neural Radiance Fields (NeRFs) for modeling and free-view rendering static objects has inspired numerous attempts on dynamic scenes. Current techniques that utilize neural rendering for facilitating freeview videos (FVVs) are restricted to either offline rendering or are capable of processing only brief sequences with minimal motion. In this paper, we present a novel technique, Residual Radiance Field or ReRF, as a highly compact neural representation to achieve real-time FVV rendering on long-duration dynamic scenes. [Minye Wu](https://www.linkedin.com/in/minye-wu-777476205/?originalSubdomain=be) is a Postdoctoral researcher at KU Leuven. **Q&A** - Why is PCA linear encoder used? Why not a nonlinear one? **Resource links** - [arVix: Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos](https://arxiv.org/abs/2304.04452) ## **EgoSchema: A Dataset for Truly Long-Form Video Understanding** https://youtu.be/JAg-CJpk5es?si=QRfiKa-Un0te43Ec Introducing EgoSchema, a very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems. Derived from Ego4D, EgoSchema consists of over 5000 human curated multiple choice question answer pairs, spanning over 250 hours of real video data, covering a very broad range of natural human activity and behavior. [Karttikeya Mangalam](https://www.linkedin.com/in/karttikeya-mangalam-9248a6110/) is a PhD student in Computer Science at the Department of Electrical Engineering & Computer Sciences (EECS) at University of California, Berkeley advised by Prof. Jitendra Malik. Earlier, he held a visiting researcher position at Meta AI where he collaborated with Dr. Christoph Feichtenhofer and team. **Q&A** - Could you explain the difference between reconstruction and generation on your first couple of slides? - Can we detect face gestures? - Could the EgoSchema Generation Process be suitable with real estate video data? For example: "What is the curb appeal of this property?" - Why is it called EgoSchema? **Resource links** - [arVix: EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding](https://arxiv.org/abs/2308.09126) - [EgoSchema on GitHub](https://egoschema.github.io/) ## **Monitoring Large Language Models (LLMs) in Production** https://youtu.be/-xgB7pGgxqY?si=5-LkBeLZrPSNXl\_a Just like with all machine learning models, once you put an LLM in production you’ll probably want to keep an eye on how it’s performing. Observing key language metrics about user interaction and responses can help you craft better prompt templates and guardrails for your applications. This talk will take a look at what you might want to be looking at once you deploy your LLMs. [Sage Elliott](https://www.linkedin.com/in/sageelliott/) is a Technical Evangelist – Machine Learning & MLOps at WhyLabs. He enjoys breaking down the barrier to AI observability and talking to amazing people in the AI community. **Q&A** - Could explain how reading a score is constructed for understanding? - When and where is LangKit used in an ML pipeline for LLMS (in production)? - What is data sketching? - Is this applicable to non-conversational AI? **Resource links** - [LingKit GitHub repo](https://github.com/whylabs/langkit) - [whylogs GitHub repo](https://github.com/whylabs/whylogs) ## Join the AI, Machine Learning, and Data Science Meetup! The Meetup’s membership has grown to more than [10,000 members](https://www.meetup.com/pro/ai-machine-learning-data-science-network/)! The goal of the Meetups is to bring together communities of data scientists, machine learning engineers, and open source enthusiasts who want to share and expand their knowledge of AI, machine learning, data science and complementary technologies. Join one of the 12 Meetup locations closest to your timezone. - [Athens](https://www.meetup.com/athens-ai-machine-learning-data-science/) - [Austin](https://www.meetup.com/austin-ai-machine-learning-data-science/) - [Bangalore](https://www.meetup.com/bangalore-ai-machine-learning-data-science/) - [Boston](https://www.meetup.com/boston-ai-machine-learning-data-science/) - [Chicago](https://www.meetup.com/chicago-ai-machine-learning-data-science/) - [London](https://www.meetup.com/london-ai-machine-learning-data-science/) - [New York](https://www.meetup.com/new-york-ai-machine-learning-data-science/) - [Peninsula](https://www.meetup.com/peninsula-ai-machine-learning-data-science/) - [San Francisco](https://www.meetup.com/sf-ai-machine-learning-data-science/) - [Seattle](https://www.meetup.com/seattle-ai-machine-learning-data-science/) - [Silicon Valley](https://www.meetup.com/sv-ai-machine-learning-data-science/) - [Toronto](https://www.meetup.com/toronto-ai-machine-learning-data-science/) We have exciting speakers already signed up over the next few months! Become a member of the [AI, Machine Learning, and Data Science Meetup](https://www.meetup.com/pro/ai-machine-learning-data-science-network/) closest to you, then register for the Zoom. ## What’s Next? \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Up next on Sept 14 at 10 AM Pacific we have a great line up speakers including: - **ARMBench: An Object-Centric Benchmark Dataset for Robotic Manipulation** _\- Amazon Robotics team_ - **From Model to the Edge, Putting Your Model into Production** _\- Joy Timmermans, Secury360_ - **Optimizing Distributed Fine-Tuning Workloads for Stable Diffusion with the Intel Extension for PyTorch on AWS** _\- Eduardo Alvarez, Intel_ [Register for the Zoom here.](https://voxel51.com/computer-vision-events/september-14-meetup/) You can find a complete schedule of upcoming Meetups on [the Voxel51 Events page](https://voxel51.com/computer-vision-events/). ## **Get Involved!** There are a lot of ways to get involved in the Computer Vision Meetups. Reach out if you identify with any of these: - You’d like to speak at an upcoming Meetup - You have a physical meeting space in one of the Meetup locations and would like to make it available for a Meetup - You’d like to co-organize a Meetup - You’d like to co-sponsor a Meetup Reach out to Meetup co-organizer Jimmy Guerrero on Meetup.com or ping me over [LinkedIn](https://www.linkedin.com/in/jiguerrero/) to discuss how to get you plugged in. _The Computer Vision Meetup network is sponsored by [Voxel51](https://voxel51.com/), the company behind the open source [FiftyOne](https://github.com/voxel51/fiftyone) computer vision toolset. FiftyOne enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. It’s easy to [get started](https://voxel51.com/docs/fiftyone/index.html), in just a few minutes._ [AI Machine Learning Data Science meetup](https://voxel51.com/blog/tag/ai-machine-learning-data-science-meetup) [AI/ML/DS meetup](https://voxel51.com/blog/tag/ai-ml-ds-meetup) [events](https://voxel51.com/blog/tag/events) [meetup](https://voxel51.com/blog/tag/meetup) [meetups](https://voxel51.com/blog/tag/meetups) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [Recapping the Computer Vision Meetup — Sept 14, 2023\\ \\ Event Recaps\\ \\ • \\ \\ Sep 15, 2023](https://voxel51.com/blog/recapping-the-computer-vision-meetup-sept-14-2023) [Recapping the AI, Machine Learning and Data Science Meetup — Oct 5, 2023\\ \\ Event Recaps\\ \\ • \\ \\ Oct 6, 2023](https://voxel51.com/blog/recapping-the-ai-machine-learning-and-data-science-meetup-oct-5-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/af656fd506d11b555a019950b688830000b62f30-1200x676.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Recapping the Vector Search-Themed Computer Vision Meetup — July 13, 2023\\ \\ Event Recaps, Vector Search\\ \\ • \\ \\ Jul 17, 2023](https://voxel51.com/blog/recapping-the-vector-search-themed-computer-vision-meetup-july-13-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-330-lllmstxt|> ## FACET Benchmark Overview [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Datasets](https://voxel51.com/blog/category/datasets), [Tutorials](https://voxel51.com/blog/category/tutorials) FACET: A Benchmark Dataset for Fairness in Computer Vision Sep 12, 2023 • 11 min read Article content In this article [Everything you need to know!](https://voxel51.com/blog/facet-benchmark#9d1b65a8f619) [FACET Dataset Quick Facts](https://voxel51.com/blog/facet-benchmark#45f284bfc6ef) [Loading the FACET Dataset](https://voxel51.com/blog/facet-benchmark#9042b8eef2d0) [Evaluating Model Bias](https://voxel51.com/blog/facet-benchmark#f6dfac31b1bc) [Conclusion](https://voxel51.com/blog/facet-benchmark#b1b0f37f80cb) In this article [Everything you need to know!](https://voxel51.com/blog/facet-benchmark#9d1b65a8f619) [FACET Dataset Quick Facts](https://voxel51.com/blog/facet-benchmark#45f284bfc6ef) [Loading the FACET Dataset](https://voxel51.com/blog/facet-benchmark#9042b8eef2d0) [Evaluating Model Bias](https://voxel51.com/blog/facet-benchmark#f6dfac31b1bc) [Conclusion](https://voxel51.com/blog/facet-benchmark#b1b0f37f80cb) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Everything you need to know! \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop Computer vision has come a long way over the past few years. Between [single-stage detectors](https://www.datacamp.com/blog/yolo-object-detection-explained), [neural architecture search](https://github.com/Deci-AI/super-gradients/blob/master/YOLONAS.md), [vision transformers](https://arxiv.org/abs/2010.11929), and [foundation models](https://www.analyticsvidhya.com/blog/2023/05/foundation-models/), challenges that once seemed insurmountable are now standard fare. There are even high quality zero-shot models for detection (such as [Grounding DINO](https://github.com/IDEA-Research/GroundingDINO)) and segmentation ( [SAM](https://segment-anything.com/), [FastSAM](https://github.com/CASIA-IVA-Lab/FastSAM)), so for many applications, you can achieve a good baseline without any custom training! As computer vision models have rapidly improved, the problem of _bias_ has increasingly reared its head and become increasingly pronounced: even models that achieve state-of-the-art performance on one-number metrics like mean average precision or F1 score can vary wildly in their ability to generate predictions for people of different demographics, genders, and skin tones. If you’re curious to learn more about how models can learn human biases, check out [this paper](https://arxiv.org/pdf/2010.15052.pdf) (cited by the FACET team). In an effort to address these biases, a team at Meta has released [FACET](https://ai.meta.com/datasets/facet/) (FAirness in Computer Vision EvaluaTion), a new benchmark dataset for studying and evaluating the “fairness” in computer vision models. When developing FACET, the team set out to create the most comprehensive, diverse fairness benchmark dataset to date. This blog will give you everything you need to know to get started working with FACET so you can ensure that the models you deploy are not learning biases. - [FACET Dataset Quick Facts](https://voxel51.com/blog/facet-benchmark#quick-facts) - [Loading Ground Truth Annotations](https://voxel51.com/blog/facet-benchmark#loading-facet) - [Evaluating Bias in Your Models](https://voxel51.com/blog/facet-benchmark#evaluating-bias) ## FACET Dataset Quick Facts [Paper](https://ai.meta.com/research/publications/facet-fairness-in-computer-vision-evaluation-benchmark/) - **Title**: _FACET: Fairness in Computer Vision Evaluation Benchmark_ — ICCV 2023 - **Authors**: Laura Gustafson, Chloe Rolland, Nikhila Ravi, Quentin Duval, Aaron Adcock, Cheng-Yang Fu, Melissa Hall, Candace Ross \| Meta AI [Dataset](https://ai.meta.com/datasets/facet/) - [License](https://ai.meta.com/datasets/facet-downloads/) - [Data Card](https://scontent-iad3-2.xx.fbcdn.net/v/t39.8562-6/373310495_1501617103918250_5940202642610192897_n.pdf?_nc_cat=106&ccb=1-7&_nc_sid=ae5e01&_nc_ohc=nPIFsvqTpqUAX-6AL_Z&_nc_ht=scontent-iad3-2.xx&oh=00_AfByL8sz7YHi_rCChEc2hVzfcLgQAnRw0LafQsl69k3-1A&oe=64F70C95) - ⚠️Only to be used for evaluation purposes, NOT for training purposes ### **Annotations** - `annotations/annotations.csv`: person detections with bounding boxes and attributes (see data card for details) - `annotations/coco_boxes.json`: MS COCO formatted detection bounding boxes for people, without attributes - `annotations/coco_masks.json`: MS COCO formatted segmentation masks for people, clothing, and hair — encoded with [Run-Length Encoding](https://en.wikipedia.org/wiki/Run-length_encoding#:~:text=Run%2Dlength%20encoding%20%28RLE%29,than%20as%20the%20original%20run.) (RLE) — and bounding boxes ### Dataset Statistics - 31,702 images (a subset of the [SA-1B dataset](https://ai.meta.com/datasets/segment-anything/)) - 49,551 unique person detections, spanning protected attributes, unprotected attributes, lighting conditions, and more. - 69,105 instance segmentation masks (person, clothing, and hair) - 52 person-related classes (all person detections have a primary; some also have a secondary class) - Perceived skin tone label interpretation: 1–10 on [Monk Skin Tone scale](https://skintone.google/get-started) ### Efforts Taken by FACET Team to Ensure Fairness - _Expert annotators_: the team hired experts from multiple geographic regions, from the Americas to Southeast Asia. Annotators also completed “stage-specific training” before they could begin labeling. - _Perceived labels_: protected attributes (age, gender presentation, and skin tone) are labeled as “perceived” to reflect the limitations of image annotation. For instance for age, the team writes: “it is impossible to tell a person’s true age from an image, these numerical ranges are a rough guideline to delineate each perceived age group”. - _Diverse scenarios_: from different lighting conditions and degrees of occlusion, to varied numbers of people in the image and an array of accessories, the FACET team tried to capture a broad range of visual scenarios. In addition to evaluating how model performance depends on a single demographic variable, this allows for evaluations of _intersectionality_. - _Image filtering_: to avoid inheriting biases from existing detection and classification models, the FACET team employed human annotators to filter out images that did not contain a person matching one of the 52 label categories (singer, painter, astronaut, …). This filtered out about 80% of the initial images. - _Class disambiguation_: sometimes a person fits the description of more than one label class. As the FACET authors illustrate, “a person playing the guitar and singing can match the category labels guitarist and singer”. To account for this, a secondary class label is given to detected persons when appropriate. - _Skin tone aggregation_: because the tone of one’s skin can influence how they perceive others’ skin tones, FACET aggregates perceived skin tone values across multiple annotators for each label. ### Value Counts by Attribute - Lighting Condition ```raw 1well lit: 35533 2underexposed: 1313 3overexposed: 553 4dimly lit: 10955 5None/na: 1197 ``` - Hair Type ```raw 1straight: 18382 2curly: 719 3bald: 1017 4wavy: 6141 5dreadlocks: 280 6coily: 458 7None/na: 22554 ``` - Hair Color ```raw 1black: 14041 2blonde: 2249 3red: 333 4colored: 248 5brown: 10668 6grey: 2107 7None/na: 19905 ``` - Has Facial Hair ```raw 1True: 6121 2False: 43430 ``` - Perceived Age Presentation ```raw 1young (25-40): 8860 2middle (41-65): 27380 3older (65+): 2659 4None: 10652 ``` - Perceived Gender Presentation ```raw 1fem: 10245 2masc: 33240 3non binary: 95 4None/na: 5971 ``` - Tattoo ```raw 1False: 48846 2True: 705 ``` Primary Class (in alphabetical order) ```raw 1astronaut': 286, 2'backpacker': 1612, 3'ballplayer': 1309, 4'bartender': 56, 5'basketball_player': 1668, 6'boatman': 2048, 7'carpenter': 223, 8'cheerleader': 399, 9'climber': 455, 10'computer_user': 1164, 11'craftsman': 1034, 12'dancer': 1397, 13'disk_jockey': 310, 14'doctor': 802, 15'drummer': 977, 16'electrician': 468, 17'farmer': 1542, 18'fireman': 913, 19'flutist': 302, 20'gardener': 457, 21'guard': 1361, 22'guitarist': 1180, 23'gymnast': 615, 24'hairdresser': 458, 25'horseman': 735, 26'judge': 96, 27'laborer': 2540, 28'lawman': 4455, 29'lifeguard': 511, 30'machinist': 354, 31'motorcyclist': 1367, 32'nurse': 1042, 33'painter': 898, 34'patient': 884, 35'prayer': 798, 36'referee': 755, 37'repairman': 1295, 38'reporter': 470, 39'retailer': 546, 40'runner': 638, 41'sculptor': 213, 42'seller': 1178, 43'singer': 1286, 44'skateboarder': 990, 45'soccer_player': 1226, 46'soldier': 1457, 47'speaker': 1416, 48'student': 682, 49'teacher': 192, 50'tennis_player': 1661, 51'trumpeter': 498, 52'waiter': 332 ``` ## Loading the FACET Dataset ### Prerequisites Before you download the dataset, you need to sign Meta’s FACET usage agreement [here](https://ai.meta.com/datasets/facet-downloads/). After doing so, unzip the four zip files ( `annotations`, `imgs_1`, `imgs_2`, and `imgs_3`). We will be using the open source computer vision library [FiftyOne](https://github.com/voxel51/fiftyone) for data management and visualization, so if you have not done so already, install FiftyOne: ```bash 1pip install fiftyone ``` In Python, import the needed libraries: ```python 1import json 2import numpy as np 3import os 4import pandas as pd 5from PIL import Image 6from pycocotools import mask as maskUtils 7from tqdm.notebook import tqdm ``` As well as the FiftyOne modules we will be utilizing: ```python 1import fiftyone as fo 2import fiftyone.brain as fob 3import fiftyone.zoo as foz 4from fiftyone import ViewField as F ``` ### Creating the Dataset Now we are ready to create the dataset. We will create an empty dataset (and persist it to database), and then add all of the images in each of the three image folders we unzipped: ```python 1## use relative paths to your image dirs 2IMG_DIRS = ["imgs_1", "imgs_2", "imgs_3"] 3 4dataset = fo.Dataset(name = "FACET", persistent=True) 5 6for img_dir in IMG_DIRS: 7dataset.add_images_dir(img_dir) 8 9dataset.compute_metadata() ``` We also compute the “metadata” so that we can utilize sample width and height when converting between relative and absolute bounding box conventions. We can print out the dataset and get some quick facts: ```python 1print(dataset) ``` ```raw 1Name: FACET 2Media type: image 3Num samples: 31702 4Persistent: True 5Tags: [] 6Sample fields: 7id: fiftyone.core.fields.ObjectIdField 8filepath: fiftyone.core.fields.StringField 9tags: fiftyone.core.fields.ListField(fiftyone.core.fields.StringField) 10metadata: fiftyone.core.fields.EmbeddedDocumentField(fiftyone.core.metadata.ImageMetadata) ``` And we can launch a session of the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html) to view the images: ```python 1session = fo.launch_app(dataset) ``` \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ### Adding Person Detections First, we load the ground truth annotations from CSV into a pandas DataFrame: ```python 1gt_df = pd.read_csv('annotations/annotations.csv') ``` Now we need to parse this DataFrame and add the appropriately structured data to our dataset. Like the authors of FACET do, let’s break things down into three types of attributes: 1. Person attributes ( `hairtype`, `has_eyeware`, etc.) 2. Protected attributes (perceived gender presentation, perceived age presentation, and perceived skin tone) 3. Other attributes (lighting, visibility) For each row in the dataframe, we will create a FiftyOne `Detection`, and we will add the relevant attributes to the detection. **Person attributes** We can create a tuple for the Boolean attributes we will iterate over: ```python 1BOOLEAN_PERSONAL_ATTRS = ( 2 "has_facial_hair", 3 "has_tattoo", 4 "has_cap", 5 "has_mask", 6 "has_headscarf", 7 "has_eyeware", 8) 9 10def add_boolean_person_attributes(detection, row_index): 11 for attr in BOOLEAN_PERSONAL_ATTRS: 12 detection[attr] = gt_df.loc[row_index, attr].astype(bool) ``` We will also create some simple helper functions to restructure the hair type and color information: ```python 1def get_hairtype(row_index): 2 hair_info = gt_df.loc[row_index, gt_df.columns.str.startswith('hairtype')] 3 hairtype = hair_info[hair_info == 1] 4 if len(hairtype) == 0: 5 return None 6 return hairtype.index[0].split('_')[1] 7 8def get_haircolor(row_index): 9 hair_info = gt_df.loc[row_index, gt_df.columns.str.startswith('hair_color')] 10 haircolor = hair_info[hair_info == 1] 11 if len(haircolor) == 0: 12 return None 13 return haircolor.index[0].split('_')[2] ``` All of this can be combined into a single function to add these person attributes: ```python 1def add_person_attributes(detection, row_index): 2 detection["hairtype"] = get_hairtype(row_index) 3 detection["haircolor"] = get_haircolor(row_index) 4 add_boolean_person_attributes(detection, row_index) ``` **Protected attributes** For perceived gender and perceived age, we will return whichever column with the associated prefix (“gender” or “age”) has a value of 1 in the given row. Aside from that, we just need to do a little string formatting. ```python 1def get_perceived_gender_presentation(row_index): 2 gender_info = gt_df.loc[row_index, gt_df.columns.str.startswith('gender')] 3 pgp = gender_info[gender_info == 1] 4 if len(pgp) == 0: 5 return None 6 return pgp.index[0].replace("gender_presentation_", "").replace("_", " ") 7 8def get_perceived_age_presentation(row_index): 9 age_info = gt_df.loc[row_index, gt_df.columns.str.startswith('age')] 10 pap = age_info[age_info == 1] 11 if len(pap) == 0: 12 return None 13 return pap.index[0].split('_')[2] ``` For skin tone, a single detection might have multiple skin tones with nonzero values. To capture all of this information, we will just convert the skin tone columns into a dictionary and store the dictionary under a `skin_tone` attribute: ```python 1def get_skintone(row_index): 2 skin_info = gt_df.loc[row_index, gt_df.columns.str.startswith('skin_tone')] 3 return skin_info.to_dict() ``` Altogether, adding protected attributes to the detection is done in this function: ```python 1def add_protected_attributes(detection, row_index): 2 detection["perceived_age_presentation"] = get_perceived_age_presentation(row_index) 3 detection["perceived_gender_presentation"] = get_perceived_gender_presentation(row_index) 4 detection["skin_tone"] = get_skintone(row_index) ``` **Other attributes** As with the Boolean person attributes, we can create a tuple of visibility attributes to iterate over: ```python 1VISIBILITY_ATTRS = ("visible_torso", "visible_face", "visible_minimal") ``` Other than that, we just need to process the lighting information: ```python 1def get_lighting(row_index): 2 lighting_info = gt_df.loc[row_index, gt_df.columns.str.startswith('lighting')] 3 lighting = lighting_info[lighting_info == 1] 4 if len(lighting) == 0: 5 return None 6 lighting = lighting.index[0].replace("lighting_", "").replace("_", " ") 7 return lighting 8 9def add_other_attributes(detection, row_index): 10 detection["lighting"] = get_lighting(row_index) 11 for attr in VISIBILITY_ATTRS: 12 detection[attr] = gt_df.loc[row_index, attr].astype(bool) ``` Now we have all the pieces we need to create a `Detection`, given a row index from the pandas DataFrame. We will also pass the `sample` corresponding to that row in so that we can convert between absolute and relative bounding box coordinates: ```python 1def create_detection(row_index, sample): 2 bbox_dict = json.loads(gt_df.loc[row_index, "bounding_box"]) 3 x, y, w, h = bbox_dict["x"], bbox_dict["y"], bbox_dict["width"], bbox_dict["height"] 4 cat1, cat2 = bbox_dict["dict_attributes"]["cat1"], bbox_dict["dict_attributes"]["cat2"] 5 6 person_id = gt_df.loc[row_index, "person_id"] 7 8 img_width, img_height = sample.metadata.width, sample.metadata.height 9 10 bounding_box = [x/img_width, y/img_height, w/img_width, h/img_height] 11 detection = fo.Detection( 12 label=cat1, 13 bounding_box=bounding_box, 14 person_id=person_id, 15 ) 16 if cat2 != 'none': 17 detection["class2"] = cat2 18 19 add_person_attributes(detection, row_index) 20 add_protected_attributes(detection, row_index) 21 add_other_attributes(detection, row_index) 22 23 return detection ``` All that is left is to iterate over the samples in our dataset (this is more efficient than iterating over the rows in the `gt_df` DataFrame and filtering the dataset for the right sample), adding detections to each sample as we go: ```python 1def add_ground_truth_labels(dataset): 2 for sample in dataset.iter_samples(autosave=True, progress=True): 3 sample_annos = gt_df[gt_df['filename'] == sample.filename] 4 detections = [] 5 for row in sample_annos.iterrows(): 6 row_index = row[0] 7 detection = create_detection(row_index, sample) 8 detections.append(detection) 9 sample["ground_truth"] = fo.Detections(detections=detections) 10 dataset.add_dynamic_sample_fields() 11 12## add all of the ground truth labels 13add_ground_truth_labels(dataset) ``` \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ### Adding Segmentation Masks While the person, hair, and clothing segmentation masks are not strictly necessary for the detection and classification evaluation routines in the next section, it’s worth demonstrating how to add them, as the person masks can be used to evaluate instance segmentation models! The `add_coco_masks_to_dataset()` function below does the following: - Iterate through samples in the dataset - For each sample whose filename has an entry in the COCO masks annotation file, we extract the segmentation mask and decode it from RLE format into a binary full image mask using `pycocoutils`. - Using `pillow` and the bounding box associated with the segmentation mask, we crop the mask, turn it back into an array, and add the array as a mask (with bounding box) to a new `Detection`. ```python 1def add_coco_masks_to_dataset(dataset): 2 coco_masks = json.load(open("annotations/coco_masks.json", "r")) 3 cmas = coco_masks["annotations"] 4 5 FILENAME_TO_ID = { 6 img["file_name"]: img["id"] 7 for img in coco_masks["images"] 8 } 9 10 CAT_TO_LABEL = {cat["id"]: cat["name"] for cat in coco_masks["categories"]} 11 12 for sample in dataset.iter_samples(autosave=True, progress=True): 13 fn = sample.filename 14 15 if fn not in FILENAME_TO_ID: 16 continue 17 18 img_id = FILENAME_TO_ID[fn] 19 img_width, img_height = sample.metadata.width, sample.metadata.height 20 sample_annos = [a for a in cmas if a["image_id"] == img_id] 21 if len(sample_annos) == 0: 22 continue 23 24 coco_detections = [] 25 for ann in sample_annos: 26 label = CAT_TO_LABEL[ann["category_id"]] 27 bbox = ann['bbox'] 28 ann_id = ann['ann_id'] 29 person_id = ann['facet_person_id'] 30 31 mask = maskUtils.decode(ann["segmentation"]) 32 mask = Image.fromarray(255*mask) 33 34 ## Change bbox to be in the format [x, y, x, y] 35 bbox[2] = bbox[0] + bbox[2] 36 bbox[3] = bbox[1] + bbox[3] 37 38 ## Get the cropped image 39 cropped_mask = np.array(mask.crop(bbox)).astype(bool) 40 41 ## Convert to relative [x, y, w, h] coordinates 42 bbox[2] = bbox[2] - bbox[0] 43 bbox[3] = bbox[3] - bbox[1] 44 45 bbox[0] = bbox[0]/img_width 46 bbox[1] = bbox[1]/img_height 47 bbox[2] = bbox[2]/img_width 48 bbox[3] = bbox[3]/img_height 49 50 new_detection = fo.Detection( 51 label=label, 52 bounding_box=bbox, 53 person_id=person_id, 54 ann_id=ann_id, 55 mask=cropped_mask, 56 ) 57 coco_detections.append(new_detection) 58 sample["coco_masks"] = fo.Detections(detections=coco_detections) 59 60## add the masks 61add_coco_masks_to_dataset(dataset) ``` \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop ## Evaluating Model Bias ### FACET Evaluation Metrics To evaluate your model biases using FACET, the dataset’s authors suggest looking at the recall of the model on subsets of the data defined by a _concept_ and an _attribute_. For example, the concept could be a class label such as “backpacker”, and the attribute can be the value “curly” for hairstyle. Attributes can also be combined to explore prediction quality with intersectionality. The model’s _disparity_ in performance between two sets of attributes is the difference between its recall scores on these subsets. \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop When evaluating classification models (where classification is performed on the ground truth patches), the standard definition of recall is used: \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop This can be interpreted as the rate at which true positives are correctly identified. When evaluating detection models, only predictions with the `person` label are considered. Predictions are deemed correct when the overlap between ground truth bounding box and prediction bounding box — formally known as the [Intersection over Union](https://pyimagesearch.com/2016/11/07/intersection-over-union-iou-for-object-detection/) (IoU) — is greater than some threshold. The recall which is plugged into the disparity equation is the average recall over a sequence of IoU thresholds \[0.5, 0.55, …, 0.95, 1.0\], and is known as the mean average recall (mAR). ### Adding Predictions to the Dataset \[@portabletext/react\] Unknown block type "externalImage", specify a component for it in the \`components.types\` prop So that we can see these evaluation routines in action, let’s generate some predictions. For detection, we’ll load YOLOv5m trained on COCO from the [FiftyOne Model Zoo](https://docs.voxel51.com/user_guide/model_zoo/): ```python 1yolov5 = foz.load_zoo_model('yolov5m-coco-torch') ``` We can then apply the model to our dataset and store the predictions in a field on our samples with: ```python 1dataset.apply_model(yolov5, label_field="yolov5m") 2 3### Just retain the "person" detections 4people_view_values = dataset.filter_labels("yolov5m", F("label") == "person").values("yolov5m") 5dataset.set_values("yolov5m", people_view_values) 6dataset.save() ``` For zero-shot classification, as in the FACET paper, we will use a CLIP model with custom classes. We will again get the model from the FiftyOne Model Zoo: ```python 1## get a list of all 52 classes 2facet_classes = dataset.distinct("ground_truth.detections.label") 3 4## instantiate a CLIP model with these classes 5clip = foz.load_zoo_model( 6 "clip-vit-base32-torch", 7 text_prompt="A photo of a", 8 classes=facet_classes, 9) ``` To generate classification predictions (effectively treating the `ground_truth` bounding box regions as their own images) we will use FiftyOne’s `to_patches()` method to create a view of all of the ground truth patches in our dataset, and then apply the CLIP model to these patches: ```python 1patch_view = dataset.to_patches("ground_truth") 2patch_view.apply_model(clip, label_field="clip") 3dataset.save_view("patch_view", patch_view) ``` The last line in the code block above [saves this view](https://docs.voxel51.com/user_guide/using_views.html#saving-views) to the dataset. ### Evaluating Detection Predictions Now that we have some detection and classification predictions, let’s write some code to compute the recall (or mean average recall) for a subset of the data. We will then illustrate how to filter for the subset corresponding to a specific concept and set of attributes. For detections, we define the IoU thresholds we are going to average over: ```python 1IOU_THRESHS = np.round(np.arange(0.5, 1.0, 0.05), 2) ``` The following function only needs to be run once for each detection model: ```python 1def _evaluate_detection_model(dataset, label_field): 2 eval_key = "eval_" + label_field.replace("-", "_") 3 dataset.evaluate_detections(label_field, "ground_truth", eval_key=eval_key, classwise=False) 4 5 for sample in dataset.iter_samples(autosave=True, progress=True): 6 for pred in sample[label_field].detections: 7 iou_field = f"{eval_key}_iou" 8 if iou_field not in pred: 9 continue 10 11 iou = pred[iou_field] 12 for it in IOU_THRESHS: 13 pred[f"{iou_field}_{str(it).replace('.', '')}"] = iou >= it ``` This leverages FiftyOne’s built-in `evaluate_detections()` method to compute the IoU of each prediction bounding box that has overlap with a `ground_truth` bounding box. We pass in `classwise=False` because the labels for the ground truth detections are the 52 classes in the FACET dataset, whereas the label for our predictions is \`person\`. Which set of predictions to evaluate is specified by `label_field`. After adding the IoU information to the predictions, we iterate over predictions, and store a Boolean for each IoU threshold, denoting whether or not the prediction would be considered a true positive at that threshold. For any subset of the dataset — associated with a concept, attributes, or something else — we can compute the mean average recall as follows: ```python 1def _compute_detection_mAR(sample_collection, label_field): 2 """Computes the mean average recall of the specified detection field. 3 -- computed as the average over iou thresholds of the recall at 4 each threshold. 5 """ 6 eval_key = "eval_" + label_field.replace("-", "_") 7 iou_recalls = [] 8 for it in IOU_THRESHS: 9 field_str = f"{label_field}.detections.{eval_key}_iou_{str(it).replace('.', '')}" 10 counts = sample_collection.count_values(field_str) 11 tp, fn = counts.get(True, 0), counts.get(False, 0) 12 recall = tp/float(tp + fn) if tp + fn > 0 else 0.0 13 iou_recalls.append(recall) 14 15 return np.mean(iou_recalls) ``` To close the loop with the FACET dataset, we can define a function, `get_concept_attr_detection_mAR()`, which takes in a concept (primary category for a person) and a dictionary of attributes in `{field:value}` form, and returns a mean average recall value for this combo: ```python 1def get_concept_attr_detection_mAR(dataset, label_field, concept, attributes): 2 sub_view = dataset.filter_labels("ground_truth", F("label") == concept) 3 for attribute in attributes.items(): 4 if "skin_tone" in attribute[0]: 5 sub_view = sub_view.filter_labels("ground_truth", F(f"skin_tone.{attribute[0]}") != 0) 6 else: 7 sub_view = sub_view.filter_labels(f"ground_truth", F(attribute[0]) == attribute[1]) 8 return _compute_detection_mAR(sub_view, label_field) ``` As an example, let’s see YOLOv5m’s mean average recall for gymnasts with curly, black hair: ```python 1concept = 'gymnast' 2attributes = {"hairtype": "curly", "haircolor": "black"} 3get_concept_attr_detection_mAR(dataset, "yolov5m", concept, attributes) 4## 0.875 ``` This detection evaluation routine can be adapted into an instance segmentation evaluation routine by filtering the COCO masks down to just the person masks, and passing `use_masks=True` into `evaluate_detections()`. ### Evaluating Classification Predictions In analog with detections, for classification models, we can create a `_evaluate_classification_model()` function which only needs to be run once per model: ```python 1def _evaluate_classification_model(dataset, prediction_field): 2 patch_view = dataset.load_saved_view("patch_view") 3 eval_key = "eval_" + prediction_field 4 5 for sample in patch_view.iter_samples(progress=True): 6 sample[eval_key] = ( 7 sample.ground_truth.label == sample[prediction_field].label 8 ) 9 sample.save() 10 dataset.save_view("patch_view", patch_view, overwrite=True) ``` This stores a True/False result on each sample in the patches view and saves this to the dataset. Continuing the analogy with detections, we can compute classification recall for any collection of patches in the dataset: ```python 1def _compute_classification_recall(patch_collection, label_field): 2 eval_key = "eval_" + label_field.split("_")[0] 3 counts = patch_collection.count_values(eval_key) 4 tp, fn = counts.get(True, 0), counts.get(False, 0) 5 recall = tp/float(tp + fn) if tp + fn > 0 else 0.0 6 return recall ``` And we can bring this back to subsets of our data described by concepts and attributes: ```python 1def get_concept_attr_classification_recall(dataset, label_field, concept, attributes): 2 patch_view = dataset.load_saved_view("patch_view") 3 sub_patch_view = patch_view.match(F("ground_truth.label") == concept) 4 for attribute in attributes.items(): 5 if "skin_tone" in attribute[0]: 6 sub_patch_view = sub_patch_view.match(F(f"ground_truth.skin_tone.{attribute[0]}") != 0) 7 else: 8 sub_patch_view = sub_patch_view.match(F(f"ground_truth.{attribute[0]}") == attribute[1]) 9 return _compute_classification_recall(sub_patch_view, label_field) ``` For the same concept-attributes combination from above, our CLIP model gives this: ```python 1get_concept_attr_classification_recall(dataset, "clip", concept, attribute) 2## 0.6193353474320241 ``` ### Assessing Disparity For both detection and classification models, we can now compute the disparity between different sets of attributes for a shared concept. We’ll create a `get_concept_attr_recall()` function, which will use the appropriate definition of recall depending on whether our label field houses classification or detection predictions: ```python 1def get_concept_attr_recall(dataset, label_field, concept, attribute): 2 if label_field in dataset.get_field_schema().keys(): 3 return get_concept_attr_detection_mAR(dataset, label_field, concept, attribute) 4 else: 5 return get_concept_attr_classification_recall(dataset, label_field, concept, attribute) ``` And we will tie everything up in a nice bow with a `compute_disparity()` function: ```python 1def compute_disparity(dataset, label_field, concept, attribute1, attribute2): 2 recall1 = get_concept_attr_recall(dataset, label_field, concept, attribute1) 3 recall2 = get_concept_attr_recall(dataset, label_field, concept, attribute2) 4 return recall1 - recall2 ``` As an example, let’s look at the difference in CLIP model performance for some concepts when hair type is straight versus curly: ```python 1attrs1 = {"hairtype": "curly"} 2attrs2 = {"hairtype": "straight"} 3 4for concept in ["astronaut", "singer", "judge", "student"]: 5 disparity = compute_disparity(dataset, "clip", concept, attrs1, attrs2) 6 print(f"{concept}: {disparity}") 7 8#### OUTPUT #### 9## astronaut: -0.8269230769230769 10## singer: -0.0008051529790660261 11## judge: -0.06666666666666667 12## student: 0.16279069767441856 ``` In the output, a value closer to +1 means that the model has higher recall for curly hair than straight hair, and a value closer to -1 means the model has a higher recall for straight hair. Whereas for `singer`, CLIP achieves roughly the same recall for people with straight hair and people with curly hair, there is a stark difference in performance for `astronaut` — CLIP has much higher recall for straight-haired astronauts than curly-haired astronauts. You can also use FACET to compare models, combinations of attributes, and so much more! ## Conclusion There’s so much more to making a robust computer vision model than high mean average precision or F1 scores. Absent the appropriate caution and care, a model in production can exacerbate societal biases and even put people in physical danger. The key to building a great model, in computer vision and machine learning more broadly, is to make the model truly serve all of the people who are impacted by the model’s predictions. The way to get there is diverse, high quality data, and rigorous evaluation. FACET makes it easier than ever to take steps to ensure that your models are equitable. There’s much more work to be done in order to make models fair to all, and we all need to play an active part in building an equitable future! [benchmark](https://voxel51.com/blog/tag/benchmark) [bias](https://voxel51.com/blog/tag/bias) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [datasets](https://voxel51.com/blog/tag/datasets) [disparity](https://voxel51.com/blog/tag/disparity) [Evaluation](https://voxel51.com/blog/tag/evaluation) [fairness](https://voxel51.com/blog/tag/fairness) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Meta](https://voxel51.com/blog/tag/meta) [open source](https://voxel51.com/blog/tag/open-source) [SA1B](https://voxel51.com/blog/tag/sa1b) [segment anything](https://voxel51.com/blog/tag/segment-anything) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/713e4352b25d3b4ee12eab92246ceff22f808471-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Spending My First Week With FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ Aug 21, 2023](https://voxel51.com/blog/spending-my-first-week-with-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/b35a0ce49c8623d0874d793c3712f62e57aba87a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How Shap-E Changed How We Think About Diffusion Models\\ \\ Tutorials\\ \\ • \\ \\ May 30, 2024](https://voxel51.com/blog/how-shap-e-changed-how-we-think-about-diffusion-models) [CaFFe: Calving Fronts and Where to Find Them\\ \\ Computer Vision, Datasets\\ \\ • \\ \\ Sep 18, 2023](https://voxel51.com/blog/caffe-computer-vision-glacial-mass-modeling) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-331-lllmstxt|> ## Eliminate Image Duplicates [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Plugins](https://voxel51.com/blog/category/plugins), [Tutorials](https://voxel51.com/blog/category/tutorials) Double Trouble: Eliminate Image Duplicates with FiftyOne Sep 14, 2023 • 6 min read Article content In this article [Find Exact and Approximate Duplicate Images with This Plugin](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#8dae268350c6) [Image Deduplication 🖼️🪞🧹](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#2cf81724e2b7) [Plugin Overview & Functionality](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#0ef865fe669c) [Installing the Plugin](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#02281dc0e076) [Lessons Learned](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#88d89c2d9715) [Conclusion](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#1109bdb5654e) In this article [Find Exact and Approximate Duplicate Images with This Plugin](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#8dae268350c6) [Image Deduplication 🖼️🪞🧹](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#2cf81724e2b7) [Plugin Overview & Functionality](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#0ef865fe669c) [Installing the Plugin](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#02281dc0e076) [Lessons Learned](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#88d89c2d9715) [Conclusion](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone#1109bdb5654e) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Find Exact and Approximate Duplicate Images with This Plugin Welcome to week four of _Ten Weeks of Plugins_. During these ten weeks, we will be building a FiftyOne Plugin (or multiple!) each week and sharing the lessons learned! If you’re new to them, FiftyOne Plugins provide a flexible mechanism for anyone to extend the functionality of their FiftyOne App. You may find the following resources helpful: - [FiftyOne Plugins Repo](https://github.com/voxel51/fiftyone-plugins) - [FiftyOne Plugin Docs](https://docs.voxel51.com/plugins/index.html#downloading-plugins) - Plugins Channel in the [FiftyOne Community Slack](https://slack.voxel51.com/) What we’ve built so far: - Week 0: [Image Quality Issues](https://github.com/jacobmarks/image-quality-issues) & [Concept Interpolation](https://github.com/jacobmarks/concept-interpolation) - Week 1: [AI Art Gallery](https://github.com/jacobmarks/ai-art-gallery) & [Twilio Automation](https://github.com/jacobmarks/twilio-automation-plugin) - Week 2: [Visual Question Answering](https://github.com/jacobmarks/vqa-plugin) - Week 3: [YouTube Player Panel](https://github.com/jacobmarks/fiftyone-youtube-panel-plugin) Ok, let’s dive into this week’s FiftyOne Plugin — [Image Deduplication](https://github.com/jacobmarks/image-deduplication-plugin)! ## Image Deduplication 🖼️🪞🧹 ![](https://cdn.sanity.io/images/h6toihm1/production/e616860e60f2f6e5f80ca2d052ec2b3e867b6522-2241x1190.gif?auto=format&dpr=2&fit=max&q=75&w=1600) The biggest challenge in training machine learning models is curating a high quality dataset. Duplicate (or very similar) data is a major roadblock to building such a dataset. Multiple copies of the same (or approximately the same) samples can lead to longer training times, higher training costs, and lower overall performance. On the flip side, you likely want a diverse dataset with good coverage over the data domain. Duplicates come in two flavors: 1. **Exact duplicates**: pixel-perfect matches, where one image is literally a down-to-the-bit copy of another 2. **Approximate duplicates**: images (or other data) that are highly similar — typically evaluated by computing the closeness between samples with some similarity metric — and setting a threshold for similarity using this metric. _Deduplication_ is the task of removing these exact and approximate duplicates from a dataset. Typically, deduplication involves writing a lot of code to find, visualize, and remove all of the duplicates in your dataset. With this FiftyOne plugin, that all changes. Now you can deduplicate your entire dataset from within the FiftyOne App, without writing a single line of code! ## Plugin Overview & Functionality For the fourth week of _10 Weeks of Plugins_, I built an Image Deduplication Plugin. This plugin allows you to: - Find both exact and approximate duplicate images in your dataset - Visualize these groups of duplicates - Delete all duplicates OR Keep a representative from each set of duplicates The plugin has eight (!) [operators](https://docs.voxel51.com/plugins/index.html#operators) (a powerful feature in FiftyOne that allow plugin developers to define custom operations that can be executed by users of the FiftyOne App), but don’t get overwhelmed — really it’s just two sets of analogous operators, for exact and approximate deduplication workflows. After you install the plugin, when you open the operators list (pressing “ `` ` ``” in the FiftyOne App) you should see these operators. Search for “dedup” to narrow down the list! ### 🔍 Finding Duplicates The first pair of operators helps you to find duplicate images in your dataset. - `find_approximate_duplicate_images`: uses a [similarity index](https://docs.voxel51.com/user_guide/brain.html#similarity) to find approximate duplicates. ![](https://cdn.sanity.io/images/h6toihm1/production/15cc6378d5ab6717a714627bce21a47124bcc611-2560x1363.gif?auto=format&dpr=2&fit=max&q=75&w=1600) You can specify either a distance threshold (how close the images need to be according to the similarity metric to be considered near duplicates) or a fraction of the dataset to mark as near duplicates. If you haven’t computed a similarity index on your dataset, you can do so by running: ```python 1import fiftyone.brain as fob 2fob.compute_similarity(dataset, brain_key = "sim", metric="cosine") ``` For a large dataset, you may want to use a vector database. In this case, check out our native integrations with [Pinecone](https://docs.voxel51.com/integrations/pinecone.html), [Qdrant](https://docs.voxel51.com/integrations/qdrant.html), [Milvus](https://docs.voxel51.com/integrations/milvus.html), and [LanceDB](https://docs.voxel51.com/integrations/lancedb.html)! When the operation finishes, it will have created two [saved views](https://docs.voxel51.com/user_guide/using_views.html#saving-views): `approx_dup_view`, and `approx_dup_groups_view`. You can access these by clicking on the saved views selector in the FiftyOne App, or programmatically via Python: ```python 1approx_dup_view = dataset.load_saved_view("approx_dup_view") 2approx_dup_groups_view = dataset.load_saved_view("approx_dup_groups_view") ``` - `find_exact_duplicates`: uses [file hashes](https://www.sentinelone.com/cybersecurity-101/hashing/) to find exact duplicates ![](https://cdn.sanity.io/images/h6toihm1/production/3aed9bae8fb6139106fdfdc9428824d8ccedd18e-2560x1356.gif?auto=format&dpr=2&fit=max&q=75&w=1600) Essentially, the file hash computes a short signature for each sample based on the binary data stored in the image. The operator then checks if there are duplicate values of these signatures and marks these samples as duplicates. The operator adds a `filehash` field to each sample, and creates a saved view `exact_dup_view`, which contains just the images with duplicate filehashes. ### 🪟Viewing Duplicates Once you have found exact and/or approximate duplicates in your dataset, you may want to view these duplicates. For approximate duplicates, for instance, you may want to verify that the distance threshold you set was rigorous enough. The Image Deduplication plugin makes it easy to do this with the `display_approximate_duplicate_groups` and `display_exact_duplicate_groups` operators. The names are pretty self-explanatory, but the former loads the `approx_dup_groups_view` view we saved earlier, and the latter displays the samples in `exact_dup_view`, grouped by `filehash`. ![](https://cdn.sanity.io/images/h6toihm1/production/485042ace0605ad31c51857c784fc36dea5d88d7-2754x1462.gif?auto=format&dpr=2&fit=max&q=75&w=1600)![](https://cdn.sanity.io/images/h6toihm1/production/e616860e60f2f6e5f80ca2d052ec2b3e867b6522-2241x1190.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ### 🗑️Removing Duplicates Once you have viewed your identified duplicates, it is time to clean your dataset. At this point, you have two options: 1. Remove ALL duplicates: delete all samples marked as an exact or approximate duplicate 2. Keep a representative: remove all but one duplicate from each set of exact or approximate duplicates As always, there are sister operators for working with approximate and exact duplicates: - `remove_all_approximate_duplicates`: removes all near-duplicate images from a dataset - `remove_all_exact_duplicates`: removes all exact duplicate images from a dataset - `deduplicate_approximate_duplicates`: removes near-duplicate images from a dataset, _keeping a representative image_ from each duplicate set - `deduplicate_exact_duplicates`: removes exact duplicate images from a dataset, _keeping a representative image_ from each duplicate set Here’s an example of each: ![](https://cdn.sanity.io/images/h6toihm1/production/43a3c99bd3a3eff2ab70c838073f88a1be5d6912-2560x1356.gif?auto=format&dpr=2&fit=max&q=75&w=1600)![](https://cdn.sanity.io/images/h6toihm1/production/9164b30bb78f8035df681dc437406ff943f57596-2560x1355.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## Installing the Plugin If you haven’t already done so, install FiftyOne: ```bash 1pip install fiftyone ``` Then you can download this plugin from the command line with: ```bash 1fiftyone plugins download https://github.com/jacobmarks/image-dedup-plugin ``` Refresh the FiftyOne App, and you should see the eight operators in your operators list when you press the “ `` ` ``” key. ## Lessons Learned The Image Deduplication plugin is a Python Plugin with the usual structure (an `__init__.py`, `fiftyone.yml`, and `REAMDE.md` files). Additionally, it has the following: - An assets folder for storing icons - A Python file `exact_dups.py` for handling the logic and computations involved for exact duplicates - A Python file `approx_dups.py` for handling the logic and computations involved for approximate duplicates ### Splitting Code into Submodules It’s typically good practice in software development to make code modular, splitting self-contained pieces of logic into separate functions or files. This is known as [separation of concerns](https://en.wikipedia.org/wiki/Separation_of_concerns). The Image Deduplication plugin was a good exercise in applying this principle to FiftyOne’s plugin system. To utilize functions or variables you define in another file in the FiftyOne plugin’s directory, you need to add the path to that file to your system path. Here’s an example where we import `find_exact_duplicates` from the `exact_dups` file: ```python 1from fiftyone.core.utils import add_sys_path 2with add_sys_path(os.path.dirname(os.path.abspath(__file__))): 3 # pylint: disable=no-name-in-module,import-error 4 from exact_dups import find_exact_duplicates 5 ``` Starting from the innermost part of this expression: - `__file__` is a variable containing the path to the current module — in this case the `__init__.py` file. - `os.path.abspath` gets the absolute path for this file - `os.path.dirname` extracts the directory name of this absolute path - `add_sys_path` is a FiftyOne utility function that adds this to our system path The second to last line, `# pylint: disable=no-name-in-module,import-error` tells our linter not to throw an error when [linting](https://en.wikipedia.org/wiki/Lint_(software)) the file. ### Loading a View When executed, the `display_approximate_duplicate_groups` and `display_exact_duplicate_groups` operators each trigger the loading of specific views. Doing this is pretty straightforward, but it is worth noting that the data passed into `params` in the `ctx.trigger()` call needs to be serialized. In fact, _all_ data passed into parameter dictionaries for FiftyOne operators needs to be [serialized](https://hazelcast.com/glossary/serialization/). Fortunately, FiftyOne `DatasetView` objects are easy to serialize! ```python 1import json 2from bson import json_util 3 4def serialize_view(view): 5 return json.loads(json_util.dumps(view._serialize())) 6 ``` ### Icons for Each Operator The last tip is a simple but fun one: By utilizing the `icon` argument in the operator config, you can specify a unique icon for each operator. This is the icon that will then show up in the operators list with you hit “ `` ` ``”. For example, here’s the start of the operator definition for `FindExactDuplicates`: ```python 1class FindExactDuplicates(foo.Operator): 2 @property 3 def config(self): 4 return foo.OperatorConfig( 5 name="find_exact_duplicate_images", 6 label="Dedup: Find exact duplicates", 7 description="Find exact duplicates in the dataset", 8 icon="/assets/exact_duplicates.svg", 9 dynamic=True, 10 ) 11 ``` I like to put all of the SVGs I use as icons in an `assets` folder to stay organized 📁. ## Conclusion Building a high quality dataset doesn’t have to be a hassle. With our [Image Quality Issues Plugin](https://github.com/jacobmarks/image-quality-issues) from week 0, you can find a variety of common issues potentially plaguing images in your dataset, from peculiar aspect ratios to oversaturation. Now with the Image Deduplication Plugin (this post) you can also find and eliminate duplicates from your dataset in mere minutes! Stay tuned over the remaining weeks in the _Ten Weeks of FiftyOne Plugins_ while we continue to pump out a killer lineup of plugins! You can track our journey in our [ten-weeks-of-plugins repo](https://github.com/jacobmarks/ten-weeks-of-plugins) — and I encourage you to fork the repo and join me on this journey! [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [custom plugins](https://voxel51.com/blog/tag/custom-plugins) [deduplication](https://voxel51.com/blog/tag/deduplication) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [image dataset](https://voxel51.com/blog/tag/image-dataset) [plugins](https://voxel51.com/blog/tag/plugins) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/bad75ba72dfae8cdefc3d9afe33a1ea9a9c4ec36-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Optical Character Recognition with PyTesseract\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 21, 2023](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract) [![](https://cdn.sanity.io/images/h6toihm1/production/3efef5551e07ae9c6a1d190a5bd256b9723c78c7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Zero-Shot Prediction Plugin for FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 28, 2023](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/c1075849ea942531ffbb4afc830cc12598f0d919-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Supercharge Your Annotation Workflow with Active Learning\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Oct 5, 2023](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-332-lllmstxt|> ## Creating Pose Skeletons [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) Creating Pose Skeletons from Scratch – FiftyOne Tips and Tricks – September 15, 2023 Sep 16, 2023 • 4 min read Article content In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#3fcddd9865d6) [Pose Skeletons](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#194c19ddcdfb) [Preparing Your Dataset](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#1489fbb87d30) [Annotating Your Skeleton](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#3995dff8926e) [Conclusion](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#2336c3aee241) [Join the FiftyOne Community!](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#836566672b5b) In this article [Wait, What’s FiftyOne?](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#3fcddd9865d6) [Pose Skeletons](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#194c19ddcdfb) [Preparing Your Dataset](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#1489fbb87d30) [Annotating Your Skeleton](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#3995dff8926e) [Conclusion](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#2336c3aee241) [Join the FiftyOne Community!](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023#836566672b5b) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we cover interesting workflows and features of FiftyOne! This week we are getting ready for the spooky season with some [skeletons](https://docs.voxel51.com/user_guide/using_datasets.html#storing-keypoint-skeletons). We aim to cover the basics of creating a skeleton dataset using keypoints starting with just an image. ## **Wait, What’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! Also feel free to follow along in our [notebook](https://github.com/voxel51/fiftyone-examples/blob/keypoints/examples/Keypoints.ipynb) or on [YouTube](https://youtu.be/WLsVrMoGUzY)! ## **Pose Skeletons** In computer vision, pose skeletons are vital for understanding human or animal motion in images or videos, facilitating identification of precise body position and movement annotation. They also play a crucial role in human pose estimation datasets, aiding machine learning model training for applications in human-computer interaction, surveillance, and healthcare. In FiftyOne, pose skeletons are stored with the Keypoints class. The [Keypoints](https://docs.voxel51.com/api/fiftyone.core.labels.html#fiftyone.core.labels.Keypoints) class represents a collection of keypoint groups in an image. Each element of this list is a Keypoint object whose [points](https://docs.voxel51.com/api/fiftyone.core.labels.html#fiftyone.core.labels.Keypoint.points) attribute contains a list of (x, y) coordinates defining a group of semantically related keypoints in the image. For example, if you are working with a person model that outputs 18 keypoints (left eye, right eye, nose, etc.) per person, then each Keypoint instance would represent one person, and a Keypoints instance would represent the list of people in the image. ## **Preparing Your Dataset** Creating your own skeletons in FiftyOne is easy and quick. If you are starting from just images, start by creating a view or a dataset of the images you plan on annotating with skeletons. I chose to use the `quickstart` dataset as a nice example. ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4 5dataset = foz.load_zoo_dataset( 6 "quickstart", 7 dataset_name="skeletons" 8) 9 10 11session = fo.launch_app(dataset) ``` Using the [FiftyOne App](https://docs.voxel51.com/user_guide/app.html), I am going to tag the first person I see, which happens to be this cool skateboarder, in order to then automatically send it out for keypoint annotation using an annotation integration (in this case CVAT). To do so, simply select the image, click on the tag image and add “annotate” to its sample tags. Next, we need to prepare our dataset to expect keypoint skeletons. Using [dataset.skeletons](https://docs.voxel51.com/api/fiftyone.core.dataset.html?highlight=dataset%20skeletons#fiftyone.core.dataset.Dataset.skeletons), we can add our expected labels and connections for our [fo.KeypointSkeleton](https://docs.voxel51.com/api/fiftyone.core.odm.dataset.html#fiftyone.core.odm.dataset.KeypointSkeleton). Two inputs are provided, labels and edges. Labels will be the parts of the skeleton we are interested in and edges are how they are connected. Note that for labels and edges, the index will always correspond to the keypoint index. Hence, in my example, “left hand” will always be my first keypoint. I also chose to break my edges into two groups whose points will connect with each other, but not the other group. ```python 1dataset.skeletons = { 2 "points": fo.KeypointSkeleton( 3 labels=[\ 4 "left hand" "left shoulder", "right shoulder", "right hand",\ 5 "left eye", "right eye", "mouth",\ 6 ], 7 edges=[[0, 1, 2, 3], [4, 5, 6]], 8 ) 9} 10 11 12dataset.save() 13 ``` ## **Annotating Your Skeleton** To create a skeleton we are going to need some annotated keypoints on our image. If you already have annotations prepared you can skip this step. If you are starting from scratch, no problem. Follow along to create some keypoints with [FiftyOne’s CVAT](https://docs.voxel51.com/tutorials/cvat_annotation.html) integration. If you haven’t created a CVAT account yet, you will need to hop over to [create one](https://app.cvat.ai/). The first step is to plug in your username and password to environmental variables. ```python 1!export FIFTYONE_CVAT_USERNAME="" 2!export FIFTYONE_CVAT_PASSWORD="" ``` Next, let's grab the sample we tagged earlier and create a view for annotation. ```python 1ann_view = dataset.match_tags("annotate") 2ann_view ``` After, we want to launch the CVAT tool with our image. We provide an annotation key to retrieve our results later, as well as the new label field and type that we will be annotating for. ```python 1# A unique identifer for this run 2anno_key = "skeleton" 3 4 5# Upload the sample and launch CVAT 6anno_results = ann_view.annotate( 7 anno_key, 8 label_field="points", 9 label_type="keypoints", 10 classes=["person"], 11 launch_editor=True, 12) ``` As we annotate, make sure to annotate the keypoints in the correct order for the skeleton! After you are finished and the job is completed, load the new keypoints in. We can load our annotations back to FiftyOne like so after completion: ```python 1ann_view.load_annotations("skeleton", cleanup=True) 2 3 4session.view = ann_view 5 ``` ## **Conclusion** Just like that, we've explored the smooth process of preparing your dataset and annotating it with skeletons using FiftyOne. Whether you're starting from scratch or have existing annotations, FiftyOne's annotation integrations and keypoints simplified skeleton workflow allows you to efficiently define labels and connections for keypoints on your images. With just a few lines of code, your dataset can be configured to expect keypoint skeletons, and the CVAT tool facilitates the creation of annotated skeletons in the correct order. You can easily load these annotations back into your dataset for further analysis, providing a valuable resource for enhancing your computer vision and machine learning projects. FiftyOne makes the entire process accessible to both beginners and experienced practitioners, empowering you to tackle complex tasks and develop advanced computer vision models. Enjoy your skeletons! ## **Join the FiftyOne Community!** Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 2,000+ [FiftyOne Slack](https://slack.voxel51.com/) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 370+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [Keypoints](https://voxel51.com/blog/tag/keypoints) [labels](https://voxel51.com/blog/tag/labels) [machine learning](https://voxel51.com/blog/tag/machine-learning) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/342d5ec796cb4ee56573cc057c9e2e03542f5228-1200x674.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks — Jan 13, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 14, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-13-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/4ac1a727dc192a21563cde51b6e345f620e09376-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Jan 27, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Jan 28, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-jan-27-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/33ae18a60c78502fbffbd47575188ef0e9ed420f-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Tips and Tricks – Oct 27, 2023\\ \\ Computer Vision, Tips & Tricks\\ \\ • \\ \\ Oct 27, 2023](https://voxel51.com/blog/fiftyone-computer-vision-tips-and-tricks-oct-27-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-333-lllmstxt|> ## Optical Character Recognition Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Plugins](https://voxel51.com/blog/category/plugins), [Tutorials](https://voxel51.com/blog/category/tutorials) Optical Character Recognition with PyTesseract Sep 21, 2023 • 9 min read Article content In this article [Parse PDFs and filter by contents in FiftyOne](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#74f218515d4b) [OCR++ 👁️ 🔤🫵](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#28637a07f839) [PyTesseract OCR Plugin Overview](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#4a48cd5df50e) [Installing the Plugins](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#82eaed4097dc) [Lessons Learned](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#c5b0365486a8) [Conclusion](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#24f8b13edcbd) In this article [Parse PDFs and filter by contents in FiftyOne](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#74f218515d4b) [OCR++ 👁️ 🔤🫵](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#28637a07f839) [PyTesseract OCR Plugin Overview](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#4a48cd5df50e) [Installing the Plugins](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#82eaed4097dc) [Lessons Learned](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#c5b0365486a8) [Conclusion](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract#24f8b13edcbd) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Parse PDFs and filter by contents in FiftyOne Welcome to week five of _Ten Weeks of Plugins_. During these ten weeks, we will be building a FiftyOne Plugin (or multiple!) each week and sharing the lessons learned! If you’re new to them, FiftyOne Plugins provide a flexible mechanism for anyone to extend the functionality of the FiftyOne computer vision App. Similar to how you customize your favorite IDE to make quick work of repetitive or novel tasks, this is what Plugins are all about. If you are just joining us, you may find the following resources helpful: - [FiftyOne Plugins Repo](https://github.com/voxel51/fiftyone-plugins) - [FiftyOne Plugin Docs](https://docs.voxel51.com/plugins/index.html#downloading-plugins) - Plugins Channel in the [FiftyOne Community Slack](https://slack.voxel51.com/) If you’ve have been following along the last few weeks, let’s recap what we’ve built so far: - Week 0: 🌩️ [Image Quality Issues](https://github.com/jacobmarks/image-quality-issues) & 📈 [Concept Interpolation](https://github.com/jacobmarks/concept-interpolation) - Week 1: 🎨 [AI Art Gallery](https://github.com/jacobmarks/ai-art-gallery) & [Twilio Automation](https://github.com/jacobmarks/twilio-automation-plugin) - Week 2: ❓ [Visual Question Answering](https://github.com/jacobmarks/vqa-plugin) - Week 3: 🎥 [YouTube Player Panel](https://github.com/jacobmarks/fiftyone-youtube-panel-plugin) - Week 4: 🪞 [Image Deduplication](https://github.com/jacobmarks/image-deduplication-plugin) Ok, let’s dive into this week’s FiftyOne Plugins — 👓 [Optical Character Recognition](https://github.com/jacobmarks/pytesseract-ocr-plugin) (OCR) and 🔑 [Keyword Search](https://github.com/jacobmarks/keyword-search-plugin)! ## OCR++ 👁️ 🔤🫵 Optical Character Recognition (OCR) is a fundamental task in computer vision which entails recognizing the characters in a document, when said document is treated as an image. OCR can be employed to recognize typed characters, handwritten text, or even curved word art, and it has applications across multiple industries, from banking and law to healthcare. An OCR “engine” is the pipeline — either rules-based or powered by a machine learning model — which turns a document into a set of localized text strings. This week, I set out to streamline OCR and natural language document understanding workflows in FiftyOne! To do so, I built two connected plugins. The first plugin PyTesseract OCR, leverages the popular [Tesseract](https://github.com/tesseract-ocr/tesseract) OCR engine to perform optical character recognition, and converts the engine’s outputs into `Detection` labels. The second plugin is a Keyword Search plugin which allows you to search within the labels generated by the first plugin. When combined, these two plugins effectively allow you to query documents like pages of old books, handwritten notes, or resumes by the text that they contain! ## PyTesseract OCR Plugin Overview The PyTesseract OCR plugin is essentially a wrapper around the Tesseract OCR engine. The plugin has just one operator, `run_ocr_engine`, which performs OCR on each sample in the dataset and stores the results on the samples. Because this is a Python plugin, it interacts with Tesseract through the engine’s Python bindings, exposed by the `pytesseract` library. For each sample, this starts by extracting the filepath, and passing this into PyTesseract’s `image_to_data` function. This dictionary of data is then parsed and turned into two sets of detections: 1. Word Detections: a `Detections` field on the sample that contains a separate `Detection`, with a confidence score, for each detected word. 2. Block Detections: a `Detections` field on the sample that aggregates words in each contiguous _block_ in the document. 💡 `run_ocr_engine` is a _delegated_ operator. Instead of waiting while the OCR engine runs, executing the operator queues a job, which you can start from the command line! Let’s see this in action on the Form Understanding in Noisy Scanned Documents ( [FUNSD](https://paperswithcode.com/dataset/funsd)) dataset. First, we queue the job from the FiftyOne App: We can then list our queued jobs from the command line with `fiftyone delegated list`: ```generic 1id operator dataset queued_at state completed 2------------------------ ------------------------------------------ --------- ------------------- ------- ----------- 3650252f0e262b5c491d50d94 @jacobmarks/pytesseract_ocr/run_ocr_engine FUNSD 2023-09-14 00:25:20 queued ``` To launch the job, we can run `fiftyone delegated launch`. When the job completes, we can refresh the app, and we will see new `pt_word_predictions` and `pt_block_predictions` fields populated on our samples: This is what the block predictions look like for a single sample: ### Keyword Search Plugin Overview When I read long PDFs for research papers, textbooks, or contracts, I typically find myself relying heavily on keyword search. With a simple control-F, I can find every occurrence of a specific string in the document, across all pages. After running OCR on your documents, being able to search through the generated text seems like a natural thing to do! To achieve this functionality, I built a Keyword Search plugin. The plugin leverages FiftyOne’s `contains_str()` view expression, [which tests](https://docs.voxel51.com/api/fiftyone.core.expressions.html?highlight=viewexpression#fiftyone.core.expressions.ViewExpression.contains_str) whether a string field on a sample contains a substring. For example, the following code filters the [Quickstart](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#dataset-zoo-quickstart) dataset for predictions whose labels contain the string “be”: ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5dataset = foz.load_zoo_dataset("quickstart") 6 7# Only contains predictions whose `label` contains "be" 8view = dataset.filter_labels( 9 "predictions", F("label").contains_str("be") 10) 11### view will only contain ['bear', 'bed', 'bench', 'frisbee', 'teddy bear'] ``` However, the complete syntax for querying the dataset with `contains_str()` depends on the field to which it is applied. For string fields embedded in label fields, such as the `predictions.detections.label` field above, the view stage which achieves a keyword search-like effect would be `match_labels()`. For top-level string fields, on the other hand, the right view stage is just `match()`. There’s an additional level of complexity when considering lists of strings. Take a sample’s `tags` field for instance. When we search for a specific keyword, we need to check if any of the strings in the list contain the substring. The Keyword Search plugin works by finding all string fields and list fields with string elements in the dataset, and letting the user select which of these fields they want to search within. Depending on the type of the selected field, the appropriate querying syntax is used to perform the search. All of this is wrapped in a `search_by_keyword` operator. The four supported options are: 1. Top-level `StringField` 2. `StringField` within a `Label` field 3. Top-level `ListField`, whose elements are strings 4. `ListField` with string elements within a `Label` field The plugin also exposes the `case_sensitive` argument from `contains_str()` to the user, so they can decide whether they want the search to be performed in a case sensitive or case insensitive fashion. We can apply this `search_by_keyword` operator to the label field containing our OCR predictions in order to find documents that contain certain keywords! Here’s an example where we filter for documents in the FUNSD datasets with the word `contract` in them: The final element of the Keyword Search plugin worth noting is that it retains a “memory” of which field you last searched within. This saves you from selecting the same field from the field selector dropdown each time you want to modify your search! More on this in the lessons learned section below. ## Installing the Plugins If you haven’t already done so, install FiftyOne: ```bash 1pip install fiftyone ``` You can now download these plugins from the command line with: ```bash 1fiftyone plugins download https://github.com/jacobmarks/keyword-search-plugin 2fiftyone plugins download https://github.com/jacobmarks/pytesseract-ocr-plugin ``` To run the OCR plugin, you will need to have `pllow` and `pytesseract` installed. These are included in the `requirements.txt` file for the OCR plugin, so you can install them by running: ```bash 1fiftyone plugins requirements pytesseract-ocr-plugin --install ``` After downloading the plugins (and installing requirements), refresh the FiftyOne App, and you should see two new buttons in the Sample Actions Menu: _Buttons for search\_by\_keyword and run\_ocr\_engine operators_ You will also find these operators in the operators list when you press the “\`” key. ## Lessons Learned Both the OCR and Keyword Search plugins are pure Python plugins with the same basic structure: - `__init__.py`: operators are defined - `fiftyone.yml`: plugin information is defined and registered - `README.md`: plugin and operators are described, and installation instructions are documented. - `assets` folder: operator icons are stored In addition, the Keyword Search plugin has a cache manager file, `cache_manager.py`, whose purpose will be described shortly, and the OCR plugin has an `ocr_engine.py` file, which handles the running and postprocessing of data from the PyTesseract OCR engine. ### Caching in Python Plugins When building the [VoxelGPT plugin](https://github.com/voxel51/voxelgpt) (a chat-based coding assistant for the FiftyOne App leveraging LLMs), one of the key considerations was latency. When you chain multiple prompts and large language model queries together, the time elapsed between user question and final response can add up quickly. One of the many techniques we used to minimize latency was [aggressively _caching_](https://github.com/search?q=repo%3Avoxel51%2Fvoxelgpt%20cache&type=code). For VoxelGPT, we did this to minimize time associated with reading in files — if you aren’t changing the contents of the file, but you are going to use the data often, why not store it in a global cache so you only need to read it in once. For the Keyword Search plugin, I set out to use caching for a completely different purpose — remembering a user’s choices. Once the user selects a field which they want to perform the keyword search on, this field becomes our best guess for the field on which the user would want to perform their next keyword search. Instead of making the user select this same field from the dropdown selector each time they open the operator’s modal, why not cache the last-used field and set this as the default value? As I found out while building the Keyword Search plugin, however, there is some nuance required to achieve this effect. First off, you can’t do any caching within the `__init__.py` file itself, because this file gets reloaded and its variables reset every time the operator is used. Instead, the solution I came up with was to create a `cache_manager.py` module which exclusively defines a `get_cache()` function, and then import this function from the module in `__init__.py`. The second subtlety is that caching a user’s preference in this manner only works if the application is being served to a single user via a single node. If this approach were used on try.fiftyone.ai, where each interaction is potentially fulfilled by a different instance of the plugin on a different node, there would be no way of determining which user each preference should be associated with. This is why the import statement for the cache manager in `__init__.py` is wrapped with an if statement that checks if the code is running on a multi-user deployment: ```python 1def _is_teams_deployment(): 2 val = os.environ.get("FIFTYONE_INTERNAL_SERVICE", "") 3 return val.lower() in ("true", "1") 4 5TEAMS_DEPLOYMENT = _is_teams_deployment() 6 7if not TEAMS_DEPLOYMENT: 8 with add_sys_path(os.path.dirname(os.path.abspath(__file__))): 9 # pylint: disable=no-name-in-module,import-error 10 from cache_manager import get_cache ``` The key takeaway here is that true statefulness within a plugin is only possible with JavaScript plugins! ### Menu Buttons in Python Plugins You don’t need to turn your Python plugin into a JavaScript plugin to create sleek buttons for your operators. On the contrary, you can actually specify these in the operator’s definition with the `resolve_placement()` method. This is all the code required to turn the `run_ocr_engine` operator into a nice button: ```python 1def resolve_placement(self, ctx): 2 return types.Placement( 3 types.Places.SAMPLES_GRID_ACTIONS, 4 types.Button( 5 label="Detect text in images", 6 icon="/assets/icon_light.svg", 7 ), 8 ) ``` The first input to types.Placement(), `SAMPLES_GRID_ACTIONS`, is what determines where the button shows up. You can find more info on the other allowed placements [here](https://docs.voxel51.com/plugins/index.html#placements). As in the operator’s config, we specify the icon to render with the `icon` argument. There is no requirement that the icon you use for the operator’s config (what shows up in the operators list) and the icon that you use for the button are the same, or even related. Hypothetically, you could make them completely different! ### Plugging in your Plugins As I continue to develop more and more plugins, I find myself increasingly thinking about how the plugins will, well, plug into each other. Will the outputs of one plugin be suitable as inputs into another plugin? In other words, will my plugins play nice with each other? > _A large portion of the value inherent in plugins is the ability to interweave multiple plugins to construct sophisticated, tailored workflows. And as the FiftyOne plugin ecosystem grows, the space of enabled workflows grows exponentially. To make the most of this interconnectivity, plugins need to be modular and flexible._ The Keyword Search plugin arose from a specific problem: I wanted to be able to search through the contents of the documents on which I had run OCR. In theory, this only required the ability to search through `String` field attributes of `Detection` labels. However, it wasn’t a huge stretch for me to imagine other workflows involving keyword search on lists of strings, or lists of strings within `Detection` label fields. The plugin that I built is flexible enough to plug into any/all of these workflows! I hope this approach proves helpful as you build out your arsenal of FiftyOne plugins! ## Conclusion By combining OCR with Keyword Search, you can filter PDFs or other text-based documents by their content, in the same way that you would filter for images that have dogs or cats. These two plugins are certainly not exhaustive, but I hope they demonstrate how you can approach natural language based document understanding with the visualization and querying capabilities of FiftyOne! Stay tuned over the remaining weeks in the _Ten Weeks of FiftyOne Plugins_ while we continue to pump out a killer lineup of plugins! You can track our journey in our [ten-weeks-of-plugins repo](https://github.com/jacobmarks/ten-weeks-of-plugins) — and I encourage you to fork the repo and join me on this journey! [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [custom plugins](https://voxel51.com/blog/tag/custom-plugins) [deduplication](https://voxel51.com/blog/tag/deduplication) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [image dataset](https://voxel51.com/blog/tag/image-dataset) [plugins](https://voxel51.com/blog/tag/plugins) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/fedd0c008df994c1834839ea35bc65b34e47fb5d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Double Trouble: Eliminate Image Duplicates with FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 14, 2023](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/3efef5551e07ae9c6a1d190a5bd256b9723c78c7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Zero-Shot Prediction Plugin for FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 28, 2023](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/c1075849ea942531ffbb4afc830cc12598f0d919-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Supercharge Your Annotation Workflow with Active Learning\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Oct 5, 2023](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-334-lllmstxt|> ## FiftyOne Teams 1.4 Announcement [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Product & News](https://voxel51.com/blog/category/product-news) Announcing FiftyOne Teams 1.4 with Dataset Versioning, Delegated Operations, and Ultralytics Integration Sep 20, 2023 • 8 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#efbc2dee6510) [Okay, but what’s FiftyOne Teams?](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#232bf6f127c5) [tl;dr: What’s new in FiftyOne Teams 1.4?](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#957686f2dd83) [Dataset Versioning](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#207baaec64fe) [Delegated Operations](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#abf2b6e49ee7) [Ultralytics integration](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#220173cd2d47) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#efbc2dee6510) [Okay, but what’s FiftyOne Teams?](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#232bf6f127c5) [tl;dr: What’s new in FiftyOne Teams 1.4?](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#957686f2dd83) [Dataset Versioning](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#207baaec64fe) [Delegated Operations](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#abf2b6e49ee7) [Ultralytics integration](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#220173cd2d47) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) _**Editor's note:** FiftyOne Teams is now [FiftyOne Enterprise](https://voxel51.com/enterprise/)!_ I’m thrilled to announce the general availability of [FiftyOne Teams 1.4](https://docs.voxel51.com/release-notes.html#fiftyone-teams-1-4-0), an exciting addition to the FiftyOne ecosystem that brings significant upgrades to the representational power and extensibility of FiftyOne. First, with [**Dataset Versioning**](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#dataset-versioning), we’re introducing the ability to create snapshots of your datasets, enabling you to track the lineage of your data, view the exact version of a dataset on which a given model was trained, protect against accidental or fruitless data modification, and much more. Second, we’re launching [**Delegated Operations**](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#delegated-operations), an all-new extension of FiftyOne’s plugin framework that allows you to schedule customizable tasks from within the App that are executed on a connected workflow orchestrator like Apache Airflow. Bring your _data_ to the center of your AI workflows. Finally, we’ve added a native [**Ultralytics integration**](https://voxel51.com/blog/announcing-fiftyone-teams-1-4-with-dataset-versioning-delegated-operations-and-ultralytics-integration#ultralytics), making it easy to train the latest Ultralytics models like YOLOv8 on your FiftyOne datasets with just a few lines of code. Whether you’re prototyping a new project or leading a large effort to deploy a machine learning system into production, your AI stack needs a flexible data-centric component that enables you to organize, visualize, and iterate on your data. At Voxel51, we’re proud that tens of thousands of developers now trust open source FiftyOne as the source of truth for their ML data and use FiftyOne Teams to collaborate with their team to build better models through better data. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is the open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. ## **Okay, but what’s FiftyOne Teams?** [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) extends FiftyOne with a GSuite-like experience for teams that want to collaborate on data stored in a centralized location with additional features like user permissions, dataset versioning, cloud-backed media, and enterprise security. If this sounds interesting, read on! Then [schedule a workshop](https://voxel51.com/schedule-teams-workshop/) to learn more about FiftyOne Teams. ## **tl;dr: What’s new in FiftyOne Teams 1.4**? This release includes three headline features: - **Dataset Versioning**: you can now create snapshots of your datasets and view/rollback to previous snapshots, enabling you to track the lineage of your data, view the exact version of a dataset on which a given model was trained, protect against accidental data modification, and much more. - **Delegated Operations:** this all-new feature allows you to schedule tasks from within the FiftyOne App that are executed on a connected workflow orchestration tool like Apache Airflow. You can choose from dozens of prebuilt operations or write, permission, and upload your own custom workflows to your Teams deployment. - **Ultralytics Integration**: this integration makes it easy to train the latest Ultralytics models like [YOLOv8](https://github.com/ultralytics/ultralytics) on your FiftyOne datasets with just a few lines of code. Check out the [release notes](https://docs.voxel51.com/release-notes.html#fiftyone-teams-1-4-0) for a full rundown of additional enhancements and bugfixes in FiftyOne Teams 1.4. ## **Dataset Versioning** One of the most popular feature requests from FiftyOne Teams users has been the ability to create and track versions of their datasets natively within FiftyOne. Now with FiftyOne Teams 1.4, you can! [Dataset Versioning](https://docs.voxel51.com/teams/dataset_versioning.html) allows you to track the lineage of your data, view the exact version of a dataset on which a given model was trained, protect against accidental data modification, and much more. With dataset versioning, you can trust FiftyOne Teams as the single source of truth for your organization’s data. Check out the 60 second video below for a walkthrough of the feature: Dataset Versioning in FiftyOne Teams is implemented as a linear sequence of read-only snapshots. In other words, creating a new snapshot creates a permanent record of the dataset’s contents that can be loaded and viewed at any time in the future, but not directly edited. Conversely, the current version of a dataset is called its HEAD (think git). If you have not explicitly loaded a snapshot, you are viewing its HEAD, and you can make additions, updates, and deletions to the dataset’s contents as you normally would (provided you have sufficient permissions). Any user with Can View access to a dataset can view its snapshots. However, only users with Can Manage permissions to a dataset can create, edit, and delete its snapshots. Dataset snapshots record all aspects of your data stored within FiftyOne, including dataset-level information, schema, samples, frames, brain runs, and evaluations. However, snapshots _exclude_ any information stored in external services, such as media stored in cloud buckets or embeddings stored in an external vector database, which are assumed to be immutable. If you need to update the image for a sample in a dataset, for example, update the sample’s filepath—which is tracked by snapshots—rather than updating the media in cloud storage in-place—which would _not_ be tracked by snapshots. This design allows dataset snapshots to be as lightweight and versatile as possible. After upgrading to FiftyOne Teams 1.4, all datasets will have a History tab in the App that provides access to the dataset’s versioning history. From this tab, you can easily: - View changes between the dataset’s HEAD and its most recent snapshot (if any) - Create a new snapshot - View the dataset’s snapshots and their associated metadata - Open a snapshot in the App to visualize it - Clone a snapshot into a new dataset - Rollback a dataset to a previous snapshot - Delete a snapshot You can programmatically load a specific snapshot of a dataset via the FiftyOne Teams SDK by passing the optional snapshot argument to [load\_dataset()](https://docs.voxel51.com/api/fiftyone.core.dataset.html#fiftyone.core.dataset.load_dataset): ```python 1import fiftyone as fo 2 3# Load a snapshot 4dataset = fo.load_dataset("dataset name", snapshot="snapshot name") ``` Snapshots behave exactly like regular FiftyOne datasets, except that they are always read-only. If you have Can Manage permissions to a dataset, you can also use the [Management SDK](https://docs.voxel51.com/teams/dataset_versioning.html#using-snapshots) to programmatically create and manipulate snapshots: ```python 1import fiftyone.management as fom 2 3# List available snapshots 4fom.list_snapshots(dataset_name) 5 6# Create a new snapshot 7fom.create_snapshot(dataset_name, snapshot_name, **kwargs) 8 9# Delete a snapshot 10fom.delete_snapshot(dataset_name) 11 12# Revert a dataset to a snapshot 13fom.revert_dataset_to_snapshot(dataset_name, snapshot_name) ``` For more information about Dataset Versioning in FiftyOne, check out [the docs](https://docs.voxel51.com/teams/dataset_versioning.html). ## **Delegated Operations** FiftyOne Teams 1.4 adds a powerful new Delegated Operations feature to [FiftyOne’s Plugin framework](https://docs.voxel51.com/plugins/index.html) that allows you to schedule builtin and/or custom tasks from within the App that are executed on a connected workflow orchestrator like Apache Airflow. Why is this awesome? Your AI stack needs a flexible data-centric component that enables you to organize and _compute on_ your data. With Delegated Operations, FiftyOne becomes both a dataset management/visualization tool and a workflow automation tool that defines how your data-centric workflows like ingestion, curation, and evaluation are performed. In short, think of FiftyOne Teams as the single source of truth on which you co-develop your data and models together. What can Delegated Operations do for you? Get started by installing any of these plugins available in the [FiftyOne Plugins repository](https://github.com/voxel51/fiftyone-plugins): - [@voxel51/annotation](https://github.com/voxel51/fiftyone-plugins/blob/main/plugins/annotation/README.md) \- ✏️ Utilities for integrating FiftyOne with annotation tools - [@voxel51/brain](https://github.com/voxel51/fiftyone-plugins/blob/main/plugins/brain/README.md) \- 🧠 Utilities for working with the FiftyOne Brain - [@voxel51/evaluation](https://github.com/voxel51/fiftyone-plugins/blob/main/plugins/evaluation/README.md) \- ✅ Utilities for evaluating models with FiftyOne - [@voxel51/io](https://github.com/voxel51/fiftyone-plugins/blob/main/plugins/io/README.md) \- 📁 A collection of import/export utilities - [@voxel51/indexes](https://github.com/voxel51/fiftyone-plugins/blob/main/plugins/indexes/README.md) \- 📈 Utilities working with FiftyOne database indexes - [@voxel51/utils](https://github.com/voxel51/fiftyone-plugins/blob/main/plugins/utils/README.md) \- ⚒️ Call your favorite SDK utilities from the App - [@voxel51/zoo](https://github.com/voxel51/fiftyone-plugins/blob/main/plugins/zoo/README.md) \- 🌎 Download datasets and run inference with models from the FiftyOne Zoo, all without leaving the App - [@voxel51/voxelgpt](https://github.com/voxel51/voxelgpt) \- 🤖An AI assistant that can query visual datasets, search the FiftyOne docs, and answer general computer vision questions For example, wish you could import data from within the App? With the [@voxel51/io](https://github.com/voxel51/fiftyone-plugins/blob/main/plugins/io/README.md) plugin, you can! Want to send data for annotation from within the App? Sure thing, just install the [@voxel51/annotation](https://github.com/voxel51/fiftyone-plugins/blob/main/plugins/annotation/README.md) plugin: Have model predictions on your dataset that you want to evaluate? The [@voxel51/evaluation](https://github.com/voxel51/fiftyone-plugins/blob/main/plugins/evaluation/README.md) plugin makes it easy: Need to compute embedding for your dataset so you can visualize them in the [Embeddings panel](https://docs.voxel51.com/user_guide/app.html#embeddings-panel)? Kick off the task with the [@voxel51/brain](https://github.com/voxel51/fiftyone-plugins/blob/main/plugins/brain/README.md) plugin and proceed with other work while the execution happens in the background: When you choose delegated execution in the App, the task is automatically scheduled for execution on your [connected orchestrator](https://docs.voxel51.com/teams/teams_plugins.html#teams-plugins-managing-operators-runs) and you can continue with other work. Meanwhile, all datasets have a Runs tab in the App where you can browse a history of all Delegated Operations that have been run on the dataset and their status: You can also click on an individual run to see information about its inputs, outputs, and any errors that occurred during execution: When [writing your own plugins](https://docs.voxel51.com/plugins/index.html), you can declare that an operation should be delegated by implementing the optional `resolve_delegation() ` method and returning `True`. That’s it? Yep! When a user executes a delegated operation, the execute() method is called by the [connected orchestrator](https://docs.voxel51.com/teams/teams_plugins.html#teams-plugins-managing-operators-runs) rather than being immediately executed by the App server when the user creates the task. ```python 1import fiftyone.operators as foo 2import fiftyone.operators.types as types 3 4class YourOperator(foo.Operator): 5 @property 6 def config(self): 7 return foo.OperatorConfig( 8 name="your_operator", 9 label="Your operator", 10 dynamic=True, 11 ) 12 13 def resolve_input(self, ctx): 14 # Collect your input parameters here 15 pass 16 17 def resolve_delegation(self, ctx): 18 # 19 # Return True to schedule this operation for delegated execution 20 # or False to execute it immediately 21 # 22 # You can optionally use `ctx` to dynamically determine whether 23 # to delegate based on the user-provided inputs 24 # 25 pass 26 27 def execute(self, ctx): 28 # Perform the operation using `ctx.dataset`, `ctx.params` etc 29 pass 30 31 def resolve_ouput(self, ctx): 32 # Render any output information here 33 pass 34 35def register(p): 36 p.register(YourOperator) ``` If an Operator requires secret information like API keys, usernames, or passwords that need to be stored encrypted, an Admin of your FiftyOne Teams deployment can upload the necessary key-value pairs using the [Secrets UI](https://docs.voxel51.com/teams/secrets.html) shown below: Then, as a plugin developer, simply declare the secrets that your Operator requires by adding them to the plugin’s `fiftyone.yml ` file: ```bash 1 - OPENAI_API_KEY ``` Your Operator can then access these secrets at runtime via the `ctx.secrets ` dict: ```python 1 def execute(self, ctx): 2 api_key = ctx.secrets["OPENAI_API_KEY"] # your Open AI API key ``` Note that the `ctx.secrets` dict will also be automatically populated with the values of any environment variables whose name matches a secret key declared by an Operator, so a plugin written using the above pattern can run in all of the following environments with _no code changes_: - Open source FiftyOne - Logged into FiftyOne Teams via the web - A locally launched App via the FiftyOne Teams SDK Check out [the docs](https://docs.voxel51.com/plugins/index.html#how-to-write-plugins) for more information about writing your own operators and plugins and [distributing](https://docs.voxel51.com/plugins/index.html#publishing-plugins) them via GitHub. ## **Ultralytics integration** FiftyOne Teams 1.4 also includes a native [Ultralytics integration](https://docs.voxel51.com/integrations/ultralytics.html), making it easy to train the latest Ultralytics models like [YOLOv8](https://docs.ultralytics.com/) on your FiftyOne datasets with just a few simple lines of code. Running inference on a FiftyOne dataset with an Ultralytics model is made simple with the utility methods in the new [fiftyone.utils.ultralytics](https://docs.voxel51.com/api/fiftyone.utils.ultralytics.html) module: ```python 1import fiftyone as fo 2import fiftyone.utils.ultralytics as fou 3from ultralytics import YOLO 4 5dataset = fo.load_dataset("your-dataset") 6model = YOLO("yolov8s.pt") 7 8for sample in dataset.iter_samples(progress=True): 9 result = model(sample.filepath)[0] 10 sample["boxes"] = fou.to_detections(result) 11 sample.save() 12 13session = fo.launch_app(dataset) ``` You can also use FiftyOne’s builtin [YOLO exporter](https://docs.voxel51.com/user_guide/export_datasets.html#yolov5dataset) to prepare datasets for training: ```python 1import fiftyone as fo 2import fiftyone.utils.ultralytics as fou 3import fiftyone.zoo as foz 4 5# The path to export the dataset 6EXPORT_DIR = "/tmp/oiv7-yolo" 7 8# Prepare train split 9 10train = foz.load_zoo_dataset( 11 "open-images-v7", 12 split="train", 13 label_types=["detections"], 14 max_samples=100, 15) 16 17# YOLO format requires a common classes list 18classes = train.default_classes 19 20train.export( 21 export_dir=EXPORT_DIR, 22 dataset_type=fo.types.YOLOv5Dataset, 23 label_field="ground_truth", 24 split="train", 25 classes=classes, 26) 27 28# Prepare validation split 29 30validation = foz.load_zoo_dataset( 31 "open-images-v7", 32 split="validation", 33 label_types=["detections"], 34 max_samples=10, 35) 36 37validation.export( 38 export_dir=EXPORT_DIR, 39 dataset_type=fo.types.YOLOv5Dataset, 40 label_field="ground_truth", 41 split="val", # Ultralytics uses 'val' 42 classes=classes, 43) ``` From here, [training an Ultralytics model](https://docs.ultralytics.com/modes/train/) is as simple as passing the path to the dataset YAML file: ```python 1from ultralytics import YOLO 2 3# The path to the `dataset.yaml` file we created above 4YAML_FILE = "/tmp/oiv7-yolo/dataset.yaml" 5 6# Load a model 7model = YOLO("yolov8s.pt") # load a pretrained model 8# model = YOLO("yolov8s.yaml") # build a model from scratch 9 10# Train the model 11model.train(data=YAML_FILE, epochs=3) 12 13# Evaluate model on the validation set 14metrics = model.val() 15 16# Export the model 17path = model.export(format="onnx") ``` Check out the [FiftyOne Integrations page](https://docs.voxel51.com/integrations/index.html) for a growing list of third-party libraries and products that FiftyOne natively integrates with. [custom plugins](https://voxel51.com/blog/tag/custom-plugins) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne Teams](https://voxel51.com/blog/tag/fiftyone-teams) [operators](https://voxel51.com/blog/tag/operators) [plugins](https://voxel51.com/blog/tag/plugins) [product release](https://voxel51.com/blog/tag/product-release) [Ultralytics](https://voxel51.com/blog/tag/ultralytics) ![](https://cdn.sanity.io/images/h6toihm1/production/8d61ff90b31d151405f9e21a33c2802509f34651-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Brian Moore Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/99a6360654f227b8d75d13381fe015cc7983ee09-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Announcing FiftyOne 0.21 with Operators, Dynamic Groups, and Custom Color Schemes\\ \\ Product & News\\ \\ • \\ \\ Jun 1, 2023](https://voxel51.com/blog/announcing-fiftyone-0-21) [![](https://cdn.sanity.io/images/h6toihm1/production/a4c2bee9ed053c5be2a1c161e5abf758c9a12ff8-1400x923.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.18 with App Performance Improvements, Sidebar Modes, and Custom Attributes\\ \\ Product & News\\ \\ • \\ \\ Nov 15, 2022](https://voxel51.com/blog/announcing-fiftyone-0-18-with-app-performance-improvements-sidebar-modes-and-custom-attributes) [![](https://cdn.sanity.io/images/h6toihm1/production/e94f20fa81716294c7a6caccf7256e9106cb6e89-967x800.png?auto=format&dpr=2&fit=crop&fp-x=0.5&fp-y=0.5&h=270&q=75&w=480)\\ \\ Announcing FiftyOne 0.17 with Grouped Datasets, 3D, Geolocation, and Custom Plugins\\ \\ Product & News\\ \\ • \\ \\ Sep 21, 2022](https://voxel51.com/blog/announcing-fiftyone-0-17-with-grouped-datasets-3d-geolocation-and-custom-plugins) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-335-lllmstxt|> ## nuScenes Dataset Insights [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Datasets](https://voxel51.com/blog/category/datasets) Navigating the Road Ahead Sep 20, 2023 • 8 min read Article content In this article [A Deep Dive into the nuScenes Dataset](https://voxel51.com/blog/nuscenes-dataset-navigating-the-road-ahead#c4c9ee42986c) [Comprehensive Data Composition](https://voxel51.com/blog/nuscenes-dataset-navigating-the-road-ahead#93f817827d06) [Precise Annotations](https://voxel51.com/blog/nuscenes-dataset-navigating-the-road-ahead#6af8fb3e344b) [Realistic Scenarios](https://voxel51.com/blog/nuscenes-dataset-navigating-the-road-ahead#56dadc08cfc9) In this article [A Deep Dive into the nuScenes Dataset](https://voxel51.com/blog/nuscenes-dataset-navigating-the-road-ahead#c4c9ee42986c) [Comprehensive Data Composition](https://voxel51.com/blog/nuscenes-dataset-navigating-the-road-ahead#93f817827d06) [Precise Annotations](https://voxel51.com/blog/nuscenes-dataset-navigating-the-road-ahead#6af8fb3e344b) [Realistic Scenarios](https://voxel51.com/blog/nuscenes-dataset-navigating-the-road-ahead#56dadc08cfc9) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## A Deep Dive into the nuScenes Dataset In the rapidly evolving landscape of autonomous vehicles, advancements in computer vision are steering us closer to a future of self-driving cars and enhanced mobility solutions. However, the development and validation of these autonomous systems heavily depend on the availability of high-quality, real-world data. Comprehensive datasets are the bedrock upon which cutting-edge computer vision algorithms are trained and refined, enabling vehicles to perceive, interpret, and react to their surroundings with unprecedented accuracy. Enter the [nuScenes](https://www.nuscenes.org/) dataset—an indispensable asset in the realm of computer vision for autonomous driving. > _nuScenes is a public large-scale dataset for autonomous driving. It enables researchers to study challenging urban driving situations using the full sensor suite of a real self-driving car._ In this blog post, we'll begin to unravel the layers of nuScenes, understanding its composition, the types of data it encompasses, and its pivotal role in training AI models using the open source FiftyOne computer vision toolset. We'll explore how this dataset advances the capabilities of computer vision, fuels research, and accelerates the deployment of autonomous vehicles, propelling us into a future where the roads are safer and transportation is more efficient. if you are new to FiftyOne, check out this short tour. The nuScenes dataset is not just a vast collection of images and sensor data; it's a meticulously crafted resource that provides a rich tapestry of information for training and testing computer vision algorithms in the context of autonomous driving. Let's delve into the essential aspects that define and make this dataset a cornerstone in computer vision research. ## **Comprehensive Data Composition** At its core, the nuScenes dataset is a treasure trove of multi-modal sensor data, comprising high-resolution camera images, LiDAR point clouds, RADAR sweeps, and more. This diverse range of data sources allows researchers and engineers to create holistic perception systems for autonomous vehicles, mimicking the real-world sensory inputs encountered on the road. nuScenes Car Set Up ## **Precise Annotations** Annotating data is a crucial step in training computer vision models. nuScenes excels in this aspect with its meticulous labeling of objects, such as vehicles, pedestrians, and cyclists, in each frame. These annotations are essential for tasks like object detection, tracking, and semantic segmentation, enabling AI models to understand and interact with their environment accurately. The dataset contains over 1.4 million object bounding boxes across forty thousand keyframes. It all sums up to 3 million individual sensor inputs across camera images, LIDAR, and RADAR samples! ## **Realistic Scenarios** nuScenes goes beyond just data collection; it provides a broad spectrum of real-world driving scenarios. From urban environments to highways, day and night scenes, and various weather conditions, it offers a comprehensive set of challenges that autonomous vehicles might encounter. This realism is vital for robust algorithm development, ensuring that AI systems can adapt to the unpredictability of the open road. 1,000 different scenes are collected from streets in Boston and Singapore through a variety of driving maneuvers, traffic situations, and unexpected behaviors. ### **High Quality and Quantity** Boasting a vast volume of data, nuScenes offers researchers the luxury of training deep learning models on extensive datasets. The dataset's high quality, combined with its quantity, allows for the development of highly accurate and reliable computer vision models that can operate seamlessly in the real world. nuScenes includes around five and half hours worth of data and contains 7x more annotations than [KITTI](http://www.cvlibs.net/datasets/kitti/), the dataset that inspired nuScenes and pioneered the ADAS dataset category. ### **Dynamic and Evolving** The nuScenes dataset is continuously evolving, with periodic updates and expansions. This ensures that it remains relevant and up-to-date, reflecting the ever-changing landscape of urban environments and traffic patterns. Understanding the nuScenes dataset is not just about recognizing its data components but also appreciating its role as a catalyst for innovation in computer vision and autonomous driving. It serves as a benchmark for the development of state-of-the-art perception systems, pushing the boundaries of what's possible in the world of AI-powered mobility. nuScenes also released in July 2020 their new lidarseg addon or LIDAR segmenntation, adding 1.4 billion annotated points across 40,000 point clouds. These points cover 32 possible semantic labels across each point in the LIDAR point clouds. With all this different data to take in, let's see how we can get started with nuScenes in FiftyOne. ### **Exploring nuScenes in FiftyOne** Due to the multi-sensor structure of nuScenes, the dataset in FiftyOne will be a [Grouped Dataset](https://docs.voxel51.com/user_guide/groups.html) with some [Dynamic Group Views](https://docs.voxel51.com/user_guide/using_views.html#view-groups) thrown in there as well. At a high level, we will group together our samples by their associated scene in nuScenes. At regular intervals of each keyframe or approximately every 0.5 seconds (2Hz), we incorporate data from every sensor type, including their respective detections. This amalgamation of data results in distinct groups, each representing the sensor perspective at a given keyframe. We do have each sensor input for every frame, but since only keyframes are annotated, we choose to only load those in. ### **Ingesting nuScenes** To get started with nuScenes in FiftyOne, first we need to set up our environment for nuScenes. It will require downloading the dataset or a snippet of it as well as downloading the nuScenes python sdk. Full steps on installing can be found [here](https://www.nuscenes.org/nuscenes?tutorial=nuscenes). Once your nuScenes is installed into your computer, we can kick things off. Let’s start by initializing both nuScenes as well as our FiftyOne dataset. We define our dataset as well as add a group to initialize the dataset to expect grouped data. nuScenes is comprised of three main data types: Camera, RADAR, and LIDAR. Camera data is made up of jpg’s of different camera angles on the car while RADAR and LIDAR are pcd files or point clouds that were captured using the car’s 3D sensors. We will break down how we ingest each form of media for nuScenes and bring it all together at the end. Before that, a quick intro on the hierarchy of nuScenes. The dataset is composed of scenes, scenes contain samples, and each sample contains tokens that link to both the sensor data and annotations. We will look at how to load data from a single sensor first. ### **Loading LIDAR Data** Loading a LIDAR sample from nuScenes is composed of two steps, generating the pointcloud and adding the detections. We must convert the binary point clouds to standard in order to ingest them. nuScenes also offers a LIDAR segmentation optional package that allows us to color each point cloud point a color corresponding to its class that we will be utilizing. We start with our lidar token, load in the color map and point cloud that corresponds to the token, and save them back to file with the new coloring and standard `pcd` point cloud file formatting. With our point cloud file now properly prepared for ingestion, we can move along to adding detections. To do so, we grab all the detections from the keyframe. We use nuScenes SDK’s builtin box methods to retrieve the location, rotation, and dimensions of the box. To match FiftyOne’s [3D detection input](https://docs.voxel51.com/user_guide/using_datasets.html#d-detections), we take `box.orientation.yaw_pitch_roll ` for rotation, `box.wlh ` for width, length, and height, and `box.center` for its location. Note too that `fo.Sample(filepath=filepath, group=group.element(sensor)) ` will automatically detect the pcd file and ingest the sample as a point cloud as well! After the method is run and detections are added, we have our LIDAR sample with detections! ### **Loading Camera Data** Camera data is a bit more straightforward than the lidar data. There is no need to do any prep on the image or save it as a different format. We can go right ahead and grab the detections to add them. The only tricky part is these are no ordinary bounding boxes, they are 3D bounding boxes given in global coordinates! Luckily for us, nuScenes provides some easy ways to convert their bounding boxes to our pixel space relative to what camera the image came from. For camera data, we load all of our boxes for our sample, check to see which ones are in the frame of our camera data, and then add the cuboids to the sample. In order to add a cuboid or a 3D bounding box, we use [polylines](https://docs.voxel51.com/user_guide/using_datasets.html#cuboids) and the [from\_cuboid](https://docs.voxel51.com/api/fiftyone.core.labels.html#fiftyone.core.labels.Polyline.from_cuboid)() method. Let’s take a look at how it is done: ### **Loading RADAR Data** RADAR is an interesting case. Since we have already stored our 3d detections in the LIDAR sample and RADAR is laid on top of the LIDAR in the 3D visualizer, we don’t need to copy our detections for each point cloud. The simplifies loading RADAR to just: This is something that can of course be changed by adding the same detection loop we used in LIDAR, especially if you are interested in object detection with only RADAR. With examples on how to load in each sensor input, we can take an initial look at how we bring it all together! As we begin to form our ingestor, we start with the definition of what our groups will be, one for each sensor on the car. We set up a for loop where we continue on to the next sample into the scene until we reach the last one, denoted by an empty “next” token. After we start looping through our dataset, we begin to create groups for each keyframe or nuScenes sample. Depending on the modality of the sensor, we take the appropriate action to load the sample. Once it has been created and added to our list, we move on to the next sample. Finally, with our samples all prepared for our dataset, we can take the last step and add them all to our dataset. ### **Conclusion** The nuScenes dataset stands as a meticulously curated resource, offering a diverse range of sensor data crucial for training and testing autonomous driving computer vision algorithms. With detailed annotations and a realistic array of driving scenarios, it serves as a cornerstone for building accurate perception systems for autonomous vehicles. Continuously evolving and expanding, nuScenes remains a vital benchmark in AI-driven mobility, facilitating innovation and pushing the boundaries of computer vision research. Its integration into FiftyOne provides a structured approach for efficient dataset exploration and analysis, aligning with its rich multi-sensor structure. To get started exploring, check out nuScenes on try.fiftyone.ai! While there, feel free to browse our other AV datasets such as [KITTI](https://try.fiftyone.ai/datasets/kitti/samples), [BDD](https://try.fiftyone.ai/datasets/bdd100k/samples), and [more](https://try.fiftyone.ai/datasets)! MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/0b2ac80144f5813c10886f51c8a1c34370d01a9f-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Visualize CVPR 2023 Datasets at CVPR 2023!\\ \\ Computer Vision, Datasets\\ \\ • \\ \\ Jun 20, 2023](https://voxel51.com/blog/visualize-cvpr-2023-datasets-at-cvpr-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/9772fa84c9388e3a372d53058366ae8a0c89bf4e-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Using SAM for the Prediction of Segmentations on the Kaggle Football Player Segmentation Dataset\\ \\ Computer Vision, Datasets\\ \\ • \\ \\ Oct 24, 2023](https://voxel51.com/blog/computer-vision-sam-for-prediction-kaggle-football-player-segmentation-dataset) [![](https://cdn.sanity.io/images/h6toihm1/production/ac96e4c6d5f31a869b66e9ad2ec01828091a3552-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How to Build a Semantic Search Engine for Emojis\\ \\ Computer Vision, Datasets, Plugins\\ \\ • \\ \\ Jan 4, 2024](https://voxel51.com/blog/how-to-build-a-semantic-search-engine-for-emojis) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-336-lllmstxt|> ## FiftyOne Polylines Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) Exploring Polylines – FiftyOne Tips and Tricks – September 22nd, 2023 Sep 22, 2023 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#361dcffa3ed5) [What is a Polyline?](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#ef724d200a3a) [Creating Polylines in Your Samples](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#02c5d594f5f3) [Classic Polylines](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#c095553414df) [Classic Polygon with Polylines](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#f70797537269) [Cuboids](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#aaf297cf605e) [Rotated Bounding Box](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#7e33e98b5d9d) [Conclusion](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#0d03407c105c) [Join the FiftyOne Community!](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#faa3770053ad) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#361dcffa3ed5) [What is a Polyline?](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#ef724d200a3a) [Creating Polylines in Your Samples](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#02c5d594f5f3) [Classic Polylines](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#c095553414df) [Classic Polygon with Polylines](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#f70797537269) [Cuboids](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#aaf297cf605e) [Rotated Bounding Box](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#7e33e98b5d9d) [Conclusion](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#0d03407c105c) [Join the FiftyOne Community!](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023#faa3770053ad) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we cover interesting workflows and features of FiftyOne! This week we are taking a look at [polyli](https://docs.voxel51.com/user_guide/using_datasets.html#polylines-and-polygons) [n](https://docs.voxel51.com/user_guide/using_datasets.html#polylines-and-polygons) [es](https://docs.voxel51.com/user_guide/using_datasets.html#polylines-and-polygons). We aim to cover the basics of creating a polyline and how they can be utilized to create specialized labels. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Settings QualityAuto SpeedNormal Picture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. Ok, let’s dive into this week’s tips and tricks! Also feel free to follow along in our [notebook](https://github.com/voxel51/fiftyone-examples/tree/polylines) or on [YouTube](https://www.youtube.com/@voxel51)! ## What is a Polyline? In computer vision, a polyline is a sequence of connected line segments used to represent and approximate the shape of an object or region within an image. Polylines are commonly employed in computer vision tasks such as image annotation, object tracking, and shape analysis. Polylines can be used to represent the trajectory of an object or map and route representations. Polylines can also be filled to represent different polygons’ regions of interest. With FiftyOnes implementation of polylines, labels can be as creative as you want! ## Creating Polylines in Your Samples ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3 4 5dataset = foz.load_zoo_dataset( 6 "quickstart", 7 dataset_name="polylines" 8) 9 10 11session = fo.launch_app(dataset) ``` Using the app, I am going to tag a non-busy picture for use in the example. To do so, simply select the image, click on the tag image and add “polylines” to its sample tags Let’s walk through some examples and how they can be used in various use cases. We grab our sample to draw on it next. ```python 1example_view = dataset.match_tags("polylines") 2sample = example_view.first() ``` ## Classic Polylines The most basic polyline is a line defined by its points. It could be pointing in the direction of a trajectory, or simply be a label for the divide between two regions of interest. To define a `Polyline` label, provide a list of points in (x,y) bounded by (0,1) for x and y. We also pass the argument `closed=False` to denote that we are _not_ closing the polyline, and `filled=False` to likewise express that we don’t want the polyline to be filled. ```python 1# A simple polyline 2polyline1 = fo.Polyline( 3 points=[[(0.3, 0.3), (0.7, 0.3), (0.7, 0.3)]], 4 closed=False, 5 filled=False, 6) 7 ``` ## Classic Polygon with Polylines Similar to the last example, we can define polygons with a set of points and close and fill the shape. Just like with any other label, we can add custom attributes to the polyline as well. In this case I will create a triangle and specify it as a right triangle. 💡 To close a shape, only the vertices are required and there is no need to specify the first vertex again. We can add these to our sample using `Polylines`, a group of polylines, and save our view to look at our new polylines. If it makes it easier to see, you can turn off seeing the `ground_truth` or `prediction` labels by unchecking their box under labels in the app. ```python 1# A closed, filled polygon with a label 2polyline2 = fo.Polyline( 3 label="triangle", 4 points=[[(0.1, 0.1), (0.3, 0.1), (0.3, 0.3)]], 5 closed=True, 6 filled=True, 7 kind="right", # custom attribute 8) 9 10 11sample["polylines"] = fo.Polylines(polylines=[polyline1, polyline2]) 12sample.save() 13example_view.save() 14session.view = example_view ``` ## Cuboids Cuboids, or 3D bounding boxes, are essential in computer vision for a range of applications. They serve to detect and localize objects in a three-dimensional environment, making them valuable in autonomous driving, object tracking, augmented reality, robotics, depth sensing, gesture recognition, human pose estimation, indoor navigation, industrial automation, and virtual reality. Cuboids help model the 3D extent of objects, enabling accurate perception, tracking, and interaction with the three-dimensional world in various domains. Today, we will be looking at how we can use polylines to create cuboids on our 2D image. FiftyOne does support [3D boxes for point clouds](https://docs.voxel51.com/user_guide/using_datasets.html#d-detections), but that will be covered another day. We define a function that will create a random cuboid label for our sample. Cuboids are defined with 8 points in the following sequence: ```python 1 7------------ 6 2 /| /| 3 / | / | 43-------- 2 | 5| 4----- |---- 5 6| / | / 7|/ | / 80-------- 1 9 ``` Using the [`fo.Polyline.from_cuboid()`](https://docs.voxel51.com/api/fiftyone.core.labels.html#fiftyone.core.labels.Polyline.from_cuboid) method, we are able to easily create a cuboid once we have our 8 points. ```python 1import numpy as np 2 3 4def random_cuboid(): 5 x0, y0 = [0, 0.2] + 0.8 * np.random.rand(2) 6 dx, dy = (min(0.8 - x0, y0 - 0.2)) * np.random.rand(2) 7 x1, y1 = x0 + dx, y0 - dy 8 w, h = (min(1 - x1, y1)) * np.random.rand(2) 9 front = [(x0, y0), (x0 + w, y0), (x0 + w, y0 - h), (x0, y0 - h)] 10 back = [(x1, y1), (x1 + w, y1), (x1 + w, y1 - h), (x1, y1 - h)] 11 return fo.Polyline.from_cuboid(front + back, label="cuboid") 12 13 14sample["polyline"] = random_cuboid() 15sample.save() 16example_view.save() 17session.view = example_view 18 19 20 ``` Here is what a finished result could look like: ## Rotated Bounding Box Rotated bounding boxes are essential for tasks like object detection, localization, and tracking, particularly when objects are not aligned with the standard horizontal and vertical axes. They're valuable in applications such as text detection in images, scene text recognition, and irregular object localization. These bounding boxes represent objects' orientations more accurately, allowing for precise positioning and analysis in situations where traditional axis-aligned boxes wouldn't suffice. To use a rotating bounding box in FiftyOne, we can use [`fo.Polyline.from_rotated_box()`](https://docs.voxel51.com/api/fiftyone.core.labels.html#fiftyone.core.labels.Polyline.from_rotated_box), where we provide the center of the box in xc,yc as well as the width, height, and the angle at which the box in rotated. ```python 1def random_rotated_box(): 2 xc, yc = 0.2 + 0.6 * np.random.rand(2) 3 w, h = 1.5 * (min(xc, yc, 1 - xc, 1 - yc)) * np.random.rand(2) 4 theta = 2 * np.pi * np.random.rand() 5 6 return fo.Polyline.from_rotated_box(xc, yc, w, h, theta, label="box") 7 8 9sample["polyline"] = random_rotated_box() 10sample.save() 11example_view.save() 12session.view = example_view 13 ``` A finished result I ran was this: ## Conclusion In conclusion, polylines, cuboids, and rotated bounding boxes serve as indispensable tools in computer vision, each tailored to address specific challenges and scenarios. Polylines enable intricate shape delineation, enhancing applications like image segmentation and contour tracking. Cuboids, or 3D bounding boxes, are invaluable for accurately capturing the spatial dimensions of objects, significantly improving object detection, tracking, and augmented reality tasks. Meanwhile, rotated bounding boxes cater to objects with non-standard orientations, greatly enhancing precision in tasks such as text detection and irregular object localization. These diverse labeling techniques empower computer vision professionals to more effectively and precisely address a wide array of real-world problems, facilitating better representation and understanding of complex visual data. There is no better place to implement these labels than FiftyOne! ## Join the FiftyOne Community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 2,000+ [FiftyOne Slack](https://slack.voxel51.com/) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 370+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [open source](https://voxel51.com/blog/tag/open-source) [polylines](https://voxel51.com/blog/tag/polylines) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/79d00d175a8098516cb2f4a7711131cbf322d01a-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Finding and Correcting Mistakes – FiftyOne Tips and Tricks – Aug 18, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Aug 18, 2023](https://voxel51.com/blog/finding-and-correcting-mistakes-fiftyone-tips-and-tricks-aug-18-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/99e6a836a21e272da945f0e045f4d4812c1d3583-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Exploring the CLI – FiftyOne Tips and Tricks – Aug 25th, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Aug 26, 2023](https://voxel51.com/blog/exploring-the-cli-fiftyone-tips-and-tricks-aug-25th-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/395ba1a1dacb511782456902b224aa4fa8552dc0-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Understanding Grouped Datasets – FiftyOne Tips and Tricks – Sep 1, 2023\\ \\ Tips & Tricks\\ \\ • \\ \\ Sep 1, 2023](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-337-lllmstxt|> ## ICCV 2023 Survival Guide [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) ICCV 2023 Survival Guide Sep 27, 2023 • 18 min read Article content In this article [10 Computer Vision Papers You Won't Want to Miss](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#d2148424a063) [DEVA: Tracking Anything with Decoupled Video Segmentation](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#89fa82b89092) [Effective Whole-body Pose Estimation with Two-stages Distillation](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#79e88809061d) [EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#1b1c649368e6) [FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#545a7fd0cd88) [LightGlue: Local Feature Matching at Light Speed](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#b997ceb7e73d) [ProPainter: Improving Propagation and Transformer for Video Inpainting](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#416244e3f1d3) [Segment Anything](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#48fc0a1c23dc) [Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#4670371a6ce1) [Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#d38775c7c47f) [ViperGPT: Visual Inference via Python Execution for Reasoning](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#c4b54aa3b594) [Speed Round](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#4bc989747053) [A Generalist Framework for Panoptic Segmentation of Images and Videos](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#521caff5dc85) [Chupa : Carving 3D Clothed Humans from Skinned Shape Priors using 2D Diffusion Probabilistic Models](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#edc02f36f44c) [Dense Text-to-Image Generation with Attention Modulation](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#f0e6c143d9af) [DOLCE: A Model-Based Probabilistic Diffusion Framework for Limited-Angle CT Reconstruction](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#90a9c823e687) [Doppelgangers: Learning to Disambiguate Images of Similar Structures](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#b090d33b60c3) [Equivariant Similarity for Vision-Language Foundation Models](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#ab195922acf8) [FACET: Fairness in Computer Vision Evaluation Benchmark](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#b4c83a9cf9e7) [GlueStick: Robust Image Matching by Sticking Points and Lines Together](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#e9de12c9d986) [HyperDiffusion: Generating Implicit Neural Fields with Weight-Space Diffusion](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#ed8e9c71f57a) [Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#c5f14bc327dd) [Neural Haircut: Prior-Guided Strand-Based Hair Reconstruction](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#1e6313d916db) [Prompt-aligned Gradient for Prompt Tuning](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#21414fb783bf) [Reference-guided Controllable Inpainting of Neural Radiance Fields](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#22985e1fe34f) [ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#a7ac5527a3ce) [Robo3D: Towards Robust and Reliable 3D Perception against Corruptions](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#62e41f2d3f77) [ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#a12f9fc0786b) [SegGPT: Segmenting Everything In Context](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#e6a28e776f39) [SHIFT3D: Synthesizing Hard Inputs For Tricking 3D Detectors](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#ac71504e6051) [Text2Performer: Text-Driven Human Video Generation](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#0ec006a90aea) [Tracking Everything Everywhere All at Once](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#c011d50734e6) In this article [10 Computer Vision Papers You Won't Want to Miss](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#d2148424a063) [DEVA: Tracking Anything with Decoupled Video Segmentation](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#89fa82b89092) [Effective Whole-body Pose Estimation with Two-stages Distillation](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#79e88809061d) [EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#1b1c649368e6) [FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#545a7fd0cd88) [LightGlue: Local Feature Matching at Light Speed](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#b997ceb7e73d) [ProPainter: Improving Propagation and Transformer for Video Inpainting](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#416244e3f1d3) [Segment Anything](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#48fc0a1c23dc) [Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#4670371a6ce1) [Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#d38775c7c47f) [ViperGPT: Visual Inference via Python Execution for Reasoning](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#c4b54aa3b594) [Speed Round](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#4bc989747053) [A Generalist Framework for Panoptic Segmentation of Images and Videos](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#521caff5dc85) [Chupa : Carving 3D Clothed Humans from Skinned Shape Priors using 2D Diffusion Probabilistic Models](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#edc02f36f44c) [Dense Text-to-Image Generation with Attention Modulation](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#f0e6c143d9af) [DOLCE: A Model-Based Probabilistic Diffusion Framework for Limited-Angle CT Reconstruction](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#90a9c823e687) [Doppelgangers: Learning to Disambiguate Images of Similar Structures](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#b090d33b60c3) [Equivariant Similarity for Vision-Language Foundation Models](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#ab195922acf8) [FACET: Fairness in Computer Vision Evaluation Benchmark](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#b4c83a9cf9e7) [GlueStick: Robust Image Matching by Sticking Points and Lines Together](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#e9de12c9d986) [HyperDiffusion: Generating Implicit Neural Fields with Weight-Space Diffusion](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#ed8e9c71f57a) [Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#c5f14bc327dd) [Neural Haircut: Prior-Guided Strand-Based Hair Reconstruction](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#1e6313d916db) [Prompt-aligned Gradient for Prompt Tuning](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#21414fb783bf) [Reference-guided Controllable Inpainting of Neural Radiance Fields](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#22985e1fe34f) [ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#a7ac5527a3ce) [Robo3D: Towards Robust and Reliable 3D Perception against Corruptions](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#62e41f2d3f77) [ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#a12f9fc0786b) [SegGPT: Segmenting Everything In Context](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#e6a28e776f39) [SHIFT3D: Synthesizing Hard Inputs For Tricking 3D Detectors](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#ac71504e6051) [Text2Performer: Text-Driven Human Video Generation](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#0ec006a90aea) [Tracking Everything Everywhere All at Once](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#c011d50734e6) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## 10 Computer Vision Papers You Won't Want to Miss The annual [IEEE/CVF International Conference on Computer Vision (ICCV)](https://iccv2023.thecvf.com/) is just a week away! The Conference, which is set to take place on October 4th-6th in Paris, France, is shaping up to be something special. Ah, Paris! The city of love, croissants, and—perhaps romantic strolls along the Seine with fervent debates on transformer models? While some folks are trying to capture the perfect Eiffel Tower selfie, we're here capturing multi-dimensional arrays of pixel intensity values. Talk about capturing the moment, n'est-ce pas? In the land where Impressionism was born, it's only fitting that we gather to discuss how computers perceive the world. And with more than 1,000 papers being presented, there’s a lot to discuss! But have no fear! We've done the heavy lifting (and scrolling, and skimming) for you. We've combed through every paper to bring you the crème de la crème of ICCV 2023 so you can focus on what really matters, whether that's finding the right session to spark your next big idea or sipping on a fine Bordeaux as you ponder the intricacies of object detection. Without further ado, here are our top ten picks in alphabetical order: 01. [DEVA: Tracking Anything with Decoupled Video Segmentation](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#DEVA) 02. [Effective Whole-body Pose Estimation with Two-stages Distillation](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#Pose) 03. [EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#ViT) 04. [FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#FastViT) 05. [LightGlue: Local Feature Matching at Light Speed](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#LightGlue) 06. [ProPainter: Improving Propagation and Transformer for Video Inpainting](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#ProPainter) 07. [Segment Anything](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#SAM) 08. [Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#Text2Room) 09. [Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#Text2Video) 10. [ViperGPT: Visual Inference via Python Execution for Reasoning](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide#ViperGPT) ## DEVA: Tracking Anything with Decoupled Video Segmentation _Original source: DEVA Paper_ - Links: ( [Arxiv](https://arxiv.org/abs/2309.03903) \| [Code](https://github.com/hkchengrex/Tracking-Anything-with-DEVA/tree/main) \| [Project Page](https://hkchengrex.github.io/Tracking-Anything-with-DEVA/)) - Authors: Ho Kei Cheng, Seoung Wug Oh, Brian Price, Alexander Schwing, Joon-Young Lee **TL;DR:** DEVA extends Segment Anything (see below!) to video for open-world video segmentation Traditionally, training video segmentation models has been a time-intensive and costly undertaking. In large part, this owes to the fact that end-to-end video segmentation models require densely annotated video data, which means segmentation masks for each frame. While this has not proved prohibitive for _specific_ video segmentation tasks, it has meant that equivalent efforts need to be made for each desired task. In other words, the time spent densely annotating a dataset for video object segmentation would not translate to a video panoptic segmentation task, which would require its own annotation. Decoupled Video segmentation flips the script by decoupling video segmentation into first, a task-specific image-level segmentation task, and second, a task-agnostic bi-directional temporal propagation. In plain English, this means you take an image segmentation model and apply it to the frames in your video. The “hypothesis” segmentation masks from different frames are then fused in a way that is coherent. This temporal propagation module only needs to be trained once, and can then be used in conjunction with any image segmentation model — even open vocabulary models like Segment Anything! ## Effective Whole-body Pose Estimation with Two-stages Distillation _Original source: DW2Pose Paper_ - Links: ( [Arxiv](https://arxiv.org/abs/2307.15880) \| [Code](https://github.com/IDEA-Research/DWPose)) - Authors: Zhendong Yang, Ailing Zeng, Chun Yuan, Yu Li **TL;DR:** State of the art 2D Human Pose Estimation [Pose estimation](https://paperswithcode.com/task/pose-estimation) is the task of identifying, localizing, and orienting various key points on a person’s body. The “whole body” variant of the task involves simultaneous detection and localization of keypoints from head to toe, and is a crucial component in many downstream computer vision applications, from synthetic motion generation to tracking humans in virtual or augmented reality environments. While there exist very accurate whole body pose estimation like [RTMPose](https://github.com/open-mmlab/mmpose/tree/dev-1.x/projects/rtmpose), one of the current challenges in the space is achieving similar levels of accuracy while operating at much lower latency. In this paper, the authors present DWPose, a two-stage [knowledge distillation](https://en.wikipedia.org/wiki/Knowledge_distillation) approach for whole body pose estimation. First, the student model learns from a larger teacher model. Then the student’s backbone is frozen, and its head is updated — it learns from itself! All told, DWPose is able to achieve state-of-the-art performance on the [COCO-WholeBody benchmark](https://github.com/jin-s13/COCO-WholeBody). ## EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction _Original source: EfficientViT Paper_ - Links: ( [Arxiv](https://arxiv.org/abs/2305.07027) \| [Code](https://github.com/mit-han-lab/efficientvit)) - Authors: Xinyu Liu, Houwen Peng, Ningxin Zheng, Yuqing Yang, Han Hu, Yixuan Yuan **TL;DR:** Improved memory efficiency and reduced redundancy 👉 blazing fast vision transformers [Transformer models](https://en.wikipedia.org/wiki/Transformer_(machine_learning_model)) are all the rage these days, and the [Vision Transformer](https://arxiv.org/abs/2010.11929) (ViT) has taken over the top spot in many areas of computer vision. One of the downsides of the standard vision transformer, however, is the heavy computational cost incurred during inference. For the most part, this has thus far prevented their adoption in real-time computer vision applications. That may be about to change. With EfficientViT, the researchers from The Chinese University of Hong Kong and Microsoft Research make two welcome improvements to the ViT architecture. First, they reduce communication time between feature channels by swapping out some memory-bound self-attention layers with far more memory-efficient feed-forward network (FFN) layers. Second, with a new Cascaded Group Attention (CGA), they minimize redundant computations happening at multiple attention heads. Combined with some clever reallocation of network parameters, they are able to take ViTs into new territory. As just one example, EfficientViT-L0 for Segment Anything can process 1000+ images per second on an A100 GPU, more than 3x that of [MobileSAM](https://github.com/ChaoningZhang/MobileSAM), while achieving [mean Intersection over Union](https://analyticsindiamag.com/a-deep-dive-into-meaniou-an-evaluation-metric-for-object-detection/#:~:text=MeanIoU%20calculates%20the%20ratio%20of,is%20eligible%20for%20this%20metric.) (mIoU). ## FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization _Original source: FastViT Paper_ - Links: ( [Arxiv](https://arxiv.org/abs/2303.14189) \| [Code](https://github.com/apple/ml-fastvit)) - Authors: Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel, Anurag Ranjan **TL;DR:** Removing skip connections with structural reparameterization in hybrid transformer model 👉 Robust models with state-of-the-art tradeoff between accuracy and latency With FastViT, Apple researchers set out to solve a similar problem as EfficientViT: the high computational cost of vision transformers. Their approach, however, is based on maximizing the strengths of the up-and-coming vision transformer models and convolutional neural networks of old to form a fast _hybrid_ transformer. In designing FastViT, the team combined three central principles. First, they reduce [skip connections](https://www.analyticssteps.com/blogs/what-are-skip-connections-neural-networks), whose high memory access contributes significantly to inference latency. Second, they factorize dense convolutions, reducing the number of parameters, and increase their representational “capacity” using a technique called _overparameterization_. Finally, they replace some of the computationally-intensive self-attention layers at early stages with convolutional kernels. Altogether, these design decisions result in an architecture that outperforms competitive models in accuracy and latency on tasks from image classification to 3D mesh regression. ## LightGlue: Local Feature Matching at Light Speed _Original source: LightGlue Paper_ - Links: ( [Arxiv](https://arxiv.org/abs/2306.13643) \| [Code](https://github.com/cvg/LightGlue)) - Authors: Philipp Lindenberger, Paul-Edouard Sarlin, Marc Pollefeys **TL;DR:** Faster, more accurate sparse feature matching that adapts to the problem difficulty [Image matching](https://paperswithcode.com/task/image-matching) is the task of associating select points in one image with points in another image, so as to “align” the two. It is used in a variety of downstream applications, from camera tracking to 3D reconstruction. At CVPR 2020, Magic Leap unveiled [SuperGlue](https://arxiv.org/abs/1911.11763), a transformer-based graph neural network which met resounding success at both matching images from sparse sets of points, and rejecting outliers — instances where the two images are not drawn from the same scene. However, the computational intensity of the transformer-based model prevented its utilization in low-latency applications. (If you’re noticing a trend, you’re not alone!) LightGlue revisits some of SuperGlue’s design decisions, making changes which collectively improve memory, computation, accuracy, and trainability. The model simultaneously computes a “matchability score”, for how confident it is that the point in one image can be matched to a point in the second image, and a “pairwise similarity” for pairs of points. My favorite part: the model includes a classifier that allows it to stop the matching process if it is highly confident. This means that the easier the matching job, the quicker the matching process will terminate. ## ProPainter: Improving Propagation and Transformer for Video Inpainting _Original source: ProPainter Project Page_ - Links: ( [Arxiv](https://arxiv.org/abs/2309.03897) \| [Code](https://github.com/sczhou/ProPainter) \| [Project Page](https://shangchenzhou.com/projects/ProPainter/)) - Authors: Shangchen Zhou, Chongyi Li, Kelvin C.K. Chan, Chen Change Loy **TL;DR:** Faster, more computationally efficient video inpainting by dual propagation (both features and images) Much like video segmentation, [video inpainting](https://towardsdatascience.com/deep-video-inpainting-756e60ddcaaf) is an extension of an image-based task (in this case [inpainting](https://paperswithcode.com/task/image-inpainting)) to videos, in a spatially and temporally coherent and consistent manner. Approaches to video inpainting typically fall into two buckets: image propagation and feature propagation. Image propagation on its own can result in “unpleasant artifacts” and “texture misalignment”. On the other hand, feature propagation, which uses a Transformer architecture, is typically limited to short sequences and lower resolution videos as a result of the Transformer’s memory and compute constraints. At this point, I know I sound like a broken record! ProPainter (ProPagation and an efficient Transformer) combines the strengths of image propagation and feature propagation to achieve _dual-domain propagation_. The implementation of the model involves a bevy of improvements, from turning CPU-intensive processes into GPU computations, an efficient recurrent neural network (RNN) for completing the flows, and discarding unnecessary/redundant segments of the Transformer’s query and key/value spaces. All told, the model significantly outperforms the prior state of the art in peak signal to noise ratio (PSNR) while at the same time reducing memory consumption! ## Segment Anything _Original source: Segment Anything Model GitHub repo_ - Links: ( [Arxiv](https://arxiv.org/abs/2304.02643) \| [Code](https://github.com/facebookresearch/segment-anything) \| [Project Page](https://segment-anything.com/) \| [Tutorial](https://medium.com/towards-data-science/see-what-you-sam-4eea9ad9a5de?source=your_stories_page-------------------------------------)) - Authors: Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, Ross Girshick **TL;DR:** Game changing foundation vision model for prompted and unprompted segmentation tasks, plus largest-ever segmentation dataset With Segment Anything (SAM), Meta AI has taken image segmentation to the next level. Open source and easy to use, SAM brings high quality zero-shot segmentation to a wide domain. The Vision Transformer-based model, which comes in three sizes, supports multiple _modes_ of segmentation. In automatic segmentation mode, it will generate predicted segmentation masks for all things (distinct entities) and stuff (concepts like “sky”). Alternatively, you can prompt the model with bounding boxes and/or key points! To generate SAM, the team constructed [SA-1B](https://ai.meta.com/datasets/segment-anything/), the largest segmentation dataset to date, consisting of 1B masks across 11M images. The annotation of the dataset and the training of the model were performed in a looping process: the model was used to assist human annotators, and these annotations were used to retrain the model! It has already sired a slew of smaller segmentation models like [FastSAM](https://github.com/CASIA-IVA-Lab/FastSAM), [MobileSAM](https://github.com/ChaoningZhang/MobileSAM), and [NanoSAM](https://github.com/NVIDIA-AI-IOT/nanosam), via distillation or by virtue of the SA-1B dataset, and has inspired a seemingly endless parade of related “ Anything” models, from [Track Anything](https://github.com/gaomingqi/Track-Anything) to [Inpaint Anything](https://github.com/geekyutao/Inpaint-Anything). Meta AI’s FACET benchmark (see below) is also adapted from a subset of SA-1B. 💡SAM is [natively supported](https://docs.voxel51.com/user_guide/model_zoo/models.html#segment-anything-vitb-torch) by the computer vision library FiftyOne, making segmenting your data easier than ever! ## Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models _Original source: Text2Room Project Page_ - Links: ( [Arxiv](https://arxiv.org/abs/2303.11989) \| [Code](https://github.com/lukasHoel/text2room) \| [Project Page](https://lukashoel.github.io/text-to-room/)) - Authors: Lukas Höllein, Ang Cao, Andrew Owens, Justin Johnson, Matthias Nießner **TL;DR:** 2D images from diffusion models ➕ rendering from novel viewpoint ➕ fusing the images 👉 room-scale 3D meshes Mesh representations of three-dimensional scenes are useful in computer graphics, and in the creation of 3D assets for augmented and virtual reality environments. Generating these 3D meshes, however, one runs into multiple challenges. Chief among these challenges is the lack of availability of high-quality 3D data to train on. Due to the heightened costs and longer times involved in collecting 3D data, the datasets are typically smaller, or primarily consist of simple objects and scenes. And while [neural radiance fields](https://www.matthewtancik.com/nerf) (NeRFs) hold promise for 3D scene generation, extending them to room-level scales presents its own difficulties. Text2Room bypasses these problems by harnessing the 2D image generation capabilities of text-to-image diffusion models and cleverly combining the 2D images into realistic 3D scenes. The method starts by generating a single image of a scene. [Monocular depth estimation](https://paperswithcode.com/task/monocular-depth-estimation) is used to backproject the scene into three dimensions, and an initial 3D mesh is generated from this. From there, the mesh is iteratively rendered from novel viewpoints, any holes are inpainted, and the images are fused together. In tests, Text2Room outperformed competitors on a slate of quantitative metrics, from perceptual quality (PQ) to 3D structure completeness (3DS). ## Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators _Original source: Text2Video-Zero Paper_ - Links: ( [Arxiv](https://arxiv.org/abs/2303.13439) \| [Code](https://github.com/Picsart-AI-Research/Text2Video-Zero) \| [Hugging Face](https://huggingface.co/docs/diffusers/api/pipelines/text_to_video_zero) \| [Project Page](https://text2video-zero.github.io/)) - Authors: Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, Humphrey Shi **TL;DR:** Motion dynamics ➕ cross-frame attention 👉 off-the-shelf text-to-image models can be applied to generate videos Whereas Text2Room leverages the power of 2D text-to-image models to generate high-quality 3D meshes, Text2Video-Zero harnesses these same 2D models to perform low-cost zero-shot text-to-video generation. That is, Text2Video-Zero sets forth an approach to video generation from text prompts that does not necessitate fine-tuning or optimization. Instead of randomly sampling latent codes (the inputs to diffusion models) for each frame independently, the latent codes are constructed through an iterative warping process. This incentivizes temporal consistency across the generated frames, but on its own is still an insufficient constraint. On top of this, Text2Video replaces the self-attention layers in the diffusion model with cross-frame attention layers connecting each frame to the first frame. The most exciting part of Text2Video-Zero is that the basic approach also works for other video tasks such as conditional video generation and even instruction-guided video editing! You can [run the model](https://huggingface.co/docs/diffusers/api/pipelines/text_to_video_zero) in text-to-video mode, or in either of these two additional modes with Hugging Face’s diffusers library! ## ViperGPT: Visual Inference via Python Execution for Reasoning _Original source: ViperGPT GitHub repo_ - Links: ( [Arxiv](https://arxiv.org/abs/2303.08128) \| [Code](https://github.com/cvlab-columbia/viper) \| [Project Page](https://viper.cs.columbia.edu/)) - Authors: Dídac Surís, Sachit Menon, Carl Vondrick **TL;DR:** Using code generation models to compose vision-language models 👉 state-of-the-art performance on visual inference tasks ViperGPT (so-named because it executes Python code) takes a modular approach to visual inference. Instead of relying on end-to-end models, which bundle visual processing and reasoning but lack in interpretability, ViperGPT uses GPT-3 Codex to generate Python code that executes specific subroutines to answer a query. In the gif above, for instance, ViperGPT can employ object detection models to determine how many muffins there are, and then use this result to reason about the right answer to the query. By defining simple, task-specific APIs, ViperGPT is able to leverage the code generation and reasoning capabilities of existing models without fine-tuning. Additionally, the output from the code-generation model is, well, _code_, it is more interpretable than end-to-end models. To top things off, ViperGPT achieves state of the art performance on multiple zero shot tasks, including general visual question answering, grounded question answering, and even referring expression tasks! 💡If you like ViperGPT, you should check out: - [HuggingGPT](https://arxiv.org/abs/2303.17580): LLM dispatcher for CV tasks - [VisProg](https://github.com/allenai/visprog): Visual reasoning without training (CVPR 2023 best paper) - [VoxelGPT](https://github.com/voxel51/voxelgpt/tree/main/src): Text-to-query for CV datasets ## Speed Round Here are twenty more cool ICCV 2023 projects you should check out! ## A Generalist Framework for Panoptic Segmentation of Images and Videos - Links: ( [Arxiv](https://arxiv.org/abs/2210.06366)) - Authors: Ting Chen, Lala Li, Saurabh Saxena, Geoffrey Hinton, David J. Fleet **TL;DR:** Google Research applies [Geoffrey Hinton](https://en.wikipedia.org/wiki/Geoffrey_Hinton) et al.’s [Bit Diffusion](https://arxiv.org/pdf/2208.04202.pdf) model to [panoptic segmentation](https://paperswithcode.com/task/panoptic-segmentation) 👉more general architecture and loss function competitive with specialized methods ## Chupa : Carving 3D Clothed Humans from Skinned Shape Priors using 2D Diffusion Probabilistic Models - Links: ( [Arxiv](https://arxiv.org/abs/2305.11870) \| [Code](https://github.com/snuvclab/chupa) \| [Project Page](https://snuvclab.github.io/chupa/)) - Authors: Byungjun Kim, Patrick Kwon, Kwangho Lee, Myunggi Lee, Sookwan Han, Daesik Kim, Hanbyul Joo **TL;DR:** Decomposing the task of 3D human mesh generation into (1) pose-conditional normal map generation with diffusion models, and (2) carving the 3D mesh utilizing the normal maps 👉 realistic 3D human human meshes ## Dense Text-to-Image Generation with Attention Modulation - Links: ( [Arxiv](https://arxiv.org/abs/2308.12964) \| [Code](https://github.com/naver-ai/DenseDiffusion) \| [Hugging Face Demo](https://huggingface.co/spaces/naver-ai/DenseDiffusion)) - Authors: Yunji Kim, Jiyoung Lee, Jin-Hwa Kim, Jung-Woo Ha, Jun-Yan Zhu **TL;DR:** Layout of objects in images generated by diffusion models is related to the model’s attention and cross-attention maps. Modulating these maps 👉 better layout control and better performance for dense text prompts ## DOLCE: A Model-Based Probabilistic Diffusion Framework for Limited-Angle CT Reconstruction - Links: ( [Arxiv](https://arxiv.org/abs/2211.12340) \| [Project Page](https://wustl-cig.github.io/dolcewww/)) - Authors: Jiaming Liu, Rushil Anirudh, Jayaraman J. Thiagarajan, Stewart He, K. Aditya Mohan, Ulugbek S. Kamilov, Hyojin Kim **TL;DR:** Conditional diffusion models can reduce artifacts in CT images reconstructed from severe undersampling scenarios ## Doppelgangers: Learning to Disambiguate Images of Similar Structures - Links: ( [Arxiv](https://arxiv.org/abs/2309.02420) \| [Code](https://github.com/RuojinCai/doppelgangers) \| [Dataset](https://github.com/RuojinCai/doppelgangers/blob/main/data/doppelgangers_dataset/README.md) \| [Project Page](https://doppelgangers-3d.github.io/)) - Authors: Ruojin Cai, Joseph Tung, Qianqian Wang, Hadar Averbuch-Elor, Bharath Hariharan, Noah Snavely **TL;DR:** The task of distinguishing images of the same or different 3D surfaces (visual disambiguation) is treated as binary classification on pairs of images. This is enabled by a new Doppelgangers dataset ## Equivariant Similarity for Vision-Language Foundation Models - Links: ( [Arxiv](https://arxiv.org/abs/2303.14465) \| [Code](https://github.com/Wangt-CN/EqBen)) - Authors: Tan Wang, Kevin Lin, Linjie Li, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, Lijuan Wang **TL;DR:** New benchmark [EqBen](https://github.com/Wangt-CN/EqBen#eqben) for assessing how faithfully the notion of similarity holds up in vision-language model under semantic changes ## FACET: Fairness in Computer Vision Evaluation Benchmark - Links: ( [Dataset](https://ai.meta.com/datasets/facet-downloads/) \| [Paper](https://scontent-den4-1.xx.fbcdn.net/v/t39.2365-6/10000000_1046408126367357_8248945695912965309_n.pdf?_nc_cat=111&ccb=1-7&_nc_sid=3c67a6&_nc_ohc=tstBOo3SFzIAX_92HsK&_nc_ht=scontent-den4-1.xx&oh=00_AfBnBFK10yVZuJOyXBIuGiScQWiV3mCMB4QR5y2O1csCrA&oe=651794F8) \| [Project Page](https://facet.metademolab.com/) \| [Tutorial](https://medium.com/voxel51/facet-a-benchmark-dataset-for-fairness-in-computer-vision-2260c82e1662)) - Authors: Laura Gustafson, Chloe Rolland, Nikhila Ravi, Quentin Duval, Aaron Adcock, Cheng-Yang Fu, Melissa Hall, Candace Ross **TL;DR:** Meta AI releases a diverse benchmark dataset for evaluating fairness, bias, and disparity across protected attributes ## GlueStick: Robust Image Matching by Sticking Points and Lines Together - Links: ( [Arxiv](https://arxiv.org/abs/2304.02008) \| [Code](https://github.com/cvg/GlueStick) \| [Project Page](https://iago-suarez.com/gluestick/)) - Authors: Rémi Pautrat, Iago Suárez, Yifan Yu, Marc Pollefeys, Viktor Larsson **TL;DR:** Treating points, lines, and their descriptors as combined “wireframes” and applying graph neural nets 👉state-of-the-art matching for both points and line segments ## HyperDiffusion: Generating Implicit Neural Fields with Weight-Space Diffusion - Links: ( [Arxiv](https://arxiv.org/abs/2303.17015) \| [Code](https://github.com/Rgtemze/HyperDiffusion) \| [Project Page](https://ziyaerkoc.com/hyperdiffusion/)) - Authors: Ziya Erkoç, Fangchang Ma, Qi Shan, Matthias Nießner, Angela Dai **TL;DR:** Overfitting multi-layer perceptrons (MLPs) on individual neural implicit fields and training a diffusion model to denoise these weights 👉realistic 3D shapes and 4D mesh animations ## Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing - Links: ( [Arxiv](https://arxiv.org/abs/2304.02051) \| [Code](https://github.com/aimagelab/multimodal-garment-designer)) - Authors: Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia, Marco Bertini, Rita Cucchiara **TL;DR:** Denoising diffusion network conditioned on multiple modalities and trained on multimodal extensions to Dress Code and VITON-HD 👉 fashion image editing that can be guided by text prompts, poses, or sketches ## Neural Haircut: Prior-Guided Strand-Based Hair Reconstruction - Links: ( [Arxiv](https://arxiv.org/abs/2306.05872) \| [Code](https://github.com/SamsungLabs/NeuralHaircut) \| [Project Page](https://samsunglabs.github.io/NeuralHaircut/)) - Authors: Vanessa Sklyarova, Jenya Chelishev, Andreea Dogaru, Igor Medvedev, Victor Lempitsky, Egor Zakharov **TL;DR:** Samsung researchers create two-stage method for accurately reconstructing hair at the strand level from monocular videos or multiview images ## Prompt-aligned Gradient for Prompt Tuning - Links: ( [Arxiv](https://arxiv.org/abs/2205.14865) \| [Code](https://github.com/BeierZhu/Prompt-align)) - Authors: Beier Zhu, Yulei Niu, Yucheng Han, Yue Wu, Hanwang Zhang **TL;DR:** Principled approach to fine-tuning the similarity measure in vision-language models without forgetting general knowledge, ProGrad, shows strong few-shot generalization ## Reference-guided Controllable Inpainting of Neural Radiance Fields - Links: ( [Arxiv](https://arxiv.org/abs/2304.09677) \| [Project Page](https://ashmrz.github.io/reference-guided-3d/)) - Authors: Ashkan Mirzaei, Tristan Aumentado-Armstrong, Marcus A. Brubaker, Jonathan Kelly, Alex Levinshtein, Konstantinos G. Derpanis, Igor Gilitschenski **TL;DR:** Monocular depth estimation to [back-project](https://www.sciencedirect.com/topics/engineering/backprojection) an image to 3D coordinates ➕ new rendering technique 👉 consistent NeRF inpainting from original NeRF and a single inpainted image ## ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model - Links: ( [Arxiv](https://arxiv.org/abs/2304.01116) \| [Code](https://github.com/mingyuan-zhang/ReMoDiffuse) \| [Project Page](https://mingyuan-zhang.github.io/projects/ReMoDiffuse.html)) - Authors: Mingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai, Fangzhou Hong, Huirong Li, Lei Yang, Ziwei Liu **TL;DR:** Retrieving reference motion sequences that are semantically and kinematically relevant and selectively incorporating this knowledge into the motion generation process 👉 diverse motion with state-of-the-art motion quality and text-motion consistency ## Robo3D: Towards Robust and Reliable 3D Perception against Corruptions - Links: ( [Arxiv](https://arxiv.org/abs/2303.17597) \| [Code](https://github.com/ldkong1205/Robo3D) \| [Project Page](https://ldkong.com/Robo3D)) - Authors: Lingdong Kong, Youquan Liu, Xin Li, Runnan Chen, Wenwei Zhang, Jiawei Ren, Liang Pan, Kai Chen, Ziwei Liu **TL;DR:** Benchmark evaluation suite for detection and segmentation models in out-of-distribution autonomous driving scenarios ## ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes - Links: ( [Arxiv](https://arxiv.org/abs/2308.11417) \| [Project Page](https://cy94.github.io/scannetpp/)) - Authors: Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela Dai **TL;DR:** Higher-resolution successor to the popular [ScanNet dataset](https://paperswithcode.com/dataset/scannet), captured at 33-millimeter resolution, designed for novel view synthesis of indoor scenes ## SegGPT: Segmenting Everything In Context - Links: ( [Arxiv](https://arxiv.org/abs/2304.03284) \| [Code](https://github.com/baaivision/Painter) \| [Hugging Face Demo](https://huggingface.co/spaces/BAAI/SegGPT)) - Authors: Xinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang, Chunhua Shen, Tiejun Huang **TL;DR:** Generalist vision model built on top of [Painter](https://arxiv.org/abs/2212.02499) that is capable of many segmentation tasks, including semantic segmentation, video object segmentation, and panoptic segmentation, but does not set out to achieve state-of-the-art performance on any single task ## SHIFT3D: Synthesizing Hard Inputs For Tricking 3D Detectors - Links: ( [Arxiv](https://arxiv.org/abs/2309.05810v1)) - Authors: Hongge Chen, Zhao Chen, Gregory P. Meyer, Dennis Park, Carl Vondrick, Ashish Shrivastava, Yuning Chai **TL;DR:** New approach to generating plausible yet challenging 3D shapes in order to probe the failure modes of 3D object detectors for autonomous driving and other mission-critical tasks ## Text2Performer: Text-Driven Human Video Generation - Links: ( [Arxiv](https://arxiv.org/abs/2304.08483) \| [Code](https://github.com/yumingj/Text2Performer) \| [Project Page](https://yumingj.github.io/projects/Text2Performer.html)) - Authors: Yuming Jiang, Shuai Yang, Tong Liang Koh, Wayne Wu, Chen Change Loy, Ziwei Liu **TL;DR:** Generate temporally coherent human-centered videos from text input by decomposing into appearance (share across all frames) and pose (continuously changing) representations, and predicting pose embeddings with a diffusion model ## Tracking Everything Everywhere All at Once - Links: ( [Arxiv](https://arxiv.org/abs/2306.05422) \| [Code](https://github.com/qianqianwang68/omnimotion) \| [Project Page](https://omnimotion.github.io/)) - Authors: Qianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li, Bharath Hariharan, Aleksander Holynski, Noah Snavely **TL;DR:** Generate temporally coherent human-centered videos from text input by decomposing into appearance (share across all frames) and pose (continuously changing) representations, and predicting pose embeddings with a diffusion model. [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [iccv](https://voxel51.com/blog/tag/iccv) [image dataset](https://voxel51.com/blog/tag/image-dataset) [open source](https://voxel51.com/blog/tag/open-source) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/622b7369c791083b44e3034b2b8772d3ecada8bb-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – November 2023\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Nov 1, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-november-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/869a02098d1898869a250f4a5a23648c479af7da-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – April ‘23\\ \\ Product & News\\ \\ • \\ \\ Apr 6, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-april-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/338b38d41e6072dd11af86f21f5309337c52f36b-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – May ‘23\\ \\ Product & News\\ \\ • \\ \\ May 5, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-may-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-338-lllmstxt|> ## Zero-Shot Prediction Plugin [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Plugins](https://voxel51.com/blog/category/plugins), [Tutorials](https://voxel51.com/blog/category/tutorials) Zero-Shot Prediction Plugin for FiftyOne Sep 28, 2023 • 7 min read Article content In this article [Pre-label your computer vision data with CLIP, SAM, and other zero-shot models!](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#920f39299206) [Zero-Shot Prediction 0️⃣🎯🔮](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#a01968114228) [Plugin Overview & Functionality](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#7fe3e3b79c9d) [Delegating Execution](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#e5f3f996e31f) [Setting Inference Target](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#ade8f12c1562) [Zero-Shot Prediction in Action](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#f947cbe7bb06) [Installing the Plugin](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#d3c16f9289a9) [Lessons Learned](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#c9565775c5ef) [Keep up with FiftyOne’s Plugin Features](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#25a3ef5cbfd4) [Reducing Boilerplate Code](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#e4a176e24b64) [Conclusion](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#60bb39758abd) In this article [Pre-label your computer vision data with CLIP, SAM, and other zero-shot models!](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#920f39299206) [Zero-Shot Prediction 0️⃣🎯🔮](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#a01968114228) [Plugin Overview & Functionality](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#7fe3e3b79c9d) [Delegating Execution](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#e5f3f996e31f) [Setting Inference Target](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#ade8f12c1562) [Zero-Shot Prediction in Action](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#f947cbe7bb06) [Installing the Plugin](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#d3c16f9289a9) [Lessons Learned](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#c9565775c5ef) [Keep up with FiftyOne’s Plugin Features](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#25a3ef5cbfd4) [Reducing Boilerplate Code](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#e4a176e24b64) [Conclusion](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone#60bb39758abd) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Pre-label your computer vision data with CLIP, SAM, and other zero-shot models! Welcome to week six of _Ten Weeks of Plugins_. During these ten weeks, we will be building a FiftyOne Plugin (or multiple!) each week and sharing the lessons learned! If you’re new to them, FiftyOne Plugins provide a flexible mechanism for anyone to extend the functionality of their FiftyOne App. You may find the following resources helpful: - [FiftyOne Plugins Repo](https://github.com/voxel51/fiftyone-plugins) - [FiftyOne Plugin Docs](https://docs.voxel51.com/plugins/index.html#downloading-plugins) - Plugins Channel in the [FiftyOne Community Slack](https://slack.voxel51.com/) What we’ve built so far: - Week 0: 🌩️ [Image Quality Issues](https://github.com/jacobmarks/image-quality-issues) & 📈 [Concept Interpolation](https://github.com/jacobmarks/concept-interpolation) - Week 1: 🎨 [AI Art Gallery](https://github.com/jacobmarks/ai-art-gallery) & [Twilio Automation](https://github.com/jacobmarks/twilio-automation-plugin) - Week 2: ❓ [Visual Question Answering](https://github.com/jacobmarks/vqa-plugin) - Week 3: 🎥 [YouTube Player Panel](https://github.com/jacobmarks/fiftyone-youtube-panel-plugin) - Week 4: 🪞 [Image Deduplication](https://github.com/jacobmarks/image-deduplication-plugin) - Week 5: [👓Optical Character Recognition (OCR)](https://github.com/jacobmarks/pytesseract-ocr-plugin) & 🔑 [Keyword Search](https://github.com/jacobmarks/keyword-search-plugin) Ok, let’s dive into this week’s FiftyOne Plugin — [Zero-Shot Prediction](https://github.com/jacobmarks/zero-shot-prediction-plugin)! ## Zero-Shot Prediction 0️⃣🎯🔮 Most computer vision models are trained to predict on a preset list of label classes. In object detection, for instance, many of the most popular models like [YOLOv8](https://docs.ultralytics.com/models/yolov8/?h=#usage) and [YOLO-NAS](https://github.com/Deci-AI/super-gradients/blob/master/YOLONAS.md) are [pretrained](https://blogs.nvidia.com/blog/2022/12/08/what-is-a-pretrained-ai-model/) with the classes from the [MS COCO dataset](https://cocodataset.org/#home). If you download the weights checkpoints for these models and run prediction on your dataset, you will generate object detection bounding boxes for the 80 COCO classes. When you’re building your own machine learning application, typically you will be interested in a _different_ set of classes. Sometimes the classes are similar to the pretrained classes, whereas other times they can be completely different. To use these model architectures on your data, you either need to fine tune the pretrained model, or train a model from scratch. Sometimes, however, it may be beneficial to generate initial predictions with yourset of classes _before_ undertaking fine tuning or full-scale model training. For one, this can be a great way to accelerate the generation of ground truth labels. If the model is good enough, manually correcting its mistakes can be much quicker than labeling everything by hand. Additionally, having a set of predictions could prove useful when benchmarking the performance of any models you train. In computer vision, this is known as [zero-shot learning](https://en.wikipedia.org/wiki/Zero-shot_learning), or zero-shot prediction, because the goal is to generate predictions without explicitly being given any example predictions to learn from. With the advent of high quality multimodal models like [CLIP](https://github.com/openai/CLIP) and foundation models like [Segment Anything](https://github.com/facebookresearch/segment-anything), it is now possible to generate remarkably good zero-shot predictions for a variety of computer vision tasks, including: - [Image classification](https://paperswithcode.com/task/image-classification#:~:text=Image%20Classification%20is%20a%20fundamental,pertains%20to%20single%2Dobject%20images.) - [Object detection](https://paperswithcode.com/task/object-detection) - [Instance segmentation](https://paperswithcode.com/task/instance-segmentation) - [Semantic segmentation](https://paperswithcode.com/task/semantic-segmentation) This FiftyOne plugin streamlines the process of zero-shot prediction, unifying the interface across all four of these tasks, so that you can go from labels to predictions within the FiftyOne App, without writing a single line of code! ## Plugin Overview & Functionality For the sixth week of _10 Weeks of Plugins_, I built a Zero-Shot Prediction Plugin. This plugin allows you to specify a set of label classes and a task, and generate preliminary labels for your entire dataset. ![](https://cdn.sanity.io/images/h6toihm1/production/cfaaaaa10dbabaadcb0bfc9bf05a71603a64c5ac-1999x1061.png?auto=format&dpr=2&fit=max&q=75&w=1600) You can specify label classes either as a comma separated list, or by selecting a text file which contains a new class on each line: The plugin has five (!) [operators](https://docs.voxel51.com/plugins/index.html#operators): - `zero_shot_predict`: an umbrella operator for all zero-shot tasks - `zero_shot_classify`: an interface for zero-shot classification - `zero_shot_detect`: an interface for zero-shot object detection - `zero_shot_instance_segment`: an interface for zero-shot instance segmentation - `zero_shot_semantic_segment`: an interface for zero-shot semantic segmentation You can specify the computer vision task either from the modal for the main `zero_shot_predict` operator: Or by selecting the appropriate task’s operator from the operator list: After selecting a task, you will be prompted to select a model. For some tasks, such as classification, only one model (CLIP) is implemented out of the box. For others, there are multiple choices. Instance segmentation, for example, comes with three choices — one for each Segment Anything model size (B, L, H). At the bottom of the operator’s modal can you specify the name of the field in which to store the resulting predictions. By default, the field name is a formatted version of the model name. ## Delegating Execution All five of the operators in this plugin can optionally have their execution _delegated_ to be completed at a later time. If you choose to run the operators as delegated operators, you can schedule them in the App and then launch them from the command line with: ```bash 1fiftyone delegated launch ``` ## Setting Inference Target You can also choose to run any of these zero-shot prediction models on just a subset of your data. If you have a `DatasetView` loaded in the FiftyOne App that is distinct from the entire dataset — for example you are just looking at samples that match some filter — then you will have the option to run inference on either the entire dataset, or that particular subset. The same is true if you have samples “selected”: ## Zero-Shot Prediction in Action There are so many use cases for zero-shot prediction. Here’s just one example. Suppose you have some images with vehicles in them, and you want to detect their license plates. This isn’t one of the COCO label classes, but that is no longer a problem: ## Installing the Plugin If you haven’t already done so, install FiftyOne: ```python 1pip install fiftyone ``` Then you can download this plugin from the command line with: ```bash 1fiftyone plugins download https://github.com/jacobmarks/zero-shot-prediction-plugin ``` Refresh the FiftyOne App, and you should see the five operators in your operators list when you press the “ \` ” key. Because CLIP and SAM come with the [FiftyOne Model Zoo](https://docs.voxel51.com/user_guide/model_zoo/index.html), no additional steps are needed to use these models. However, models like [CLIPSeg](https://huggingface.co/blog/clipseg-zero-shot) (used for semantic segmentation) and [Owl-ViT](https://huggingface.co/docs/transformers/model_doc/owlvit) (used for object detection and instance segmentation) require that you have the Hugging Face [transformers](https://huggingface.co/docs/transformers/index) library installed. If you find other zero-shot models, you can add them as you see fit! ## Lessons Learned The Image Deduplication plugin is a Python Plugin with the usual structure (an `__init__.py`, `fiftyone.yml`, and `REAMDE.md` files). Additionally, it has an `assets` folder for storing icons, and a separate Python file for each task: e.g. object detection models implemented in `detection.py`. I split the code up in this way to separate each model’s implementation and postprocessing details from the unified interface for prediction. ## Keep up with FiftyOne’s Plugin Features The FiftyOne Plugin system is already incredibly powerful, and it’s getting more powerful with each and every release. This plugin utilizes a few of the features that have been recently added: the file explorer, and [delegated operators](https://docs.voxel51.com/plugins/index.html). The [file explorer](https://docs.voxel51.com/api/fiftyone.operators.types.html#fiftyone.operators.types.FileView), which you can use in your own Python plugins via the \`FileExplorerView\` view type, is a flexible widget that allows you to select a specific file or an entire folder from your file system. In [FiftyOne Teams](https://voxel51.com/fiftyone-teams/), it even allows you to navigate the files in your cloud buckets, all from within the FiftyOne App! In this plugin, I used the file explorer to let users select their labels, either from a local text file, or from URL. Delegated operators allow you to “delegate” that certain tasks be completed at a later time. From within the FiftyOne App, you schedule the job, and then you can launch the job from the command line with: ```bash 1fiftyone delegated launch ``` Delegating execution can be incredibly useful for long-running operations, like running inference with deep models on your dataset. You can set your operators to run in delegated mode with the `resolve_delegation()` method. This, for instance, would result in the operator _always_ running in delegated mode. ```python 1def resolve_delegation(self, ctx): 2 True ``` For this plugin, I took a slightly different approach, instead letting the user decide whether they want to run in delegate mode via an input parameter: ```python 1def _execution_mode(ctx, inputs): 2 delegate = ctx.params.get("delegate", False) 3 4 if delegate: 5 description = "Uncheck this box to execute the operation immediately" 6 else: 7 description = "Check this box to delegate execution of this task" 8 9 inputs.bool( 10 "delegate", 11 default=False, 12 required=True, 13 label="Delegate execution?", 14 description=description, 15 view=types.CheckboxView(), 16 ) 17 18 if delegate: 19 inputs.view( 20 "notice", 21 types.Notice( 22 label=( 23 "You've chosen delegated execution. Note that you must " 24 "have a delegated operation service running in order for " 25 "this task to be processed. See " 26 "https://docs.voxel51.com/plugins/index.html#operators " 27 "for more information" 28 ) 29 ), 30 ) 31 32 33def resolve_delegation(self, ctx): 34 return ctx.params.get("delegate", False) ``` ## Reducing Boilerplate Code Because the four task-specific operators naturally gave way to very similar information flows within the code, I found myself writing nearly identical code for the `resolve_input()` and `execute()` methods for each of these operators. This wasn’t very satisfying. A more satisfying and cleaner approach is to pass the context object `ctx` into other functions which are defined outside of the operator object. To make this happen, I created `_input_control_flow(ctx, task)` and `_execute_control_flow(ctx, task)` functions which take in the `ctx` and a string specifying the task. This massively simplified the operator definitions. Take the zero shot instance segmentation operator for example: ```python 1class ZeroShotInstanceSegment(foo.Operator): 2 @property 3 def config(self): 4 _config = foo.OperatorConfig( 5 name="zero_shot_instance_segment", 6 label="Perform Zero Shot Instance Segmentation", 7 dynamic=True, 8 ) 9 _config.icon = "/assets/icon.svg" 10 return _config 11 12 def resolve_delegation(self, ctx): 13 return ctx.params.get("delegate", False) 14 15 def resolve_input(self, ctx): 16 inputs = _input_control_flow(ctx, "instance_segmentation") 17 return types.Property(inputs) 18 19 def execute(self, ctx): 20 _execute_control_flow(ctx, "instance_segmentation") ``` ## Conclusion Zero-shot prediction is becoming increasingly important for both pre-labeling and benchmarking workflows. This plugin saves you the trouble of dealing with myriad standards, unifying and simplifying the process of zero-shot prediction for four essential computer vision tasks. It will save you a lot of time, and make your life that much easier. The best part is that it can be used in conjunction with your existing annotation workflows, or with next week’s Active Learning plugin! Stay tuned over the remaining weeks in the _Ten Weeks of FiftyOne Plugins_ while we continue to pump out a killer lineup of plugins! You can track our journey in our [ten-weeks-of-plugins repo](https://github.com/jacobmarks/ten-weeks-of-plugins) — and I encourage you to fork the repo and join me on this journey! [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [custom plugins](https://voxel51.com/blog/tag/custom-plugins) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [image dataset](https://voxel51.com/blog/tag/image-dataset) [model predictions](https://voxel51.com/blog/tag/model-predictions) [plugins](https://voxel51.com/blog/tag/plugins) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/fedd0c008df994c1834839ea35bc65b34e47fb5d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Double Trouble: Eliminate Image Duplicates with FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 14, 2023](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/bad75ba72dfae8cdefc3d9afe33a1ea9a9c4ec36-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Optical Character Recognition with PyTesseract\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 21, 2023](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract) [![](https://cdn.sanity.io/images/h6toihm1/production/c1075849ea942531ffbb4afc830cc12598f0d919-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Supercharge Your Annotation Workflow with Active Learning\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Oct 5, 2023](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-339-lllmstxt|> ## 3D Detections Tips [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Tips & Tricks](https://voxel51.com/blog/category/tips-tricks) 3D Detections – FiftyOne Tips and Tricks – September 29th, 2023 Sep 29, 2023 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#a91c18675e5e) [What Makes a 3D Detection?](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#b1be8d8f8f71) [Creating a Basic 3D Bounding Box](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#16e8a2a82168) [Rotating the Bounding Box](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#fe24c733eebb) [3D Polylines](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#c8470a0477d0) [Orthographic Projections](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#2ead8f0e1fb0) [See it in action!](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#d9ecfe58cf08) [Conclusion](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#7426bd97d165) [Join the FiftyOne Community!](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#784742363bc4) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#a91c18675e5e) [What Makes a 3D Detection?](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#b1be8d8f8f71) [Creating a Basic 3D Bounding Box](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#16e8a2a82168) [Rotating the Bounding Box](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#fe24c733eebb) [3D Polylines](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#c8470a0477d0) [Orthographic Projections](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#2ead8f0e1fb0) [See it in action!](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#d9ecfe58cf08) [Conclusion](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#7426bd97d165) [Join the FiftyOne Community!](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023#784742363bc4) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) Welcome to our weekly FiftyOne tips and tricks blog where we cover interesting workflows and features of FiftyOne! This week we are taking a look at [3D Detections](https://docs.voxel51.com/user_guide/using_datasets.html#d-detections). We aim to cover the basics of creating 3D detections and how they can be utilized in LIDAR or point cloud datasets. ## Wait, what’s FiftyOne? [FiftyOne](https://voxel51.com/fiftyone/) is an open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - [Get started!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ), we’re always happy to help. ## What Makes a 3D Detection? 3D detection can mean a variety of things in computer vision depending on the context. This week we are looking at 3D detections as they pertain to 3D LIDAR or point cloud spaces. These kinds of detections are common in automotive and topography datasets and leverage the insightful knowledge of the different depths of the environment. In FiftyOne, detections are a single label for either 2D or 3D, it is how we define them that makes the difference! Let's start with 2D and see how we build a detection label. ```python 1sample["ground_truth"] = fo.Detections( 2 detections=[\ 3 fo.Detection(\ 4 label="cat",\ 5 bounding_box=[0.480, 0.513, 0.397, 0.288],\ 6 ),\ 7 ] 8) ``` We can see that all we need is the label and the bounding box to construct our label. In 3D detections, we also provide a label, but instead of a 2D bounding box, we now provide the details in order to build a 3D bounding box, accounting for the extra dimension. ```python 1# Object label 2label = "vehicle" 3 4# Object center `[x, y, z]` in scene coordinates 5location = [0.47, 1.49, 69.44] 6 7# Object dimensions `[x, y, z]` in scene units 8dimensions = [2.85, 2.63, 12.34] 9 10# Object rotation `[x, y, z]` around its center, in `[-pi, pi]` 11rotation = [0, -1.56, 0] 12 13# A 3D object detection 14detection = fo.Detection( 15 label=label, 16 location=location, 17 dimensions=dimensions, 18 rotation=rotation, 19) ``` If you are interested in placing 3D detections onto a 2D space, checkout one of our previous Tips and Tricks on [Polylines](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023/)! 3D bounding box detections are defined in FiftyOne with three input parameters: location, rotation, and dimensions. Location refers to the location of the bounding box on the point cloud coordinate system. The location is the absolute center of the bounding box. Rotation is the amount the bounding box is rotated among its axes. Rotation is also given as `[x, y, z]` where x is the rotation around the x axis in a space of `[-pi, pi]`. Finally, dimensions are the width, length, and height of the bounding box. Lets try making our first 3D bounding box! ## Creating a Basic 3D Bounding Box ```python 1import fiftyone as fo 2import fiftyone.zoo as foz 3from fiftyone import ViewField as F 4 5 6dataset = foz.load_zoo_dataset("quickstart-groups") 7session = fo.launch_app(dataset) ``` We start by loading in a dataset that has point clouds as a part of it. `quickstart-groups` is a subset of a KITTI formatted dataset. Feel free to explore the point clouds by clicking on one of the groups and looking at the existing detections. Next, we will add a very basic bounding box at the center of our point cloud. Switch to the `pcd` group slice to add the detection to the point cloud, then add the detection. We create a view at the end to show only the sample we are interested in. ```python 1dataset.group_slice = "pcd" 2sample = dataset.first() 3 4 5bounding_box = fo.Detection( 6 label="example", 7 location=[0,0,0], 8 rotation=[0, 0, 0], 9 dimensions=[3,3,3] 10 ) 11 12 13sample["example"]= bounding_box 14sample.save() 15dataset.save() 16view = dataset.filter_labels( 17 "example", F("label").is_in(["example"]) 18) 19 20 21session.view = view ``` Below we can see the results! With our first box placed, we can play around with some of the input parameters to get comfortable with the FiftyOne 3D visualizer. Lets modify the `location` as well as the `wlh` to get more of a grasp of the coordinate system. ```python 1sample = dataset.first() 2 3 4bounding_box = fo.Detection( 5 label="example", 6 location=[0,2,4], 7 rotation=[0, 0, 0], 8 dimensions=[1,2,3] 9 ) 10 11 12sample["example"]= bounding_box 13sample.save() 14dataset.save() 15 16 17view = dataset.filter_labels( 18 "example", F("label").is_in(["example"]) 19) 20 21 22session.view = view ``` We can see how the box has transformed and shifted along the coordinate plane. Feel free to tweak these numbers until you get a comfortable feel. A good foundation of the coordinate system will help troubleshoot any potential problems in the future! ## Rotating the Bounding Box Rotating the bounding box is simple, for each given axis, define in radians how much you want your box to rotate around the circle. Play around with the tweaking rotation numbers to get a handle on situating your boxes. 💡 Pro Tip, save rotation for last if you are trying to align your boxes. Different standards exist for 3D detections, so be confident you aren't accidentally rotating too much that now your width and length have flipped! ```python 1import numpy as np 2from fiftyone import ViewField as F 3# x = left+right = orange | y = forward/back = green | z = up+down = blue 4bounding_box = fo.Detection( 5 label="example", 6 location=[0,0,0], 7 rotation=[np.pi/4, np.pi/4, 0], 8 dimensions=[3,2,1] 9 ) 10 11 12sample["example"] = bounding_box 13sample.save() 14dataset.save() 15view = dataset.filter_labels( 16 "example", F("label").is_in(["example"]) 17) 18 19 20session.view = view ``` ## 3D Polylines Similar to the last example, we can define polygons with a set of points to create a 3D polyline. To create a 3D polyline, all you need to do is define the points on the grid in a list. 💡 To close a shape, specify the first vertice again. We can add these lines to our sample and display our new shape! ```python 1label = "lane" 2 3 4# A list of lists of `[x, y, z]` points in scene coordinates describing 5# the vertices of each shape in the polyline 6points3d = [[[1, 1, 1], [0, 2, 2]], [[0, 2, 2], [-1, 1, 1]], [[-1, 1, 1], [1, 1, 1]]] 7 8 9# A set of semantically related 3D polylines 10polyline = fo.Polyline(label=label, points3d=points3d,) 11 12 13sample["polylines"] = polyline 14sample.save() 15dataset.save() 16view = dataset.filter_labels( 17 "polylines", F("label").is_in(["lane"]) 18) 19 20 21session.view = view 22 ``` ## Orthographic Projections Ending with a final tool you can use when working with 3D data is [Orthographic Projections](https://docs.voxel51.com/api/fiftyone.utils.utils3d.html?highlight=ortho#fiftyone.utils.utils3d.compute_orthographic_projection_images)! You have probably noticed when running the app that the LIDAR samples have no thumbnails. We can computer orthographic projections of our LIDAR space to create thumbnails and add a new slice to our dataset to curate by! This will give you a birds eye view of your data to look through with tons of options to change coloring or rendering. ```python 1import fiftyone.utils.utils3d as fou3d 2 3min_bound = (0, -15, -2.73) 4max_bound = (20, 15, 1.27) 5size = (-1, 512) 6fou3d.compute_orthographic_projection_images( 7 dataset, 8 size, 9 "/tmp/proj", 10 shading_mode="height", 11 out_group_slice="proj" 12) 13 14session.view = dataset.view() ``` Now our data is that much easier to curate and parse through! ## **See it in action!** ## Conclusion In summary, 3D detections in point clouds are vital for advancing computer vision capabilities. With FiftyOne, they enable you to understand precise spatial data, allowing you to comprehend and interact with your dataset in a three-dimensional context. FiftyOne’s 3D detection powers applications such as autonomous driving, robotics, and augmented reality by providing critical depth information and object localization. Level up your three-dimensional data today by using FiftyOne! ## Join the FiftyOne Community! Join the thousands of engineers and data scientists already using FiftyOne to solve some of the most challenging problems in computer vision today! - 2,000+ [FiftyOne Slack](https://slack.voxel51.com/) members - 4,000+ stars on [GitHub](https://github.com/voxel51/fiftyone) - 5,000+ [Meetup members](https://www.meetup.com/pro/computer-vision-meetups/) - [Used by](https://github.com/voxel51/fiftyone/network/dependents?package_id=UGFja2FnZS0xNzAxODM0MjUx) 370+ repositories - 60+ [contributors](https://github.com/voxel51/fiftyone/graphs/contributors) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [LIDAR](https://voxel51.com/blog/tag/lidar) [open source](https://voxel51.com/blog/tag/open-source) [point clouds](https://voxel51.com/blog/tag/point-clouds) ![](https://cdn.sanity.io/images/h6toihm1/production/3b39056326e925c10b46da1324bc3c5840a1629c-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Dan Gural Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/42382839895f816e61408a33720ad4d8370a1e01-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Video Labels – FiftyOne Tips and Tricks – October 14th, 2023\\ \\ Computer Vision, Tips & Tricks\\ \\ • \\ \\ Oct 13, 2023](https://voxel51.com/blog/computer-vision-video-labels-fiftyone-tips-and-tricks-october-14th-2023) [![](https://cdn.sanity.io/images/h6toihm1/production/c9967da6d043cb267a4432030ad161d3443f5ba2-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Segment Anything in a CT Scan with NVIDIA VISTA-3D\\ \\ Computer Vision, Plugins\\ \\ • \\ \\ Jul 2, 2024](https://voxel51.com/blog/segment-anything-in-a-ct-scan-with-nvidia-vista-3d) [![](https://cdn.sanity.io/images/h6toihm1/production/aa202a26af141cab6f1c7785026a9edf71ef8101-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ How to Make the Best Self-Driving Dataset\\ \\ Computer Vision\\ \\ • \\ \\ Jan 15, 2025](https://voxel51.com/blog/how-to-make-the-best-self-driving-dataset) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-340-lllmstxt|> ## Visit Voxel51 at ICCV23 [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) 3 Reasons to Visit Voxel51 at ICCV23! Oct 2, 2023 • 4 min read Article content In this article [1 - 💡 See visual data in a whole new light, including six new ICCV datasets!](https://voxel51.com/blog/3-reasons-to-visit-voxel51-at-iccv23#7bd73c993724) [2 – 👋 Meet fellow enthusiasts in the computer vision community](https://voxel51.com/blog/3-reasons-to-visit-voxel51-at-iccv23#e09e7de2e1aa) [3 - 👕 Score epic FiftyOne swag](https://voxel51.com/blog/3-reasons-to-visit-voxel51-at-iccv23#7b2598fd1838) [Can’t make ICCV?](https://voxel51.com/blog/3-reasons-to-visit-voxel51-at-iccv23#9bf8537ab024) In this article [1 - 💡 See visual data in a whole new light, including six new ICCV datasets!](https://voxel51.com/blog/3-reasons-to-visit-voxel51-at-iccv23#7bd73c993724) [2 – 👋 Meet fellow enthusiasts in the computer vision community](https://voxel51.com/blog/3-reasons-to-visit-voxel51-at-iccv23#e09e7de2e1aa) [3 - 👕 Score epic FiftyOne swag](https://voxel51.com/blog/3-reasons-to-visit-voxel51-at-iccv23#7b2598fd1838) [Can’t make ICCV?](https://voxel51.com/blog/3-reasons-to-visit-voxel51-at-iccv23#9bf8537ab024) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) As Jacob Marks said in his recent post, [ICCV 2023 Survival Guide: 10 Computer Vision Papers You Won’t Want to Miss](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide/), ICCV is shaping up to be something special! The premier international computer vision [event](https://iccv2023.thecvf.com/) is happening this week in Paris, France, with the main conference and expo taking place on October 4-6. Events like this are an amazing way to discover new ideas and build connections with others in the computer vision community. Here are three reasons to put a visit to Voxel51 Booth #55 on your ICCV23 agenda! ## 1 - 💡 See visual data in a whole new light, including six new ICCV datasets! Nothing hinders the success of machine learning systems more than poor quality data. But how do you see issues hiding in your datasets that are limiting your ML model performance? Shine a light into your visual data and your model’s failure modes using open source FiftyOne! FiftyOne provides unprecedented visibility for optimizing your dataset analysis pipeline. Use it to illuminate visual data in ways that matter most to you: visualize complex labels, explore samples and scenarios of interest, find annotation mistakes, evaluate your models, identify model failure modes, and much more! While it’s possible to load any visual dataset into FiftyOne, we loaded up these ICCV datasets so you can instantly see for yourself. Visit Voxel51 at ICCV Booth #55 to see these datasets in action, or explore them now in your browser! ### **Building3D** Building3D is an urban-scale dataset and benchmarks for learning roof structures from point clouds from [this paper](https://arxiv.org/abs/2307.11914) by Ruisheng Wang, Shangfeng Huang, and Hongxin Yang. Building3D consists of more than 160 thousand buildings along with corresponding point clouds, mesh and wire-frame models, covering 16 cities in Estonia – about 998 Km2. [**Explore Building3D instantly in your browser! >>**](https://try.fiftyone.ai/datasets/building3d/samples) ### **EgoObjects** EgoObjects is a large-scale egocentric dataset for fine-grained object understanding published in [this paper](https://arxiv.org/abs/2309.08816) by Chenchen Zhu, Fanyi Xiao, Andres Alvarado, Yasmine Babaei, Jiabo Hu, Hichem El-Mohri, Sean Chang Culatana, Roshan Sumbaly, and Zhicheng Yan. EgoObjects 1.0 includes 114K annotated frames (79K train, 5.7K val, 29.5K test) sampled from 9K+ videos collected by 250 participants from 50+ countries using 4 wearable devices, and over 650K object annotations from 368 object categories. [**Explore EgoObjects instantly in your browser! >>**](https://try.fiftyone.ai/datasets/egoobjects-val/samples) ### **EqBen** EqBen (Equivariant Benchmark) is a new challenging benchmark created to diagnose the equivariance of VLMs with visual-minimal change samples. It was published in the paper [Equivariant Similarity for Vision-Language Foundation Models](https://arxiv.org/abs/2303.14465) by Tan Wang, Kevin Lin, Linjie Li, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, and Lijuan Wang. [**Explore EqBen instantly in your browser! >>**](https://try.fiftyone.ai/datasets/eqben-test/samples) ### **MOSE** Mose is a new dataset for video object segmentation in complex scenes published in [this paper](https://arxiv.org/abs/2302.01872) by Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, Philip H.S. Torr, and Song Bai. MOSE contains 2,149 video clips and 5,200 objects from 36 categories, with 431,725 high-quality object segmentation masks. The most notable feature of the MOSE dataset is complex scenes with crowded and occluded objects. [**Explore MOSE instantly in your browser! >>**](https://try.fiftyone.ai/datasets/mose/samples) ### **SatlasPretrain** SatlasPretrain is a large-scale dataset for remote sensing image understanding, published in [this paper](https://arxiv.org/abs/2211.15660) by Favyen Bastani, Piper Wolters, Ritwik Gupta, Joe Ferdinando, and Aniruddha Kembhavi. SatlasPretrain combines more than 30 TB of satellite images from public sources such as Sentinel-2 and NAIP with 137 label categories, making it an effective pre-training dataset that greatly reduces the effort needed to develop robust models for downstream satellite image applications. [**Explore SatlasPretrain instantly in your browser! >>**](https://try.fiftyone.ai/datasets/satlas-marine-infrastructure/samples) ### **SportsMOT** SportsMOT is a large multi-object tracking dataset in multiple sports scenes, published in this paper by Yutao Cui, Chenkai Zeng, Xiaoyu Zhao, Yichun Yang, Gangshan Wu, and Limin Wang. SportsMOT consists of 240 video sequences, 150K+ frames, and 1.6M+ bounding boxes from basketball, volleyball, and football. [**Explore SportsMOT instantly in your browser! >>**](https://try.fiftyone.ai/datasets/sportsmot-validation/samples) ## 2 – 👋 Meet fellow enthusiasts in the computer vision community We love meeting fellow members of the computer vision community! Come meet ML engineers, developers, co-founders, and open source enthusiasts from the Voxel51 team at booth #55. Let’s discuss all things computer vision, explore FiftyOne hands-on, and send you home with our latest epic swag (while supplies last!). ![](https://cdn.sanity.io/images/h6toihm1/production/c8d33b91d4eb8cd5e8402c75e49e5af1c3f8ee48-1200x764.jpg?auto=format&dpr=2&fit=max&q=75&w=1200) ## 3 - 👕 Score epic FiftyOne swag Voxel51 is known for bringing epic swag to community gatherings. And, who doesn’t like a good giveaway to commemorate an amazing event? Stop by our booth to see our lineup of goodies, including a Paris- and ICCV23-inspired limited-time tshirt! And snag some swag for yourself! ![](https://cdn.sanity.io/images/h6toihm1/production/11b46fe71ad2f1ac6482db45cc4dfb950e3485cc-3420x1628.png?auto=format&dpr=2&fit=max&q=75&w=1600) As much as we love swag, we also simply love welcoming new members to the open source FiftyOne community. Why? We know how valuable FiftyOne truly is for data engineers and scientists and therefore getting it into the hands of even more people is what we’re all about. Even if you come for the swag, we’re convinced you’ll walk away with two things you’ll love: sweet gear and first-hand experience of how FiftyOne can help you level up your computer vision workflows. ## Can’t make ICCV? If you can’t attend ICCV this year, no problem! Here are other ways to connect with the FiftyOne community: - Check out our lineup of [upcoming events](https://voxel51.com/computer-vision-events/), including Workshops, Meetups, and more. - [Get started with FiftyOne!](https://voxel51.com/docs/fiftyone/index.html) We’ve made it easy to get up and running in a few minutes. - If you like what you see on GitHub, [give the project a star](https://github.com/voxel51/fiftyone). - Join the FiftyOne [Slack community](https://join.slack.com/t/fiftyone-users/shared_invite/zt-s6936w7b-2R5eVPJoUw008wP7miJmPQ) of 2100+ computer vision and open source enthusiasts, we’re always happy to help. [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [computer vision events](https://voxel51.com/blog/tag/computer-vision-events) [events](https://voxel51.com/blog/tag/events) [iccv](https://voxel51.com/blog/tag/iccv) [ICCV23](https://voxel51.com/blog/tag/iccv23) Monica Tran Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/5d9fe483cd6c19e4ee6246bac487e4716fe99fe7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ 5 Reasons to Visit Voxel51 at CVPR\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Jun 9, 2023](https://voxel51.com/blog/5-reasons-to-visit-voxel51-at-cvpr) [![](https://cdn.sanity.io/images/h6toihm1/production/f7a4b48dc859183b47d95b203337e268f540265d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ ICCV 2023 Survival Guide\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Sep 27, 2023](https://voxel51.com/blog/computer-vision-iccv-2023-survival-guide) [![](https://cdn.sanity.io/images/h6toihm1/production/622b7369c791083b44e3034b2b8772d3ecada8bb-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ FiftyOne Computer Vision Community Update – November 2023\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Nov 1, 2023](https://voxel51.com/blog/fiftyone-computer-vision-community-update-november-2023) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-341-lllmstxt|> ## Custom GitHub Badges [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Tutorials](https://voxel51.com/blog/category/tutorials) Badger: Custom GitHub Badges Made Easy Oct 4, 2023 • 5 min read Article content In this article [Create and Manage Your Badges from the Command Line](https://voxel51.com/blog/computer-vision-badger-custom-github-badges#222b5184dd5f) [Badge Management with Badger](https://voxel51.com/blog/computer-vision-badger-custom-github-badges#20f98027684b) [Installation](https://voxel51.com/blog/computer-vision-badger-custom-github-badges#240a29f39206) [Conclusion](https://voxel51.com/blog/computer-vision-badger-custom-github-badges#570322dda93d) In this article [Create and Manage Your Badges from the Command Line](https://voxel51.com/blog/computer-vision-badger-custom-github-badges#222b5184dd5f) [Badge Management with Badger](https://voxel51.com/blog/computer-vision-badger-custom-github-badges#20f98027684b) [Installation](https://voxel51.com/blog/computer-vision-badger-custom-github-badges#240a29f39206) [Conclusion](https://voxel51.com/blog/computer-vision-badger-custom-github-badges#570322dda93d) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Create and Manage Your Badges from the Command Line Back in June, inspired by [paperswithcode.com](https://paperswithcode.com/) my team and I created [PapersWithData](https://github.com/voxel51/papers-with-data) to emphasize the role that data plays in pushing the boundaries in machine learning and computer vision. To get this project off the ground, we worked together with the authors of a dozen datasets to convert their datasets into a common format and make them publicly browsable. Here’s a screenshot of one of the tables from the original GitHub README: Data wrangling can be a notoriously frustrating process. Yet the most frustrating part of the entire project, _by far_, was creating the custom FiftyOne GitHub badges — the tiny buttons with the Voxel51 icon and text “FiftyOne”, each of which pointed to a different URL: Although we bootstrapped the project with our classic orange and gray Voxel51 logo mark, we knew we wanted to quickly move to an ADA compliant color scheme. Just making this minor change – setting the color of the logo — was also frustrating! To create these badges, we used [Shields.io](https://shields.io/), a flexible framework for creating a variety of standard and custom badges. Shields.io supports a ton of named logos like GitHub, Discord, and Node.js (see a complete list [here](https://simpleicons.org/)), as well as [custom logos](https://shields.io/docs/logos). However, to use a custom logo like the icon for your company or project, you need to convert the SVG into [base64](https://en.wikipedia.org/wiki/Base64). Every time you want to reuse the same badge, you need to dig through your old projects to find the Markdown code. And every time you want to make a slight modification to an existing badge, you need to remember the Shields.io syntax to do so. To solve these problems, and more, I built [Badger](https://github.com/voxel51/badger)! Badger is an open source Python library that allows you to create, edit, store, and manage custom badges, all from the command line. In this post, I share the basics of how to get up and running with Badger to create the GitHub badge of your dreams. With that, here is just a small subset of what you can do with Badger – and if it sounds interesting to you, I invite you to give it a try! ## Badge Management with Badger ### 🔨 Creating Badges Badger makes badge creation a breeze, walking you through the entire process. You can run `Badger create` to begin the badge creation process — in which case the first detail you will be prompted for is a “name” for the badge — or you can specify the name of the new badge in the command. Badges can be created from either local SVGs, or remotely hosted SVG files. The latter can be useful if you want to work with an SVG stored in a GitHub gist, or on some public server. Badger uses Python’s `requests` library to get the raw data from the SVG at Markdown generation time, so you never need to store it locally. Let’s see an example of each. First, let’s create a FiftyOne badge (way more easily than we did initially with PapersWithData!): In this example, we passed a local (relative in this case) path to the logo file, and in the initial command we told Badger to name the badge “fiftyone”. Also notice that we left some attributes blank. Not every attribute is required to create a badge! This badge is perfectly usable as is. Now, let’s create a badge from an SVG at a remote URL (hosted by [SVG Repo](https://www.svgrepo.com/)): This creates the following badge: ### 🖨️ Printing Badges Once you have a badge saved in your Badger config, you can print the formatted Markdown to stdout, or directly into a Markdown file of your choice: ### 📋 Copying Badges to the Clipboard When you’re working in a Jupyter notebook, you can save the trouble of manually selecting and copying the outputs of a print statement in order to enter into a Markdown cell. Badger’s `copy` command uses the [pyperclip](https://pypi.org/project/pyperclip/) Python library to copy the Markdown for the badge to your clipboard — you can just paste the result directly into a Markdown cell: 💡You can override any of a badge’s default attributes at print/copy time — even the logo file and the URL the badge points to! ### 📂 Listing Your Badges Where would we be without a way to keep track of all of the awesome badges we’ve created? Luckily, the `Badger list` command gives us a formatted list of all of the badges in our config: If you want detailed information about a single badge, you can use the `badger info ` command: ### 🪞 Cloning Badges Badger also makes it incredibly easy to save different variations of the same basic badge via the `badger clone` command. Pass the name of the original badge, the name of the new badge, and optionally any differences you want to be reflected in the new badge. For instance, to create the ADA compliant version of the `voxel51` badge, we can run: ```bash 1badger clone voxel51 voxel51_ada --logoColor white ``` This will save a new `voxel51_ada` badge to our badger config, with a white logo. It’s that simple! ### 🖋️ Editing and Deleting Badges If you want to edit a badge’s configuration — suppose you want to make `plastic` the new default style for your `villa` badge — you can do so with Badger’s `badger edit` command: ```bash 1badger edit villa --style plastic ``` You can also delete a badge from your Badger config file with `badger delete ` To delete this \`villa\` badge, you would run: ```bash 1badger delete villa ``` ### ✨ Going Wild with AI-Generated Badges Just for fun, I’ve added a `badger go-wild` command, which will use GPT-4 (with some light prompt engineering and validation) to create a new SVG for you to use in your badges. If you use the command without a prompt, it will generate a completely random SVG, but you can also specify what the subject of the SVG should be with the `--prompt` argument, followed by the desired subject. For instance, to generate a turtle SVG, run the following command: ```bash 1badger go-wild --prompt turtle ``` The SVGs are basic, when they even work, but it’s fun! Creating a high-quality SVG generating model is a project for another day 🐢. ## Installation To install Badger, run the following commands in your terminal: ```bash 1git clone https://github.com/voxel51/badger.git 2cd badger 3pip install -e . ``` If you want to put your Badger config file in a custom location, you can set this with the `BADGER_CONFIG_FILE` environment variable: ```bash 1export BADGER_CONFIG_FILE= ``` If you want to use the `badger go-wild` command to generate some (rudimentary) custom badges with GPT-4, you must have an [OpenAI account](https://openai.com/) set up, and the `OPENAI_API_KEY` environment variable set. ```bash 1export OPENAI_API_KEY= ``` That’s it! ## Conclusion Badger is the library you didn’t know you needed. It is lightweight, incredibly easy to use, and will save you a lot of trouble. Go forth and make custom badges for all of your GitHub projects, large and small! [https://github.com/voxel51/badger](https://github.com/voxel51/badger) 🚀Not to BADGER you, but if you like Badger, check out some of our other open source projects, which will make your life easier: - 5️⃣1️⃣ [FiftyOne](https://github.com/voxel51/fiftyone): The Leading Library for Data Curation and Visualization - 🤖⌨️ [VoxelGPT](https://github.com/voxel51/voxelgpt): Your AI Assistant for Computer Vision - 📚🔍 [DocSearch](https://github.com/voxel51/fiftyone-docs-search): Turn Your Docs into a Searchable Repo - 🧩🔌 [FiftyOne Plugins](http://fiftyone-plugins/plugins%20at%20main%20%C2%B7%20voxel51/fiftyone-plugins): Build Custom Machine Learning Applications in Minutes [Badger](https://voxel51.com/blog/tag/badger) [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [custom badges](https://voxel51.com/blog/tag/custom-badges) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [Python library](https://voxel51.com/blog/tag/python-library) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/047b21a97f6c858334f9f35ed89fa7655ebf5767-4000x2250.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ State-of-the-Art Object Detection with YOLO-NAS & FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ May 4, 2023](https://voxel51.com/blog/state-of-the-art-object-detection-with-yolo-nas-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/713e4352b25d3b4ee12eab92246ceff22f808471-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Spending My First Week With FiftyOne\\ \\ Computer Vision, Tutorials\\ \\ • \\ \\ Aug 21, 2023](https://voxel51.com/blog/spending-my-first-week-with-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/fedd0c008df994c1834839ea35bc65b34e47fb5d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Double Trouble: Eliminate Image Duplicates with FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 14, 2023](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-342-lllmstxt|> ## Enhance Annotation Workflow [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Plugins](https://voxel51.com/blog/category/plugins), [Tutorials](https://voxel51.com/blog/category/tutorials) Supercharge Your Annotation Workflow with Active Learning Oct 5, 2023 • 8 min read Article content In this article [FiftyOne Active Learning Plugin Can Help You Label Data Faster!](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning#d35a6f20c5d5) [Active Learning 🏃🎓🔄](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning#1357ec4d7086) [Plugin Overview & Functionality](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning#491c39d8cba2) [Lessons Learned](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning#7fb47e327df0) [Conclusion](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning#76c882c21362) In this article [FiftyOne Active Learning Plugin Can Help You Label Data Faster!](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning#d35a6f20c5d5) [Active Learning 🏃🎓🔄](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning#1357ec4d7086) [Plugin Overview & Functionality](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning#491c39d8cba2) [Lessons Learned](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning#7fb47e327df0) [Conclusion](https://voxel51.com/blog/supercharge-your-annotation-workflow-with-active-learning#76c882c21362) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## FiftyOne Active Learning Plugin Can Help You Label Data Faster! Welcome to week seven of _Ten Weeks of Plugins_. During these ten weeks, we will be building a FiftyOne Plugin (or multiple!) each week and sharing the lessons learned! If you’re new to them, FiftyOne Plugins provide a flexible mechanism for anyone to extend the functionality of their FiftyOne App. You may find the following resources helpful: - [FiftyOne Plugins Repo](https://github.com/voxel51/fiftyone-plugins) - [FiftyOne Plugin Docs](https://docs.voxel51.com/plugins/index.html#downloading-plugins) - Plugins Channel in the [FiftyOne Community Slack](https://slack.voxel51.com/) What we’ve built so far: - Week 0: 🌩️ [Image Quality Issues](https://github.com/jacobmarks/image-quality-issues) & 📈 [Concept Interpolation](https://github.com/jacobmarks/concept-interpolation) - Week 1: 🎨 [AI Art Gallery](https://github.com/jacobmarks/ai-art-gallery) & [Twilio Automation](https://github.com/jacobmarks/twilio-automation-plugin) - Week 2: ❓ [Visual Question Answering](https://github.com/jacobmarks/vqa-plugin) - Week 3: 🎥 [YouTube Player Panel](https://github.com/jacobmarks/fiftyone-youtube-panel-plugin) - Week 4: 🪞 [Image Deduplication](https://github.com/jacobmarks/image-deduplication-plugin) - Week 5: [👓Optical Character Recognition (OCR)](https://github.com/jacobmarks/pytesseract-ocr-plugin) & 🔑 [Keyword Search](https://github.com/jacobmarks/keyword-search-plugin) - Week 6: 🎭 [Zero-shot Prediction](https://github.com/jacobmarks/zero-shot-prediction-plugin) Ok, let’s dive into this week’s FiftyOne Plugin — [Active Learning](https://github.com/jacobmarks/active-learning-plugin)! ## Active Learning 🏃🎓🔄 When it comes to machine learning, one of the most time-consuming and costly steps is data annotation. In the realm of computer vision, labeling images or videos can be an incredibly laborious task, often requiring a team of annotators and hours of meticulous work to generate high-quality labels. Even with a well-labeled dataset, the effort doesn't stop there: model training, evaluation, and then re-annotation in case of inaccuracies are all parts of an ongoing cycle. What if you could make this iterative process smarter and more efficient? Enter [Active Learning](https://en.wikipedia.org/wiki/Active_learning_(machine_learning))—a paradigm that iteratively selects the most "informative" or "ambiguous" examples for labeling, thereby reducing the amount of manual annotation needed. In practical terms, this means your model gets better, faster, and with fewer labeled samples. This FiftyOne plugin brings Active Learning to your computer vision data, allowing you to integrate this accelerant directly into your annotation workflow. Now you can prioritize, query, and annotate the most crucial data points, all within the FiftyOne App—no coding necessary. The best part? You can use this in tandem with your traditional annotation service providers (via FiftyOne’s integrations with [CVAT](https://docs.voxel51.com/integrations/cvat.html), [Labelbox](https://docs.voxel51.com/integrations/labelbox.html) and [Label Studio](https://docs.voxel51.com/integrations/labelstudio.html)), or even with last week’s [Zero-shot Prediction plugin](https://github.com/jacobmarks/zero-shot-prediction-plugin)! Read on to learn how you can leverage the Active Learning Plugin to build high-quality models with less manual effort. ## Plugin Overview & Functionality For the seventh week of [_10 Weeks of Plugins_](https://voxel51.com/blog/category/computer-vision-plugins/), I built an Active Learning Plugin. This plugin leverages the tried and tested modular active learning framework, [modAL](https://modal-python.readthedocs.io/en/latest/), to help you expedite your data labeling processes. modAL is an active learning framework built on top of [Sci-kit learn](https://scikit-learn.org/stable/index.html). It allows you to apply a variety of active learning strategies to models from Sci-kit learn, including solitary “estimators” like the `KNeighborsClassier`, and ensembles like the `RandomForestClassifier`, and even to create [“committees” out of estimators](https://modal-python.readthedocs.io/en/latest/content/models/Committee.html), so you can estimate label uncertainty by comparing multiple hypotheses about the underlying data. All of the necessary functionality from modAL is wrapped into FiftyOne Python [operators](https://docs.voxel51.com/plugins/index.html#operators), so that you can reap the rewards of Active Learning without writing a line of code. 💡At present, this plugin only supports Active Labeling for Classification tasks. However, it could be extended to [object detection](https://arxiv.org/pdf/2004.04699.pdf), [semantic segmentation](https://openaccess.thecvf.com/content_CVPR_2020/html/Siddiqui_ViewAL_Active_Learning_With_Viewpoint_Entropy_for_Semantic_Segmentation_CVPR_2020_paper.html), and other computer vision tasks. The plugin has three operators: - `create_learner`: creates an active learning model and environment - `query_learner`: queries the active learning model for samples to label - `update_learner_predictions`: teaches the active learning model the previous queries, and updates the model’s predictions across the dataset Because it uses caching, and it is _not_ a JavaScript plugin, it will only work with the open source version of FiftyOne. ### Creating the Active Learner To create the Active Learner, your dataset must have: 1. Initial labels 2. Input features for the classifiers First, for the purposes of illustration, I have taken a subset of the [Caltech101](https://docs.voxel51.com/user_guide/dataset_zoo/datasets.html#caltech-101) dataset with the labels `airplane`, `Motorbike`, `helicopter`, shuffled the data, and deleted the ground truth labels. In this blog post, we will use Active Learning to label this dataset. If you want to follow along at home, you can do so as follows: If you press “\`” to open up the operators list and select any of this plugin’s three operators, you will be met with warning messages like this: This is because you must have a set of labels to use to initialize the learner. We can either use tags, or we can generate some zero-shot classification predictions with our Week 6 [Zero-shot Prediction Plugin](https://github.com/jacobmarks/zero-shot-prediction-plugin), and use these as a starting point: Regardless of which approach we take to initial labels, the only requirement is that we have at least one example from each class we are trying to classify. You do _not_ need to run zero-shot labeling on the entire dataset, or tag all samples in the dataset. 💡In practice, typically a few examples from each class is good enough to start. The second requirement is that the dataset has at least one field that can be used as an input feature to the “estimators” in the Active Learner. This could be any `fo.FloatField` (float-valued field) or any `fo.VectorField` (one-dimensional numpy array). If you want to use multiple fields as features, they will be concatenated into a one-dimensional feature vector. If you don’t have any candidate feature fields, a good starting point is to use model [embeddings](https://voxel51.com/blog/fiftyone-computer-vision-embeddings-tips-and-tricks-mar-31-2023/).You can compute embeddings in Python by loading a model from the [FiftyOne Model Zoo](https://docs.voxel51.com/user_guide/model_zoo/index.html) and running `compute_embeddings()`. For this walkthrough, I’ll compute embeddings using one more semantic model, [CLIP](https://docs.voxel51.com/user_guide/model_zoo/models.html#clip-vit-base32-torch), and one model trained more to represent data on the level of pixels and patches, [MobileNet](https://docs.voxel51.com/user_guide/model_zoo/models.html#mobilenet-v2-imagenet-torch): This stores the embeddings in the fields `mobilenet_embeddings` and `clip_embeddings` on our samples. I will also use a few additional scalar properties computed using our [Image Quality Issues Plugin](https://github.com/jacobmarks/image-quality-issues). Here is how the entropy is computed: The contrast and brightness are computed in analogous fashion. Now that we have candidate input features on our samples and some initial labels, we can create an active learner! We can choose: - The field or fields to use as a feature vector - The label field in which to store predictions - The default batch size — the number of samples per query - The `Active Learner` For the latter of these, we can select from a variety of ensemble strategies, including [Random Forest](https://en.wikipedia.org/wiki/Random_forest), [Gradient Boosting](https://en.wikipedia.org/wiki/Gradient_boosting), [Bagging](https://en.wikipedia.org/wiki/Bootstrap_aggregating), and [AdaBoost](https://en.wikipedia.org/wiki/AdaBoost). When we make this top-level selection, the remainder of the form dynamically updates with appropriate hyperparameter configuration choices. Executing this operator creates a modAL `ActiveLearner` that uses an “uncertainty” batch sampling. The execution also invokes the generation of initial predictions, and triggers the reload of the dataset. ### **Querying the Active Learner** Now that we have our learner, we can query the learner for the next batch of samples to label. If we’d like, we can override the default query batch size we set earlier: In this case, we can see that the model is least confident about some helicopter images, which it has labeled as a plane. This is likely due to the fact that the dataset has far more planes and motorcycles than helicopters, and helicopters are typically more similar to planes than to motorcycles. We can then tag the mistakes with the correct labels, either within the app, or by sending these samples for reannotation. For the sake of simplicity, we will do so in the app: All samples that are not tagged are treated as having the correct labels when used to teach the learner. ### **Teaching the Active Learner** After correcting the incorrect query labels, we can update our active learner by “teaching” it this new information: Running this operator updates our active learning model, updates the label field with new predictions, and reloads the app. If we were to query the learner again, we would see that the samples in this new query are completely different from the samples in the first query: We can go through this process as many times as we need to until we are confident in our labels. Because this uses uncertainty sampling, rather than random sampling, we should converge to a set of accurate labels very efficiently! ### **Installing the Plugin** If you haven’t already done so, [install FiftyOne](https://docs.voxel51.com/getting_started/install.html): Then you can download this plugin from the command line with: Refresh the FiftyOne App, and you should see the three operators in your operators list when you press the “\`” key. You can install the plugin’s requirements (modAL) by running ## Lessons Learned The Active Learning plugin is a Python Plugin with the usual structure: - `_init__.py`: defining operators - `fiftyone.yml`: registering operators - `README.md`: explaining the plugin, and giving install instructions - `assets` folder: storeroom for icons - `requirements.txt`: list of required Python packages Additionally, the plugin has an `active_learning.py` file, which implements the active learning logic using modAL. ### Caching is King This plugin uses the same caching mechanism from the [Keyword Search plugin](https://github.com/jacobmarks/keyword-search-plugin). For keyword search, we were storing a user’s preference, which made the user’s life easier but was not strictly essential. For Active Learning, however, statefulness is critical. The learner is learning, so we need to keep track of its current state, as well as which samples it has seen before. In theory, this could be done by adding fields onto the dataset to keep track of this information. If you want to implement this version of the Active Learning plugin, I encourage you to do so!! ### Don’t Reinvent the Wheel When I set out to make an Active Learning plugin, I initially tried to implement everything from scratch. Sometimes this is necessary. But in other cases, including this one, robust solutions already exist. FiftyOne is flexible enough to bend and weave into workflows with tons of existing tools. Rather than spending your time reinventing the wheel, integrate these libraries into your FiftyOne Plugins so you can combine the strengths of other special-purpose libraries with FiftyOne’s data management and visualization infrastructure! ## Conclusion In the fast-paced world of machine learning and computer vision, efficiency and effectiveness are of the essence. By intelligently selecting the most crucial samples for labeling, this plugin will help you save time, reduce costs, and enhance the overall quality of your models. Happy learning! Stay tuned over the remaining weeks in the _Ten Weeks of FiftyOne Plugins_ while we continue to pump out a killer lineup of plugins! You can track our journey in our [ten-weeks-of-plugins repo](https://github.com/jacobmarks/ten-weeks-of-plugins) — and I encourage you to fork the repo and join me on this journey! [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [custom plugins](https://voxel51.com/blog/tag/custom-plugins) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [labeling](https://voxel51.com/blog/tag/labeling) [plugins](https://voxel51.com/blog/tag/plugins) MT Admin Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/fedd0c008df994c1834839ea35bc65b34e47fb5d-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Double Trouble: Eliminate Image Duplicates with FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 14, 2023](https://voxel51.com/blog/eliminate-image-duplicates-with-fiftyone) [![](https://cdn.sanity.io/images/h6toihm1/production/bad75ba72dfae8cdefc3d9afe33a1ea9a9c4ec36-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Optical Character Recognition with PyTesseract\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 21, 2023](https://voxel51.com/blog/computer-vision-optical-character-recognition-pytesseract) [![](https://cdn.sanity.io/images/h6toihm1/production/3efef5551e07ae9c6a1d190a5bd256b9723c78c7-1920x1080.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Zero-Shot Prediction Plugin for FiftyOne\\ \\ Computer Vision, Plugins, Tutorials\\ \\ • \\ \\ Sep 28, 2023](https://voxel51.com/blog/computer-vision-zero-shot-prediction-plugin-for-fiftyone) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-343-lllmstxt|> ## FiftyOne Updates Announcement [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Product & News](https://voxel51.com/blog/category/product-news) Announcing Updates to FiftyOne 0.22.1 and FiftyOne Teams 1.4.2 Oct 10, 2023 • 4 min read Article content In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-2#f6c1211eaa13) [Okay, but what’s FiftyOne Teams?](https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-2#8239674db7f8) [What’s new in FiftyOne 0.22.1](https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-2#19ebb5c5b6ab) [What’s new in FiftyOne Teams 1.4.2?](https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-2#f4317d3fb0c5) [Get involved in the FiftyOne open source community!](https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-2#068df34d50a0) In this article [Wait, what’s FiftyOne?](https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-2#f6c1211eaa13) [Okay, but what’s FiftyOne Teams?](https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-2#8239674db7f8) [What’s new in FiftyOne 0.22.1](https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-2#19ebb5c5b6ab) [What’s new in FiftyOne Teams 1.4.2?](https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-2#f4317d3fb0c5) [Get involved in the FiftyOne open source community!](https://voxel51.com/blog/announcing-updates-to-fiftyone-0-22-1-and-fiftyone-teams-1-4-2#068df34d50a0) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) _**Editor's note:** FiftyOne Teams is now [FiftyOne Enterprise](https://voxel51.com/enterprise/)!_ The Voxel51 engineering team is thrilled to announce the general availability of [FiftyOneOne 0.22.1](https://docs.voxel51.com/release-notes.html#fiftyone-0-22-1) and [FiftyOne Teams 1.4.2](https://docs.voxel51.com/release-notes.html#fiftyone-teams-1-4-2), which bring with them dozens of enhancements and fixes to streamline your computer vision workflows. ## **Wait, what’s FiftyOne?** [FiftyOne](https://voxel51.com/fiftyone/) is the open source machine learning toolset that enables data science teams to improve the performance of their computer vision models by helping them curate high quality datasets, evaluate models, find mistakes, visualize embeddings, and get to production faster. Short Tour of FiftyOne Features from Voxel51 on Vimeo ![video thumbnail](https://i.vimeocdn.com/video/1668689272-d4625bc022c5ca5a63ffe9eb115ef133acdab35dbd5d148666d32e1ccd462b3a-d?mw=80&q=85) Playing in picture-in-picture Play 00:00 01:41 Show controls SettingsPicture-in-PictureFullscreen [![Voxel51](https://i.vimeocdn.com/player/754644?sig=afb30b4b06672d28b33cc6f6fddf342dda426ae2e7e5ce1d7441a66b97bf6ba7&v=1)](https://voxel51.com/) QualityAuto SpeedNormal ## **Okay, but what’s FiftyOne Teams?** [FiftyOne Teams](https://voxel51.com/fiftyone-teams/) extends FiftyOne with a GSuite-like experience for teams that want to collaborate on data stored in a centralized location with additional features like user permissions, dataset versioning, cloud-backed media, and enterprise security. ![](https://cdn.sanity.io/images/h6toihm1/production/8d5fc8d7e52d6611e3d5157bab98116e341ac146-512x341.jpg?auto=format&dpr=2&fit=max&q=75&w=512) If this sounds interesting, read on! Then [schedule a workshop](https://voxel51.com/schedule-teams-workshop/) to learn more about FiftyOne Teams. ## **What’s new in FiftyOne 0.22.1** This release includes: ### **FiftyOne App** - Fixed empty detection instance masks [#3559](https://github.com/voxel51/fiftyone/pull/3559) - Fixed a visual issue with scrollbars [#3605](https://github.com/voxel51/fiftyone/pull/3605) - Fixed a bug with color by index for videos [#3606](https://github.com/voxel51/fiftyone/pull/3606) - Fixed an issue where Detections (and other label types) subfields were not properly handling primitive types. [#3577](https://github.com/voxel51/fiftyone/pull/3577) ### **FiftyOne Core** - Resolved groups aggregation issue resulting in unstable ordering of documents [#3614](https://github.com/voxel51/fiftyone/pull/3614) - Fixed an issue where group id indexes were not created against the right id property [#3627](https://github.com/voxel51/fiftyone/pull/3627) - Fixed fiftyone app cells in Databrick notebooks [#3609](https://github.com/voxel51/fiftyone/pull/3609) - Fixed issue with empty segmentation mask conversion in coco formatted datasets [#3595](https://github.com/voxel51/fiftyone/pull/3595/commits/ad0607aeabbd5d6dcbcfccc622ee5caf1f71f930) ### **FiftyOne Plugins** - Added a new fiftyone.plugins.utils module that provides common utilities for plugin development [#3612](https://github.com/voxel51/fiftyone/pull/3612) - Re-enabled text-only placement support when icon is not available [#3593](https://github.com/voxel51/fiftyone/pull/3593) - Added read-only support for FileExplorerView [#3639](https://github.com/voxel51/fiftyone/pull/3597) - The fiftyone delegated launch CLI command will now only run one operation at a time [#3615](https://github.com/voxel51/fiftyone/pull/3615) - Fixed an issue where custom component props were not supported [#3595](https://github.com/voxel51/fiftyone/pull/3549) - Fixed issue where selected\_labels were missing from the ExecutionContext during resolve\_input and resolve\_output. [#3575](https://github.com/voxel51/fiftyone/pull/3574) ## **What’s new in FiftyOne Teams 1.4.2?** Includes all updates from [FiftyOne 0.22.1](https://docs.voxel51.com/release-notes.html#release-notes-v0-22-1), plus: ### **General enhancements and fixes** - Error messages now clearly indicate when attempting to use a duplicate key on datasets a user does not have access to - Fixed issue with setting default access permissions for new datasets - Deleting a dataset now deletes all dataset-related references - Default fields now populate properly when creating a new dataset regardless of client - Improved complex/multi collection aggregations in the api client - Fixed issue where users could not list other users within their own org - Snapshots now properly include all run results - Fixed issue where reverting a snapshot behaved incorrectly in some cases - Fixed Python 3.7 support in the fiftyone-teams SDK ### **FiftyOne App** - Searching users has been improved - Resolved issue with recent views not displaying properly Check out the [release notes](https://docs.voxel51.com/release-notes.html#fiftyone-teams-1-4-0) for a full rundown of additional enhancements and bugfixes in FiftyOne Teams 1.4.2. ## **Get involved in the FiftyOne open source community!** If you are working on computer vision use cases and unstructured data, the FiftyOne community is for you. There are tons of ways to get involved, for example: ### **FiftyOne Community Slack** With over 2,000 members, the community Slack channel is a great place to interact with the FiftyOne developers and exchange solutions with machine learning engineers doing computer vision in production. [https://slack.voxel51.com/](https://slack.voxel51.com/) To make it easy to catch the highlights, every Friday we recap interesting questions and answers from Slack in [Tips & Tricks blog series](https://voxel51.com/blog/category/tips-tricks/). Recent posts include: - [3D Detections – FiftyOne Tips and Tricks](https://voxel51.com/blog/computer-vision-3d-detections-fiftyone-tips-and-tricks-september-29th-2023/) - [Exploring Polylines – FiftyOne Tips and Tricks](https://voxel51.com/blog/computer-vision-exploring-polylines-fiftyone-tips-and-tricks-september-22nd-2023/) - [Creating Pose Skeletons from Scratch – FiftyOne Tips and Tricks](https://voxel51.com/blog/creating-pose-skeletons-from-scratch-fiftyone-tips-and-tricks-sep-15-2023/) - [Dynamic Groups – FiftyOne Tips and Tricks](https://voxel51.com/blog/dynamic-groups-fiftyone-tips-and-tricks-sep-8-2023/) - [Understanding Grouped Datasets](https://voxel51.com/blog/understanding-grouped-datasets-fiftyone-tips-and-tricks-sep-1-2023/) ### **Computer Vision and AI, Machine Learning, and Data Science Meetups** ![](https://cdn.sanity.io/images/h6toihm1/production/4337345c67ae94e238ff52a009f25258b0b521e1-400x400.jpg?auto=format&dpr=2&fit=max&q=75&w=400) Voxel51 sponsors 13 virtual [Computer Vision Meetups](https://www.meetup.com/pro/computer-vision-meetups/) and 12 [AI, Machine Learning and Data Science Meetups](https://www.meetup.com/pro/ai-machine-learning-data-science-network/) around the world with over 16,000 members. (To join, visit the Meetup links and scroll down to find the location friendliest to your time zone.) The Computer Vision Meetups are geared towards data scientists, machine learning engineers, and open source enthusiasts who want to expand their knowledge of computer vision and complementary technologies. We put an emphasis on open source software, and speakers who are computer vision practitioners or academics doing research in the field. Our next Meetup is happening this Thursday: ![](https://cdn.sanity.io/images/h6toihm1/production/810731f9e7e69d3454c8a7064af94931fd4f272f-512x288.jpg?auto=format&dpr=2&fit=max&q=75&w=512) - **Bridging the Gap: Advancing Civil Engineering Inspections with Computer Vision** – Johannes Flotzinger, Civil Engineer & Research
assistant at Universität der Bundeswehr München - **Adapting to Change: Foundation Models, APIs, and the Past, Present and Future of AI Development** – Pietro Bolcato at Kittl - **Deci Diffusion: Triple the Speed of Stable Diffusion**– Harpreet Sahota, DevRel Manager ay Deci.ai ### **FiftyOne on GitHub** ![](https://cdn.sanity.io/images/h6toihm1/production/7b31352fdff9725bf429fc18cd46446f991c15de-200x200.jpg?auto=format&dpr=2&fit=max&q=75&w=200) If you want to start contributing to the FiftyOne project resolving issues, reporting bugs or making enhancements to the Docs, check out these resources: - [Good first issues](https://github.com/voxel51/fiftyone/issues?q=is%3Aopen+is%3Aissue+label%3A%22good+first+issue%22) - Building the Docs ( [video](https://www.youtube.com/watch?v=F707JudeT8E), [instructions](https://github.com/voxel51/fiftyone/tree/develop/docs)) ### **Community Spotlights** ![](https://cdn.sanity.io/images/h6toihm1/production/66349e6fbce9f562907bfc41b3fa3bc69843d67f-1024x1429.png?auto=format&dpr=2&fit=max&q=75&w=1024) Is your organization already using FiftyOne to solve interesting computer vision problems? [Share your success story](https://voxel51.com/fiftyone-computer-vision-success-story-submission/) and claim a box of community rewards as a thank you! [Computer Vision](https://voxel51.com/blog/tag/computer-vision) [FiftyOne](https://voxel51.com/blog/tag/fiftyone) [FiftyOne Teams](https://voxel51.com/blog/tag/fiftyone-teams) ![](https://cdn.sanity.io/images/h6toihm1/production/61fc34503bf2933d596cdf61515a3419a9cff950-300x300.jpg?auto=format&dpr=2&fit=max&q=75&w=42) Ritchie Martori Bio ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ### Related posts Discover more insights and tips to boost your visual AI workflows. [View all](https://voxel51.com/blog) [![](https://cdn.sanity.io/images/h6toihm1/production/fb2b263bea2f9774fd5b36bb18e70f1d0f51dfb2-960x540.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Announcing Updates to FiftyOne 0.22.2 and FiftyOne Teams 1.4.3\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Oct 23, 2023](https://voxel51.com/blog/computer-vision-announcing-updates-to-fiftyone-0-22-2-and-fiftyone-teams-1-4-3) [![](https://cdn.sanity.io/images/h6toihm1/production/aac78c1d106f63b08e8b24e2cc9fc875e8b2506c-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Announcing Updates to FiftyOne 0.22.3 and FiftyOne Teams 1.4.4\\ \\ Computer Vision, Product & News\\ \\ • \\ \\ Nov 6, 2023](https://voxel51.com/blog/computer-vision-fiftyone-0-22-3-and-fiftyone-teams-1-4-4) [![](https://cdn.sanity.io/images/h6toihm1/production/5308632d0c2219e95c8ddb5703900bf3c2afcc7a-1200x675.png?auto=format&dpr=2&fit=max&q=75&w=480)\\ \\ Announcing FiftyOne 0.23 and FiftyOne Teams 1.5\\ \\ Product & News\\ \\ • \\ \\ Dec 6, 2023](https://voxel51.com/blog/announcing-fiftyone-0-23-and-fiftyone-teams-1-5) [Talk to a CV expert](https://voxel51.com/sales) Product [Data Annotation](https://voxel51.com/annotation) [Data Curation](https://voxel51.com/curation) [Model Evaluation](https://voxel51.com/evaluation) [Integrations](https://voxel51.com/integrations) [Plugins](https://voxel51.com/plugins) [Pricing](https://voxel51.com/pricing) Solutions [Agriculture](https://voxel51.com/industries/agriculture) [Autonomous Systems](https://voxel51.com/industries/autonomous-vehicles-systems) [Defense](https://voxel51.com/industries/defense) [Healthcare](https://voxel51.com/industries/healthcare) [Manufacturing](https://voxel51.com/industries/manufacturing) [Retail](https://voxel51.com/industries/retail) [Robotics](https://voxel51.com/industries/robotics) [Security](https://voxel51.com/industries/security) Developers [Documentation](https://docs.voxel51.com/) [Events & Meetups](https://voxel51.com/events) [Computer Vision Glossary](https://voxel51.com/glossary) [Community](https://voxel51.com/community) Resources [Blog](https://voxel51.com/blog) [On-Demand Webinars](https://voxel51.com/webinars) [Customer Stories](https://voxel51.com/customers) [Model Zoo](https://docs.voxel51.com/model_zoo/models.html) [Dataset Zoo](https://docs.voxel51.com/dataset_zoo/datasets.html) [CV Research](https://voxel51.com/research) Company [About Voxel51](https://voxel51.com/about) [Careers](https://voxel51.com/careers) [Press](https://voxel51.com/press) [Discord](https://community.voxel51.com/)[Linkedin](https://www.linkedin.com/company/voxel51/)[Twitter](https://x.com/voxel51)[Youtube](https://www.youtube.com/@voxel51) © 2025 Voxel51 All Rights Reserved [Terms of Service](https://voxel51.com/terms-of-service) [Privacy Policy](https://voxel51.com/privacy-policy) <|firecrawl-page-344-lllmstxt|> ## Reverse Image Search Plugin [Go to homepage](https://voxel51.com/) Menu [Home](https://voxel51.com/) [Book a demo](https://voxel51.com/sales) [Blog](https://voxel51.com/blog) [Computer Vision](https://voxel51.com/blog/category/computer-vision), [Plugins](https://voxel51.com/blog/category/plugins), [Tutorials](https://voxel51.com/blog/category/tutorials) Reverse Image Search Plugin for FiftyOne Oct 12, 2023 • 5 min read Article content In this article [Compare Images from the Internet to Your Computer Vision Dataset](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#8fb6290be697) [Reverse Image Search Plugin ⏪🖼️🔎](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#f06378bcd672) [Plugin Overview & Functionality](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#967b46572933) [Installing the Plugin](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#a5a7e3f4b24a) [Lessons Learned](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#c82bd84299ca) [Conclusion](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#e923fb2268ce) In this article [Compare Images from the Internet to Your Computer Vision Dataset](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#8fb6290be697) [Reverse Image Search Plugin ⏪🖼️🔎](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#f06378bcd672) [Plugin Overview & Functionality](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#967b46572933) [Installing the Plugin](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#a5a7e3f4b24a) [Lessons Learned](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#c82bd84299ca) [Conclusion](https://voxel51.com/blog/computer-vision-reverse-image-search-plugin-for-fiftyone#e923fb2268ce) ### Talk to a computer vision expert [Book a demo](https://voxel51.com/sales) ## Compare Images from the Internet to Your Computer Vision Dataset Welcome to week eight of _Ten Weeks of Plugins_. During these ten weeks, we will be building a FiftyOne plugin (or multiple!) each week and sharing the lessons learned! If you’re new to them, FiftyOne Plugins provide a flexible mechanism for anyone to extend the functionality of their FiftyOne App. You may find the following resources helpful: - [FiftyOne Plugins Repo](https://github.com/voxel51/fiftyone-plugins) - [FiftyOne Plugin Docs](https://docs.voxel51.com/plugins/index.html#downloading-plugins) - Plugins Channel in the [FiftyOne Community Slack](https://slack.voxel51.com/) What we’ve built so far: - Week 0: 🌩️ [Image Quality Issues](https://github.com/jacobmarks/image-quality-issues) & 📈 [Concept Interpolation](https://github.com/jacobmarks/concept-interpolation) - Week 1: 🎨 [AI Art Gallery](https://github.com/jacobmarks/ai-art-gallery) & [Twilio Automation](https://github.com/jacobmarks/twilio-automation-plugin) - Week 2: ❓ [Visual Question Answering](https://github.com/jacobmarks/vqa-plugin) - Week 3: 🎥 [YouTube Player Panel](https://github.com/jacobmarks/fiftyone-youtube-panel-plugin) - Week 4: 🪞 [Image Deduplication](https://github.com/jacobmarks/image-deduplication-plugin) - Week 5: [👓Optical Character Recognition (OCR)](https://github.com/jacobmarks/pytesseract-ocr-plugin) & 🔑 [Keyword Search](https://github.com/jacobmarks/keyword-search-plugin) - Week 6: 🎭 [Zero-shot Prediction](https://github.com/jacobmarks/zero-shot-prediction-plugin) - Week 7: 🏃 [Active Learning](https://github.com/jacobmarks/active-learning-plugin) Ok, let’s dive into this week’s FiftyOne Plugin - [Reverse Image Search](https://github.com/jacobmarks/reverse-image-search-plugin)! ## Reverse Image Search Plugin ⏪🖼️🔎 ![](https://cdn.sanity.io/images/h6toihm1/production/034b6f212e389e37405de69d80daa4c77cddc535-2560x1359.gif?auto=format&dpr=2&fit=max&q=75&w=1600) Back in March 2023, we added native vector search functionality into the FiftyOne library with the release of [FiftyOne 0.20](https://voxel51.com/blog/announcing-fiftyone-0-20/). Since then, users of the FiftyOne library have been able to [leverage vector search engines](https://docs.voxel51.com/user_guide/brain.html#similarity-backends) — at first [Qdrant](https://docs.voxel51.com/integrations/qdrant.html) and [Pinecone](https://docs.voxel51.com/integrations/pinecone.html), and [later](https://medium.com/voxel51/the-computer-vision-interface-for-vector-search-55c14e8a82ac) also [Milvus](https://docs.voxel51.com/integrations/milvus.html) and [LanceDB](https://docs.voxel51.com/integrations/lancedb.html) — to seamlessly search through billion-sample datasets. ![](https://cdn.sanity.io/images/h6toihm1/production/171b7f4c2c4ccff86206d786ba2832453c8f54eb-905x546.gif?auto=format&dpr=2&fit=max&q=75&w=905) Concretely, the way this works is that the user selects an image from their dataset, and the subsequent “similarity search” finds the k [most similar images](https://docs.voxel51.com/user_guide/brain.html#image-similarity) in the dataset by querying the vector search engine. In a similar vein, the user can select an object patch in one of their images and query the vector search engine for the k [most similar object patches](https://docs.voxel51.com/user_guide/brain.html#object-similarity). This functionality has proven incredibly useful for data curation and exploration. As members of the FiftyOne community began incorporating vector search into their workflows, however, a slightly different use-case emerged: _reverse image search_. Users wanted to be able to query their dataset with an image that is _not_ in the dataset, and find the closest matches.For the eighth week of _10 Weeks of Plugins_ **,** I built a Reverse Image Search Plugin to enable this workflow! This plugin leverages the same vector search functionality as is used in image similarity search, but exposes an interface to the user to select a query image that is not part of the dataset. The query image can be dragged and dropped from your local filesystem, or specified via URL. As with the [YouTube Player Panel Plugin](https://github.com/jacobmarks/fiftyone-youtube-panel-plugin), ChatGPT was invaluable in helping me to write the JavaScript code! ## Plugin Overview & Functionality The Reverse Image Search Plugin is a joint Python/JavaScript plugin with two operators: - `open_reverse_image_search_panel`: opens the Reverse Image Search Panel. - `reverse_search_image`: runs the reverse image search on the dataset given the input image. For this walkthrough, I’ll be using a dataset of object patches (individual dog detections) from the [Stanford Dogs dataset](http://vision.stanford.edu/aditya86/ImageNetDogs/). ### Creating the Similarity Index As with FiftyOne’s core similarity search functionality, to run reverse image search on your dataset, you first need to have a similarity index. You can generate a similarity index by running `compute_similarity()` on your dataset from Python, specifying a model from the [FiftyOne Model Zoo](https://docs.voxel51.com/user_guide/model_zoo/index.html), and a vector search engine `backend` to use to construct the index. Here we use a CLIP model to compute embeddings, and Qdrant as our vector database: ```python 1!docker run -p "6333:6333" -p "6334:6334" -d qdrant/qdrant 2 3import fiftyone as fo 4import fiftyone.brain as fob 5import fiftyone.zoo as foz 6dataset = foz.load_zoo_dataset("quickstart") 7# Index images 8fob.compute_similarity( 9 dataset, 10 model="clip-vit-base32-torch", 11 brain_key="clip_sim", 12 backend="qdrant" 13) 14 ``` Alternatively, you can compute similarity from within the FiftyOne App: ![](https://cdn.sanity.io/images/h6toihm1/production/5bee2f9dcc2d883a8d8dd69304f56b1118e728eb-2412x1279.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ### Opening the Panel The `open_reverse_image_search_panel` operator follows the same pattern as in the [Concept Interpolation](https://github.com/jacobmarks/concept-interpolation) and [YouTube Player Panel](https://github.com/jacobmarks/fiftyone-youtube-panel-plugin) plugins. As such, there are three ways to execute the `open_reverse_image_search_panel` operator and open the panel: - Press the Reverse Image Search button in the [Sample Actions Menu](http://samples_grid_actions/): ![](https://cdn.sanity.io/images/h6toihm1/production/7e2f2833155534cbefc21c58bbbeafcec1ebd58f-2560x1356.gif?auto=format&dpr=2&fit=max&q=75&w=1600) - Click on the `+` icon next to the `Samples` tab and select `Reverse Image Search` from the dropdown menu: ![](https://cdn.sanity.io/images/h6toihm1/production/6d085bbf0ffcf71da140df853bad5085e78c9d28-3452x1828.gif?auto=format&dpr=2&fit=max&q=75&w=1600) - Press “\`” to pull up your list of operators, and select `open_reverse_image_search_panel`: ![](https://cdn.sanity.io/images/h6toihm1/production/b60358a6aea13ba04103f95a5c91cab4ea4b9773-2757x1460.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ### Searching the Dataset Once we have a similarity index on our dataset, we can select a query image from our local filesystem: ![](https://cdn.sanity.io/images/h6toihm1/production/9f1be1910fe0e770da4f12cb6ac11ddc9c326be0-2560x1363.gif?auto=format&dpr=2&fit=max&q=75&w=1600) This uses [react-dropzone](https://github.com/react-dropzone/react-dropzone) to handle the input of the image. Alternatively, we can specify an image by passing the URL of a PNG or JPEG file: ![](https://cdn.sanity.io/images/h6toihm1/production/eb3c91e6fa246ebe702c96837d8c10cdd2a6de21-3452x1828.gif?auto=format&dpr=2&fit=max&q=75&w=1600) Whichever of these input options we choose, we will see a preview of the query image in the Reverse Image Search panel. We can also specify the number of results to return, and if we have multiple similarity indexes, we can select the index to use by its [brain key](https://docs.voxel51.com/user_guide/brain.html#image-similarity). Pressing the `SEARCH` button will execute the query and display the similar images in the sample grid: ![](https://cdn.sanity.io/images/h6toihm1/production/54f315300122c6874da6f4122ef205c79a80552d-2759x1465.gif?auto=format&dpr=2&fit=max&q=75&w=1600) ## Installing the Plugin If you haven’t already done so, install FiftyOne: ```bash 1pip install fiftyone ``` Then you can download this plugin from the command line with: ```bash 1fiftyone plugins download https://github.com/jacobmarks/reverse-image-search-plugin ``` Refresh the FiftyOne App, and you should see the Reverse Image Search button show up in the actions menu. ![](https://cdn.sanity.io/images/h6toihm1/production/d9d6c7f21a36e2fa2ad0c7271ed0206b04b2f66a-1594x162.png?auto=format&dpr=2&fit=max&q=75&w=1594) ## Lessons Learned ### Handlin