SmolVLA

SmolVLA is a compact, open source vision-language-action (VLA) model that Hugging Face released in 2025 as part of LeRobot. With about 450 million parameters, it maps camera images, robot state, and a language instruction to robot actions. SmolVLA was pretrained on community-contributed LeRobot datasets and is designed to run on consumer hardware.

What is SmolVLA?

SmolVLA combines a small vision-language model with an action expert, a module that turns the model's understanding of the scene and the instruction into a sequence of robot actions. Its small size is deliberate. Many VLAs have billions of parameters and need large GPUs, while SmolVLA can be fine-tuned and run on a single consumer GPU, which puts VLA work within reach of more teams.
SmolVLA's pretraining data is as notable as the model. According to its paper, it was pretrained on 481 community datasets from the Hugging Face Hub, about 22,900 episodes and 10.6 million frames, most of them recorded on low-cost SO-100 arms. Before training, the authors relabeled noisy task descriptions with a vision-language model and standardized inconsistent camera names across datasets.

Key takeaways

  • SmolVLA is a roughly 450-million-parameter VLA from Hugging Face, small enough for consumer hardware.
  • It was pretrained on hundreds of community LeRobot datasets rather than a single lab's data.
  • Its training pipeline included curation steps such as relabeling task descriptions and standardizing camera names.

How it works

SmolVLA takes one or more camera images, the robot's state, and a natural-language instruction. The vision-language model encodes the images and instruction, and the action expert, trained with flow matching, predicts a chunk of future actions. The model can run with asynchronous inference, predicting the next chunk while the robot executes the current one, which keeps motion smooth on slower hardware. Teams typically fine-tune the pretrained model on a few dozen demonstrations of their own task.

Why it matters

SmolVLA matters because it showed that community data, curated well, can pretrain a useful generalist policy. It also made the cost of messy data concrete: inconsistent camera naming and noisy task descriptions across hundreds of contributors had to be fixed before the data was usable, which is why curation is part of training any model on pooled LeRobot datasets.

Frequently asked questions

How many parameters does SmolVLA have?

About 450 million, small compared with VLAs that have several billion parameters.

What data was SmolVLA trained on?

SmolVLA was pretrained on 481 community LeRobot datasets from the Hugging Face Hub, about 22,900 episodes and 10.6 million frames, according to its paper.

Can SmolVLA be fine-tuned on my own robot data?

Yes. SmolVLA is designed to be fine-tuned on a small LeRobot dataset recorded on the target robot, using LeRobot's training tools.

Related terms

black and white photo of Jesse Mostipak
Jesse Mostipak
Director of Growth
Jesse Mostipak is the Director of Growth at Voxel51, where the work is helping humans find and trust what the brand knows, and teaching the Google knowledge graph and the LLMs answering on their behalf to do the same. That question, how knowledge gets built inside a system, is one Jesse has been chasing for years. Earlier versions of it ran through a New York City high school science classroom, data science and machine learning, and developer relations at Kaggle, Posit (formerly RStudio), and Baseten. The answer doesn't change much depending on whether the learner is a teenager, a software engineer, or a knowledge graph. Jesse holds a Master's in Education from CUNY Hunter College. LinkedIn
See all articles by Jesse Mostipak
Last updated October 8, 2026

Building visual or physical AI?

Let's talk.