What is SmolVLA?
SmolVLA combines a small vision-language model with an action expert, a module that turns the model's understanding of the scene and the instruction into a sequence of robot actions. Its small size is deliberate. Many VLAs have billions of parameters and need large GPUs, while SmolVLA can be fine-tuned and run on a single consumer GPU, which puts VLA work within reach of more teams.
SmolVLA's pretraining data is as notable as the model. According to its paper, it was pretrained on 481 community datasets from the Hugging Face Hub, about 22,900 episodes and 10.6 million frames, most of them recorded on low-cost SO-100 arms. Before training, the authors relabeled noisy task descriptions with a vision-language model and standardized inconsistent camera names across datasets.
Key takeaways
- SmolVLA is a roughly 450-million-parameter VLA from Hugging Face, small enough for consumer hardware.
- It was pretrained on hundreds of community LeRobot datasets rather than a single lab's data.
- Its training pipeline included curation steps such as relabeling task descriptions and standardizing camera names.
How it works
SmolVLA takes one or more camera images, the robot's state, and a natural-language instruction. The vision-language model encodes the images and instruction, and the action expert, trained with flow matching, predicts a chunk of future actions. The model can run with asynchronous inference, predicting the next chunk while the robot executes the current one, which keeps motion smooth on slower hardware. Teams typically fine-tune the pretrained model on a few dozen demonstrations of their own task.
Why it matters
SmolVLA matters because it showed that community data, curated well, can pretrain a useful generalist policy. It also made the cost of messy data concrete: inconsistent camera naming and noisy task descriptions across hundreds of contributors had to be fixed before the data was usable, which is why curation is part of training any model on pooled LeRobot datasets.
Frequently asked questions
How many parameters does SmolVLA have?
About 450 million, small compared with VLAs that have several billion parameters.
What data was SmolVLA trained on?
SmolVLA was pretrained on 481 community LeRobot datasets from the Hugging Face Hub, about 22,900 episodes and 10.6 million frames, according to its paper.
Can SmolVLA be fine-tuned on my own robot data?
Yes. SmolVLA is designed to be fine-tuned on a small LeRobot dataset recorded on the target robot, using LeRobot's training tools.
Related terms