Task annotation

Task annotation is the practice of attaching natural-language descriptions of what a robot is doing to recorded episodes, or to segments within them. In robot learning, those descriptions become the instructions that language-conditioned policies and vision-language-action models learn to follow. Task annotations range from one instruction per episode to timed labels for each subtask.

What is task annotation?

Every episode in a robot learning dataset needs an answer to the question "what was the robot trying to do?" Task annotation supplies it. At the simplest level, an episode gets one instruction, such as "pick up the red cube and place it in the bin." In LeRobot datasets, these instructions are stored once in the task metadata and referenced from every frame by a task index.
Richer annotations break an episode into subtasks with start and end times, such as reach, grasp, lift, and place, or add free-form language describing events as they happen. These timed annotations support training on long-horizon tasks, help locate exactly where a failure occurred, and give models more language to learn from.

Key takeaways

  • Task annotations are the natural-language instructions attached to robot episodes and segments.
  • They are the training signal that connects language to behavior in VLAs and language-conditioned policies.
  • Noisy or missing task descriptions are one of the most common quality problems in community robot datasets.

How it works

Task annotations are written during recording, added afterward by people, or proposed by a vision-language model that watches the episode and suggests descriptions for a person to review. LeRobot records one task string per episode by default, and recent versions add optional language columns for persistent instructions and timed events, along with an annotation pipeline that uses a VLM to propose them.

Why it matters

Task annotation matters because a model learns the mapping between words and actions from exactly these strings. Placeholders such as "test," one-word labels, and inconsistent phrasing for the same task teach the wrong mapping or none at all. Teams training on pooled datasets often relabel task descriptions before training, as the SmolVLA team did with a vision-language model.

Frequently asked questions

Where are task annotations stored in a LeRobot dataset?

Task descriptions are stored in the dataset's task metadata and mapped to integer task indices, and each frame references its episode's task index. Optional language columns can hold additional instructions and timed events.

What is the difference between a task annotation and a subtask annotation?

A task annotation describes the goal of a whole episode. Subtask annotations divide the episode into timed steps, such as grasp and place, each with its own start time, end time, and description.

Related terms

black and white photo of Jesse Mostipak
Jesse Mostipak
Director of Growth
Jesse Mostipak is the Director of Growth at Voxel51, where the work is helping humans find and trust what the brand knows, and teaching the Google knowledge graph and the LLMs answering on their behalf to do the same. That question, how knowledge gets built inside a system, is one Jesse has been chasing for years. Earlier versions of it ran through a New York City high school science classroom, data science and machine learning, and developer relations at Kaggle, Posit (formerly RStudio), and Baseten. The answer doesn't change much depending on whether the learner is a teenager, a software engineer, or a knowledge graph. Jesse holds a Master's in Education from CUNY Hunter College. LinkedIn
See all articles by Jesse Mostipak
Last updated October 8, 2026

Building visual or physical AI?

Let's talk.