Action Chunking with Transformers (ACT)

Action Chunking with Transformers (ACT) is an imitation learning policy that predicts a chunk of future robot actions at each step instead of a single action. It was introduced alongside the low-cost ALOHA bimanual system to learn fine-grained manipulation from a small number of human demonstrations. ACT is one of the most common baseline policies in LeRobot.

What is ACT?

ACT is a transformer-based policy that takes camera images and the robot's joint positions as input and outputs joint targets for the next several time steps. Predicting a chunk of actions, rather than one action at a time, reduces the compounding errors that build up when a policy makes many small, independent decisions, and it helps the policy reproduce the timing of human demonstrations, including pauses.
ACT is trained as a conditional variational autoencoder (CVAE), which helps it model the variation in how people perform the same task. At run time, ACT can use temporal ensembling: it predicts overlapping chunks at every step and averages the actions they share, which produces smoother motion.

Key takeaways

  • ACT predicts chunks of future actions from camera images and joint states, which reduces compounding errors.
  • It was introduced with the ALOHA system to learn fine-grained manipulation from a few dozen demonstrations.
  • ACT is a standard imitation learning baseline in LeRobot, often the first policy trained on a new SO-101 dataset.

How it works

During training, ACT reads windows of consecutive actions from each demonstration episode, so it learns to predict the next chunk given the current observation. In LeRobot, the chunk length comes from the policy configuration, and the dataset's frame rate determines how far into the future each chunk reaches. At deployment, the policy runs at the robot's control rate and executes actions from the predicted chunk.

Why it matters

ACT matters because it showed that low-cost hardware and a small set of teleoperated demonstrations can produce precise manipulation policies. Because it learns from relatively few episodes, the quality of each episode carries a lot of weight: hesitations, failed grasps, and inconsistent camera placement in the training data show up directly in the policy's behavior.

Frequently asked questions

What is the difference between ACT and action chunking?

Action chunking is the general technique of predicting several future actions at once. ACT is a specific policy, built on a transformer and trained as a CVAE, that uses action chunking.

How many demonstrations does ACT need?

ACT is designed to learn from small datasets, and LeRobot's tutorials suggest starting with about 50 episodes of a simple task. The right number depends on task difficulty and how much variation the demonstrations cover.

Related terms

black and white photo of Jesse Mostipak
Jesse Mostipak
Director of Growth
Jesse Mostipak is the Director of Growth at Voxel51, where the work is helping humans find and trust what the brand knows, and teaching the Google knowledge graph and the LLMs answering on their behalf to do the same. That question, how knowledge gets built inside a system, is one Jesse has been chasing for years. Earlier versions of it ran through a New York City high school science classroom, data science and machine learning, and developer relations at Kaggle, Posit (formerly RStudio), and Baseten. The answer doesn't change much depending on whether the learner is a teenager, a software engineer, or a knowledge graph. Jesse holds a Master's in Education from CUNY Hunter College. LinkedIn
See all articles by Jesse Mostipak
Last updated October 8, 2026

Building visual or physical AI?

Let's talk.