What is ACT?
ACT is a transformer-based policy that takes camera images and the robot's joint positions as input and outputs joint targets for the next several time steps. Predicting a chunk of actions, rather than one action at a time, reduces the compounding errors that build up when a policy makes many small, independent decisions, and it helps the policy reproduce the timing of human demonstrations, including pauses.
ACT is trained as a conditional variational autoencoder (CVAE), which helps it model the variation in how people perform the same task. At run time, ACT can use temporal ensembling: it predicts overlapping chunks at every step and averages the actions they share, which produces smoother motion.
Key takeaways
- ACT predicts chunks of future actions from camera images and joint states, which reduces compounding errors.
- It was introduced with the ALOHA system to learn fine-grained manipulation from a few dozen demonstrations.
- ACT is a standard imitation learning baseline in LeRobot, often the first policy trained on a new SO-101 dataset.
How it works
During training, ACT reads windows of consecutive actions from each demonstration episode, so it learns to predict the next chunk given the current observation. In LeRobot, the chunk length comes from the policy configuration, and the dataset's frame rate determines how far into the future each chunk reaches. At deployment, the policy runs at the robot's control rate and executes actions from the predicted chunk.
Why it matters
ACT matters because it showed that low-cost hardware and a small set of teleoperated demonstrations can produce precise manipulation policies. Because it learns from relatively few episodes, the quality of each episode carries a lot of weight: hesitations, failed grasps, and inconsistent camera placement in the training data show up directly in the policy's behavior.
Frequently asked questions
What is the difference between ACT and action chunking?
Action chunking is the general technique of predicting several future actions at once. ACT is a specific policy, built on a transformer and trained as a CVAE, that uses action chunking.
How many demonstrations does ACT need?
ACT is designed to learn from small datasets, and LeRobot's tutorials suggest starting with about 50 episodes of a simple task. The right number depends on task difficulty and how much variation the demonstrations cover.
Related terms