What is Diffusion Policy?
Diffusion Policy applies the idea behind diffusion image generators to robot control. Instead of generating pixels, it generates a short sequence of future actions. Training teaches a network to remove noise from demonstration action sequences, and at run time the policy starts from noise and denoises step by step until it produces a smooth, plausible action sequence for the current observation.
Its main strength is handling multimodal action distributions, meaning cases where human demonstrators solved the same task in different ways, such as going around an obstacle on the left or the right. A policy that averages those demonstrations can end up driving straight into the obstacle. Diffusion Policy can represent both options and commit to one.
Key takeaways
- Diffusion Policy generates robot action sequences by denoising random noise, conditioned on observations.
- It represents tasks where demonstrations show more than one valid way to act.
- It predicts action sequences and executes them in a receding horizon, replanning as new observations arrive.
How it works
The policy observes a short history of camera images and robot state, then runs a fixed number of denoising steps to produce a sequence of future actions. The robot executes the first part of that sequence, the policy observes again, and the process repeats. This receding horizon approach keeps motion smooth while letting the policy react to changes. Because denoising takes several network passes, inference is slower than single-pass policies, and the number of steps trades speed against quality.
Why it matters
Diffusion Policy matters because real demonstrations are rarely consistent, and a policy that can represent several valid behaviors learns more from the same data. It also shows why dataset quality matters: diverse demonstrations help a diffusion policy, while failed or idle segments in the data become behaviors the policy can reproduce.
Frequently asked questions
How is Diffusion Policy different from ACT?
Both predict sequences of actions from demonstrations. ACT is a transformer trained as a conditional variational autoencoder that predicts an action chunk in one pass. Diffusion Policy generates an action sequence through iterative denoising, which takes more compute at run time but represents multimodal behavior well.
Can Diffusion Policy be trained in LeRobot?
Yes. LeRobot includes a Diffusion Policy implementation that trains directly from LeRobot datasets, alongside other imitation learning policies such as ACT.
Related terms