Diffusion Policy

Diffusion Policy is an imitation learning method that represents a robot's behavior as a denoising diffusion process over sequences of actions. Starting from random noise, the policy refines a candidate action sequence step by step, conditioned on recent camera images and robot state. It handles tasks where demonstrations show several valid ways to act, and it is available in LeRobot.

What is Diffusion Policy?

Diffusion Policy applies the idea behind diffusion image generators to robot control. Instead of generating pixels, it generates a short sequence of future actions. Training teaches a network to remove noise from demonstration action sequences, and at run time the policy starts from noise and denoises step by step until it produces a smooth, plausible action sequence for the current observation.
Its main strength is handling multimodal action distributions, meaning cases where human demonstrators solved the same task in different ways, such as going around an obstacle on the left or the right. A policy that averages those demonstrations can end up driving straight into the obstacle. Diffusion Policy can represent both options and commit to one.

Key takeaways

  • Diffusion Policy generates robot action sequences by denoising random noise, conditioned on observations.
  • It represents tasks where demonstrations show more than one valid way to act.
  • It predicts action sequences and executes them in a receding horizon, replanning as new observations arrive.

How it works

The policy observes a short history of camera images and robot state, then runs a fixed number of denoising steps to produce a sequence of future actions. The robot executes the first part of that sequence, the policy observes again, and the process repeats. This receding horizon approach keeps motion smooth while letting the policy react to changes. Because denoising takes several network passes, inference is slower than single-pass policies, and the number of steps trades speed against quality.

Why it matters

Diffusion Policy matters because real demonstrations are rarely consistent, and a policy that can represent several valid behaviors learns more from the same data. It also shows why dataset quality matters: diverse demonstrations help a diffusion policy, while failed or idle segments in the data become behaviors the policy can reproduce.

Frequently asked questions

How is Diffusion Policy different from ACT?

Both predict sequences of actions from demonstrations. ACT is a transformer trained as a conditional variational autoencoder that predicts an action chunk in one pass. Diffusion Policy generates an action sequence through iterative denoising, which takes more compute at run time but represents multimodal behavior well.

Can Diffusion Policy be trained in LeRobot?

Yes. LeRobot includes a Diffusion Policy implementation that trains directly from LeRobot datasets, alongside other imitation learning policies such as ACT.

Related terms

black and white photo of Jesse Mostipak
Jesse Mostipak
Director of Growth
Jesse Mostipak is the Director of Growth at Voxel51, where the work is helping humans find and trust what the brand knows, and teaching the Google knowledge graph and the LLMs answering on their behalf to do the same. That question, how knowledge gets built inside a system, is one Jesse has been chasing for years. Earlier versions of it ran through a New York City high school science classroom, data science and machine learning, and developer relations at Kaggle, Posit (formerly RStudio), and Baseten. The answer doesn't change much depending on whether the learner is a teenager, a software engineer, or a knowledge graph. Jesse holds a Master's in Education from CUNY Hunter College. LinkedIn
See all articles by Jesse Mostipak
Last updated October 8, 2026

Building visual or physical AI?

Let's talk.