Reinforcement learning

Reinforcement learning trains an agent to make sequences of decisions by rewarding good outcomes and penalizing bad ones. Through trial and error in an environment, the agent learns a policy that maximizes cumulative reward, which is central to robotics and control.

What is reinforcement learning?

Reinforcement learning is a way to train an agent to act well over time. The agent observes the state of an environment, takes an action, and receives a reward signal that indicates how good the outcome was. By repeating this loop and favoring actions that lead to higher long-term reward, it gradually learns a policy, a mapping from states to actions, without being told the correct action at each step.
Unlike supervised learning, there are no labeled examples of the right answer, only rewards that the agent must learn to maximize.

Key takeaways

  • An agent learns by interacting with an environment and receiving rewards.
  • The goal is a policy that maximizes cumulative long-term reward.
  • It suits sequential decision problems like control and robotics.

How it works

The problem is framed as states, actions, and rewards. The agent explores actions, observes the resulting states and rewards, and updates its policy to prefer actions with higher expected return. Balancing exploration of new actions against exploitation of known-good ones is a central challenge, and training is often done in simulation before transferring to the real world.

Why it matters

Reinforcement learning is a natural fit for embodied and control tasks where decisions unfold over time and outcomes depend on sequences of actions, which places it at the heart of robot learning. It also connects to policies learned by imitation and to the sim-to-real gap that physical AI systems must cross.

Frequently asked questions

How is reinforcement learning different from supervised learning?

Supervised learning trains on labeled input-output pairs, while reinforcement learning has no labels and instead learns from reward signals earned through interaction.

Why is simulation common in reinforcement learning?

Trial-and-error learning can be slow, unsafe, or costly in the real world, so agents are often trained in simulation and then transferred to real environments.

Related terms

Last updated July 9, 2026

Building visual or physical AI?

Let's talk.