Skip to main content

Module 02: Multi-Armed Bandits & Exploration

Coming soon.

The simplest RL setting - and the best place to build intuition for the exploration vs exploitation tradeoff.

Learning Objectives​

  • Understand the multi-armed bandit problem
  • Implement epsilon-greedy, UCB, and Thompson Sampling
  • Analyze regret bounds and sample complexity
  • Bridge from bandits to contextual bandits and full RL

Concept Explanation​

Coming soon.

Code Examples​

Coming soon.

Exercises​

Coming soon.

Was this page helpful?