Module 02: Multi-Armed Bandits & Exploration
Coming soon.
The simplest RL setting - and the best place to build intuition for the exploration vs exploitation tradeoff.
Learning Objectives
- Understand the multi-armed bandit problem
- Implement epsilon-greedy, UCB, and Thompson Sampling
- Analyze regret bounds and sample complexity
- Bridge from bandits to contextual bandits and full RL
Concept Explanation
Coming soon.
Code Examples
Coming soon.
Exercises
Coming soon.
Was this page helpful?