Code-First
Every concept has runnable Python. No hand-waving - you'll implement RL algorithms from scratch.
Just Enough Math
Clear explanations with rendered equations. Enough rigor to understand why things work.
Real-World Examples
From game AI to robotics to LLM alignment - see how RL solves actual problems.
16 Modules, Zero to Mastery
A structured path from bandits and MDPs to PPO, RLHF, multi-agent RL, and beyond.
Introduction
Why RL matters, setup your lab, run your first agent
Math Foundations
Probability, linear algebra, calculus for RL
Multi-Armed Bandits
Exploration vs exploitation, UCB, Thompson sampling
MDPs & Dynamic Programming
Bellman equations, value/policy iteration
MC & TD Methods
Monte Carlo, SARSA, Q-learning, eligibility traces
Function Approximation & DQN
Neural networks for value functions, replay buffers
Policy Gradients
REINFORCE, actor-critic, advantage estimation
Modern Deep RL
PPO, SAC, TD3 - the algorithms that actually work
Reward Design & Debugging
Reward shaping, diagnostics, why your agent is broken
Partial Observability
POMDPs, recurrent policies, generalization
Model-Based RL
World models, planning, Dyna, MuZero
RLHF & LLM Alignment
Reward models, PPO for language, DPO
Multi-Agent RL
Cooperation, competition, communication
Sim-to-Real & Production
Domain randomization, deployment, safety
Paper Reproduction
Pick a paper, reproduce it, understand the gaps
Capstone
Your portfolio project - design, build, evaluate
Ready to train your first agent?
All you need is Python 3.12+ and a willingness to watch training curves for longer than is healthy.
Get Started