Module 07: Modern Deep RL (PPO, SAC, TD3)
Coming soon.
The algorithms that actually get used in production. PPO for its reliability, SAC for continuous control, TD3 for when you need something between them.
Learning Objectives
- Understand trust regions and PPO's clipped objective
- Implement PPO and compare with SB3
- Learn SAC's maximum entropy framework
- Compare DDPG, TD3, and SAC on continuous control
Concept Explanation
Coming soon.
Code Examples
Coming soon.
Exercises
Coming soon.
Was this page helpful?