Skip to main content

Module 07: Modern Deep RL (PPO, SAC, TD3)

Coming soon.

The algorithms that actually get used in production. PPO for its reliability, SAC for continuous control, TD3 for when you need something between them.

Learning Objectives​

  • Understand trust regions and PPO's clipped objective
  • Implement PPO and compare with SB3
  • Learn SAC's maximum entropy framework
  • Compare DDPG, TD3, and SAC on continuous control

Concept Explanation​

Coming soon.

Code Examples​

Coming soon.

Exercises​

Coming soon.

Was this page helpful?