Module 06: Policy Gradients & Actor-Critic
Coming soon.
Instead of learning values and deriving a policy, what if we just learned the policy directly?
Learning Objectives
- Derive the policy gradient theorem
- Implement REINFORCE with baseline subtraction
- Build an Actor-Critic agent
- Understand variance reduction techniques
Concept Explanation
Coming soon.
Code Examples
Coming soon.
Exercises
Coming soon.
Was this page helpful?