Skills › Reinforcement Learning

Policy Gradient Methods

Implement policy gradient algorithms — REINFORCE, PPO, and Actor-Critic.

0%
Confidence · no data yet
Sign in to track

After this skill you can…

  • Implement REINFORCE from scratch
  • Train a PPO agent with Stable-Baselines3
  • Explain the advantage function in Actor-Critic

Prerequisites

Watch (10 videos)

Direct Preference Optimization (DPO): End-to-End Implementation
SH AI Academy · intermediate
→ Optimize DPO policies→ Improve DPO stability and efficiency
Lecture 12: Other Social Insurance Programs
MIT OpenCourseWare · intermediate
→ Evaluate public policy effectiveness→ Design social insurance programs
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 6: Q-Learning
Stanford Online · beginner
→ Learn policy gradient methods→ Understand RL without explicit policy learning
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 8: Reward Learning
Stanford Online · beginner
→ Apply inverse reinforcement learning techniques→ Implement reward shaping methods
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 4: Actor-Critic Methods
Stanford Online · intermediate hands-on
→ Implement actor-critic methods→ Use policy estimation to improve RL algorithms
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 13: Meta RL
Stanford Online · intermediate hands-on
→ Design meta-RL algorithms→ Implement black-box meta-RL methods
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 15: Hierarchical RL and IL
Stanford Online · intermediate
→ Implement policy gradient methods→ Optimize RL policies
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 3: Policy Gradients
Stanford Online · intermediate
→ Apply policy gradient methods→ Analyze policy optimization problems
Q-learning with Flow-Matching Policies
Microsoft Research · beginner
→ Optimize policies for robotic manipulation→ Apply reinforcement learning to real-world problems
Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 2: Imitation Learning
Stanford Online · intermediate
→ Implement Policy Gradient Methods→ Understand Imitation Learning Basics