Decision Making and Reinforcement Learning

External: Coursera Courses ↗ · Coursera

Open Course on External: Coursera

Free to audit · Opens on External: Coursera

Decision Making and Reinforcement Learning

Coursera · Beginner ·🎮 Reinforcement Learning ·5mo ago

Key Takeaways

Introduces decision making and reinforcement learning using utility theory and Markov decision processes

Original Description

This course is an introduction to sequential decision making and reinforcement learning. We start with a discussion of utility theory to learn how preferences can be represented and modeled for decision making. We first model simple decision problems as multi-armed bandit problems in and discuss several approaches to evaluate feedback. We will then model decision problems as finite Markov decision processes (MDPs), and discuss their solutions via dynamic programming algorithms. We touch on the notion of partial observability in real problems, modeled by POMDPs and then solved by online planning methods. Finally, we introduce the reinforcement learning problem and discuss two paradigms: Monte Carlo methods and temporal difference learning. We conclude the course by noting how the two paradigms lie on a spectrum of n-step temporal difference methods. An emphasis on algorithms and examples will be a key part of this course.
AI explanation not available for this lesson yet
This lesson is still being prepared for the AI tutor. In the meantime, explore lessons that are ready.
Browse explainer-ready lessons →

Related Reads

📰
Exploration vs. Exploitation in PPO: How the Policy Learns When to Stop Guessing
Learn how PPO reinforcement learning agents balance exploration and exploitation to optimize their strategies
Medium · Machine Learning
📰
RL 3: Bellman Equations and Markov Decision Processes (1950s–1960s)
Learn about the historical development of Reinforcement Learning, specifically Bellman Equations and Markov Decision Processes from the 1950s-1960s
Dev.to AI
📰
New ‘Reinforcement Learning For Calibrated Decisions’ Makes AI Headlines But Look Past The Hype
Learn about Reinforcement Learning for Calibrated Decisions (RLCD) and its potential impact on AI decision-making, beyond the hype
Forbes Innovation
📰
EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning
Learn how EvoRS enables on-policy self-evolution of reward systems for open-ended reinforcement learning, improving adaptability and reducing reward hacking
ArXiv cs.AI
Up next
Middle Management Meritocracy: Shockingly Naive
iBankerU
Watch →