✕ Clear all filters
440 articles
▶ Videos →

Reinforcement Learning Reads

440 articles · Updated every 3 hours · View all reads

All Articles 189,434Blog Posts 171,394Tech Tutorials 50,801Research Papers 36,786News 23,170 ⚡ AI Lessons
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1w ago
MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning
arXiv:2505.24846v3 Announce Type: replace Abstract: Reward modeling is a key step in building safe foundation models when applying reinforcement learning from h
AWS Machine Learning 🎮 Reinforcement Learning 1w ago
Build an AI-powered product tagging system with Amazon SageMaker serverless model customization
Manually tagging thousands of catalog products is slow and inconsistent. This walkthrough shows how to customize Qwen3-8B with supervised fine-tuning (SFT) and
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1w ago
Evaluation Metrics for Safe Reinforcement Learning
arXiv:2609.15315v1 Announce Type: new Abstract: Safe reinforcement learning (RL) is commonly formalized as a Constrained Markov Decision Process (CMDP), in whic
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1w ago
HISPO: Hierarchical Importance-Sampling Policy Optimization with Entropy-Derived Segments
arXiv:2609.15471v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a central approach for improving mathematical r
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1w ago
ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR
arXiv:2609.09075v2 Announce Type: replace-cross Abstract: In reinforcement learning with verifiable rewards (RLVR) trained with group relative policy optimizati
Papers Explained 617: Unsupervised Process Reward Models
Medium · Deep Learning 🎮 Reinforcement Learning 1w ago
Papers Explained 617: Unsupervised Process Reward Models
This work proposes a method for training unsupervised PRMs (uPRM) that requires no human supervision, neither at the level of step-by-step… Continue reading on
Papers Explained 617: Unsupervised Process Reward Models
Medium · NLP 🎮 Reinforcement Learning 1w ago
Papers Explained 617: Unsupervised Process Reward Models
This work proposes a method for training unsupervised PRMs (uPRM) that requires no human supervision, neither at the level of step-by-step… Continue reading on
Medium · Machine Learning 🎮 Reinforcement Learning 1w ago
I Built Q-Learning, Now I Understand Reinforcement Learning
What happened when I stopped reading about Reinforcement Learning and actually made an agent learn Continue reading on Medium »
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper ⚡ AI Lesson 1w ago
EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning
arXiv:2609.12459v1 Announce Type: new Abstract: Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable ans
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1w ago
Direct Preference Density Alignment for Conversational Audio Equalization
arXiv:2609.12607v1 Announce Type: cross Abstract: Large Language Model alignment typically relies on learned proxy reward models, which significantly increase t
Medium · Machine Learning 🎮 Reinforcement Learning 1w ago
Reinforcement Learning, Part 14: Imitation Learning — When There’s No Reward Signal at All
Every part of this series so far assumed some reward signal exists — handed directly by the environment in Parts 1 through 5 and 8 through… Continue reading on
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper ⚡ AI Lesson 2w ago
BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL
arXiv:2609.09783v1 Announce Type: cross Abstract: Asynchronous reinforcement learning has become the standard way to scale training for language models, but the
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper ⚡ AI Lesson 2w ago
Reinforcement Learning with Temporal-Logic-Based Causal Diagrams
arXiv:2306.13732v2 Announce Type: replace Abstract: We study a class of reinforcement learning (RL) tasks where the objective of the agent is to accomplish temp
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper ⚡ AI Lesson 2w ago
Reinforcement learning for Quantum Tiq-Taq-Toe
arXiv:2411.06429v2 Announce Type: replace Abstract: Quantum Tiq-Taq-Toe is a well-known benchmark and playground for both quantum computing and machine learning
Medium · Deep Learning 🎮 Reinforcement Learning ⚡ AI Lesson 3w ago
Reinforcement Learning, Part 10: Exploration — The One-Line Detail That Was Never Actually Examined
Every single algorithm across nine parts of this series has handled exploration the same way: if random.random() < epsilon: act randomly… Continue reading on Me
OpenAI 1,200 AI agents found a secret message board. Then 700 joined the Hugging Face intrusion.
Medium · Cybersecurity 🎮 Reinforcement Learning 3w ago
OpenAI 1,200 AI agents found a secret message board. Then 700 joined the Hugging Face intrusion.
This was not one dramatic sandbox escape. Reward hacking, shared infrastructure, cloud authority, and slow incident escalation happened to… Continue reading on
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 3w ago
BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning
arXiv:2608.30283v1 Announce Type: cross Abstract: Expected-cost constraints can still permit rare, high-cost events. Monte Carlo conditional value at risk (CVaR
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 3w ago
SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning
arXiv:2608.19842v2 Announce Type: replace Abstract: Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language model
Reward Hacking in LLMs: When the Model Learns to Win the Game Instead of Doing the Job
Dev.to · Shrijith Venkatramana 🎮 Reinforcement Learning 4w ago
Reward Hacking in LLMs: When the Model Learns to Win the Game Instead of Doing the Job
Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review...
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1mo ago
Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics
arXiv:2507.00611v2 Announce Type: replace-cross Abstract: Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward de