✕ Clear all filters
223 articles
▶ Videos →

Research Papers

223 articles · Updated every 3 hours · View all reads

All Articles 187,381Blog Posts 169,836Tech Tutorials 50,115Research Papers 36,757News 23,052 ⚡ AI Lessons
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1w ago
ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR
arXiv:2609.09075v2 Announce Type: replace-cross Abstract: In reinforcement learning with verifiable rewards (RLVR) trained with group relative policy optimizati
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper ⚡ AI Lesson 1w ago
EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning
arXiv:2609.12459v1 Announce Type: new Abstract: Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable ans
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1w ago
Direct Preference Density Alignment for Conversational Audio Equalization
arXiv:2609.12607v1 Announce Type: cross Abstract: Large Language Model alignment typically relies on learned proxy reward models, which significantly increase t
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper ⚡ AI Lesson 2w ago
BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL
arXiv:2609.09783v1 Announce Type: cross Abstract: Asynchronous reinforcement learning has become the standard way to scale training for language models, but the
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper ⚡ AI Lesson 2w ago
Reinforcement Learning with Temporal-Logic-Based Causal Diagrams
arXiv:2306.13732v2 Announce Type: replace Abstract: We study a class of reinforcement learning (RL) tasks where the objective of the agent is to accomplish temp
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper ⚡ AI Lesson 2w ago
Reinforcement learning for Quantum Tiq-Taq-Toe
arXiv:2411.06429v2 Announce Type: replace Abstract: Quantum Tiq-Taq-Toe is a well-known benchmark and playground for both quantum computing and machine learning
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 3w ago
BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning
arXiv:2608.30283v1 Announce Type: cross Abstract: Expected-cost constraints can still permit rare, high-cost events. Monte Carlo conditional value at risk (CVaR
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 3w ago
SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning
arXiv:2608.19842v2 Announce Type: replace Abstract: Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language model
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 4w ago
Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics
arXiv:2507.00611v2 Announce Type: replace-cross Abstract: Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward de
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1mo ago
Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping
arXiv:2608.24135v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has emerged as a pivotal technique for enhancing the code
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1mo ago
Mode-Dependent Rectification for Stable PPO Training
arXiv:2602.05619v2 Announce Type: replace-cross Abstract: Mode-dependent architectural components (layers that behave differently during training and evaluation
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1mo ago
AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
arXiv:2508.14313v4 Announce Type: replace-cross Abstract: Test-time scaling strategies for Large Language Models predominantly rely on either reinforcement lear
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1mo ago
Efficient Exploration at Scale
arXiv:2603.17378v2 Announce Type: replace-cross Abstract: We develop an online learning algorithm that dramatically improves the data efficiency of reinforcemen
MarkTechPost 🎮 Reinforcement Learning 1mo ago
Auditing Preference Biases and Fine-Tuning Language Models with Direct Preference Optimization on Anthropic HH-RLHF Using TRL and LoRA
This tutorial provides an end-to-end workflow for fine-tuning language models using Direct Preference Optimization (DPO). We demonstrate how to audit the Anthro
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1mo ago
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
arXiv:2608.17393v1 Announce Type: new Abstract: Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool inte
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1mo ago
Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation
arXiv:2608.15680v1 Announce Type: cross Abstract: Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors,
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1mo ago
Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet
arXiv:2608.11226v1 Announce Type: new Abstract: Reinforcement-learning post-training dominates modern language-model development, yet its power behavior on GPU
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper ⚡ AI Lesson 1mo ago
Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach
arXiv:2608.11245v1 Announce Type: new Abstract: Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds,
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1mo ago
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL
arXiv:2608.11669v1 Announce Type: cross Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to
ArXiv cs.AI 🎮 Reinforcement Learning 📄 Paper 1mo ago
A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
arXiv:2512.06547v4 Announce Type: replace-cross Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data s