Sample-efficient Neuro-symbolic Proximal Policy Optimization

📰 ArXiv cs.AI

Learn how to improve sample efficiency in deep reinforcement learning using neuro-symbolic proximal policy optimization, enabling better performance in sparse-reward domains

advanced Published 29 Apr 2026
Action Steps
  1. Implement Proximal Policy Optimization (PPO) algorithm
  2. Integrate symbolic guidance into PPO using partial logical policy specifications
  3. Transfer learned policies from easier instances to more challenging settings
  4. Evaluate the performance of the neuro-symbolic PPO in sparse-reward domains
  5. Compare the sample efficiency of the proposed method with traditional PPO
Who Needs to Know This

Researchers and engineers working on reinforcement learning and robotics can benefit from this technique to improve the efficiency of their algorithms, especially in complex domains with multiple sub-goals

Key Insight

💡 Neuro-symbolic proximal policy optimization can significantly improve sample efficiency in deep reinforcement learning, especially in sparse-reward domains

Share This
🤖 Improve sample efficiency in #ReinforcementLearning with neuro-symbolic PPO! 📈

Key Takeaways

Learn how to improve sample efficiency in deep reinforcement learning using neuro-symbolic proximal policy optimization, enabling better performance in sparse-reward domains

Full Article

Title: Sample-efficient Neuro-symbolic Proximal Policy Optimization

Abstract:
arXiv:2604.25534v1 Announce Type: new Abstract: Deep Reinforcement Learning (DRL) algorithms often require a large amount of data and struggle in sparse-reward domains with long planning horizons and multiple sub-goals. In this paper, we propose a neuro-symbolic extension of Proximal Policy Optimization (PPO) that transfers partial logical policy specifications learned in easier instances to guide learning in more challenging settings. We introduce two integrations of symbolic guidance: (i) H-PP
Read full paper → ← Back to Reads

Related Videos

Generative vs Discriminative Models - Explained
Generative vs Discriminative Models - Explained
DataMListic
Class 14 Machine Learning ( S 2 25 AIMLZG 565) Prof. Kiruthiga A R 20260503 133253 Meeting Recording
Class 14 Machine Learning ( S 2 25 AIMLZG 565) Prof. Kiruthiga A R 20260503 133253 Meeting Recording
Karthik Sundara Rajan
Class 13 Machine Learning ( S 2 25 AIMLZG 565) Prof. Kiruthiga A R 20260426 133418 Meeting Recording
Class 13 Machine Learning ( S 2 25 AIMLZG 565) Prof. Kiruthiga A R 20260426 133418 Meeting Recording
Karthik Sundara Rajan
Class 15 Machine Learning ( S 2 25 AIMLZG 565) Prof. Kiruthiga A R 20260510 133228 Meeting Recording
Class 15 Machine Learning ( S 2 25 AIMLZG 565) Prof. Kiruthiga A R 20260510 133228 Meeting Recording
Karthik Sundara Raajan
Class 12 Machine Learning ( S 2 25 AIMLZG 565) Prof. Kiruthiga A R 20260425 133139 Meeting Recording
Class 12 Machine Learning ( S 2 25 AIMLZG 565) Prof. Kiruthiga A R 20260425 133139 Meeting Recording
Karthik Sundara Rajan
Class 11 Machine Learning ( S 2 25 AIMLZG 565) Prof. Kiruthiga A R 20260412 133157 Meeting Recording
Class 11 Machine Learning ( S 2 25 AIMLZG 565) Prof. Kiruthiga A R 20260412 133157 Meeting Recording
Karthik Sundara Rajan