Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement Learning

📰 ArXiv cs.AI

arXiv:2602.17062v2 Announce Type: replace Abstract: Value decomposition is a core approach for cooperative multi-agent reinforcement learning (MARL). However, existing methods still rely on a single optimal action and struggle to adapt when the underlying value function shifts during training, often converging to suboptimal policies. To address this limitation, we propose Successive Sub-value Q-learning (S2Q), which learns multiple sub-value functions to retain alternative high-value actions. In

Published 21 May 2026

Full Article

Title: Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement Learning

Abstract:
arXiv:2602.17062v2 Announce Type: replace Abstract: Value decomposition is a core approach for cooperative multi-agent reinforcement learning (MARL). However, existing methods still rely on a single optimal action and struggle to adapt when the underlying value function shifts during training, often converging to suboptimal policies. To address this limitation, we propose Successive Sub-value Q-learning (S2Q), which learns multiple sub-value functions to retain alternative high-value actions. In
Read full paper → ← Back to Reads