Representation over Routing: Overcoming Surrogate Hacking in Multi-Timescale PPO

📰 ArXiv cs.AI

arXiv:2604.13517v1 Announce Type: cross Abstract: Temporal credit assignment in reinforcement learning has long been a central challenge. Inspired by the multi-timescale encoding of the dopamine system in neurobiology, recent research has sought to introduce multiple discount factors into Actor-Critic architectures, such as Proximal Policy Optimization (PPO), to balance short-term responses with long-term planning. However, this paper reveals that blindly fusing multi-timescale signals in comple

Published 16 Apr 2026
Read full paper → ← Back to Reads