Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation

📰 ArXiv cs.AI

arXiv:2605.18591v1 Announce Type: cross Abstract: Natural policy gradients improve optimization by accounting for the geometry of distribution space, but their practical use is limited by the cost of estimating and inverting the Fisher matrix. We present Randomized Advantage Transformation (RAT), a method for estimating Tikhonov-regularized natural policy gradients via direct backpropagation. By applying the Woodbury formula, we reformulate the regularized natural policy gradients as vanilla pol

Published 19 May 2026
Read full paper → ← Back to Reads