Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes
📰 ArXiv cs.AI
Learn how Regularized Emphatic Temporal-Difference Learning achieves stability under constant stepsizes in reinforcement learning
Action Steps
- Apply emphatic temporal-difference learning to stabilize off-policy TD updates
- Analyze the mean map contraction of the ETD algorithm
- Use regenerative-cycle analysis to separate the sign of the top Lyapunov exponent from the infinite variance of the follow-on trace
- Implement constant-stepsize sampled dynamics in reinforcement learning algorithms
- Evaluate the stability of ETD under constant stepsizes using Lyapunov exponents
Who Needs to Know This
Researchers and engineers working on reinforcement learning and temporal-difference learning will benefit from this article, as it provides new insights into the stability of ETD under constant stepsizes
Key Insight
💡 ETD stabilizes off-policy TD updates, but its stability under constant stepsizes depends on the contraction of the mean map and the sign of the top Lyapunov exponent
Share This
🤖 New research on Regularized Emphatic Temporal-Difference Learning: stability under constant stepsizes in #reinforcementlearning
Full Article
Title: Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes
Abstract:
arXiv:2609.19170v1 Announce Type: new Abstract: Emphatic temporal-difference learning (ETD) stabilizes the expected off-policy TD update and changes its projection geometry, but neither property determines constant-stepsize sampled dynamics. We construct an ergodic two-state counterexample in which the ETD mean map contracts while the sampled product has a positive top Lyapunov exponent. Regenerative-cycle analysis separates this sign from the infinite variance of the follow-on trace. We introdu
Abstract:
arXiv:2609.19170v1 Announce Type: new Abstract: Emphatic temporal-difference learning (ETD) stabilizes the expected off-policy TD update and changes its projection geometry, but neither property determines constant-stepsize sampled dynamics. We construct an ergodic two-state counterexample in which the ETD mean map contracts while the sampled product has a positive top Lyapunov exponent. Regenerative-cycle analysis separates this sign from the infinite variance of the follow-on trace. We introdu
Related Videos
⚡
You're 1 lesson closer to your goal
Sign in free and we'll turn this lesson into a structured roadmap — starting with ⚡30 free Sparks for your first AI explanation or skill path.
Create free account →No credit card required.
DeepCamp AI