Improving Generalization and Data Efficiency with Diffusion in Offline Multi-agent RL

📰 ArXiv cs.AI

arXiv:2307.01472v2 Announce Type: replace Abstract: We present a novel Diffusion Offline Multi-agent Model (DOM2) for offline Multi-Agent Reinforcement Learning (MARL). Different from existing algorithms that rely mainly on conservatism in policy design, DOM2 enhances policy expressiveness and diversity based on diffusion model. Specifically, we incorporate a diffusion model into the policy network and propose a trajectory-based data-reweighting scheme in training. These key ingredients signific

Published 11 Jun 2026
Read full paper → ← Back to Reads