低交互秩结构让离线多智能体强化学习更稳定高效
Exploiting Structure in Offline Multi-Agent RL: The Benefits of Low Interaction Rank
- 引入交互秩概念,发现低秩函数对分布偏移更鲁棒
- 结合正则化与无悔学习,实现高效去中心化训练
- 适合追求稳定性和计算效率的离线多智能体研究者
我们研究离线多智能体强化学习中近似均衡的学习问题。提出一个结构假设——交互秩,并证明低交互秩函数相比一般函数对分布偏移具有显著更强的鲁棒性。基于此,我们展示当使用低交互秩函数类并结合正则化与无悔学习时,可在离线MARL中实现去中心化、计算与统计高效的算法。理论结果得到实验验证,表明具有低交互秩的评价器架构在离线MARL中表现优于常见的单智能体值分解架构。
原文摘要 · Abstract (English)
We study the problem of learning an approximate equilibrium in the offline multi-agent reinforcement learning (MARL) setting. We introduce a structural assumption -- the interaction rank -- and establish that functions with low interaction rank are significantly more robust to distribution shift compared to general ones. Leveraging this observation, we demonstrate that utilizing function classes with low interaction rank, when combined with regularization and no-regret learning, admits decentralized, computationally and statistically efficient learning in offline MARL. Our theoretical results are complemented by experiments that showcase the potential of critic architectures with low interaction rank in offline MARL, contrasting with commonly used single-agent value decomposition architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。