用最优传输融合各设备奖励函数,提升跨环境协作的强化学习效果。
Can Optimal Transport Improve Federated Inverse Reinforcement Learning?
- 基于最优传输计算奖励函数的几何均值,兼顾分布差异。
- 相比传统平均法,全局奖励估计误差更低,泛化性更强。
- 适合隐私敏感、通信受限的多智能体系统应用。
在机器人与多智能体系统中,大量自主代理常在细微不同的环境中执行共同高层目标。直接聚合数据以学习共享奖励函数通常不可行,因动力学差异、隐私限制及通信带宽有限。本文提出一种基于最优传输的联邦逆强化学习方法:各客户端本地执行轻量级最大熵逆强化学习,随后通过Wasserstein均值融合所得奖励函数,充分考虑其潜在几何结构。我们进一步证明,该均值融合方式相较于联邦学习中的常规参数平均,能获得更准确的全局奖励估计。整体上,本工作提供了一个理论严谨且通信高效的框架,用于生成可在异构代理与环境间泛化的共享奖励函数。
原文摘要 · Abstract (English)
In robotics and multi-agent systems, fleets of autonomous agents often operate in subtly different environments while pursuing a common high-level objective. Directly pooling their data to learn a shared reward function is typically impractical due to differences in dynamics, privacy constraints, and limited communication bandwidth. This paper introduces an optimal transport-based approach to federated inverse reinforcement learning (IRL). Each client first performs lightweight Maximum Entropy IRL locally, adhering to its computational and privacy limitations. The resulting reward functions are then fused via a Wasserstein barycenter, which considers their underlying geometric structure. We further prove that this barycentric fusion yields a more faithful global reward estimate than conventional parameter averaging methods in federated learning. Overall, this work provides a principled and communication-efficient framework for deriving a shared reward that generalizes across heterogeneous agents and environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。