arXiv:2608.08604cs.LG2026-08中稿 · publication in IEE…

让每个智能体按自身偏好学习,无需全局奖励函数。

Multi-Agent Reinforcement Learning via Agent-Specific Preference

论文配图:Multi-Agent Reinforcement Learning via Agent-Specific Preference
图 1 · 摘自论文原文
  • 每个智能体由专属专家通过偏好信号评估,实现去中心化评价。
  • 理论证明优化局部偏好可收敛至纳什均衡策略。
  • 通过单调聚合机制整合偏好,适合异构智能体协作场景。

多智能体强化学习(MARL)是解决复杂协作任务的强大框架,但高度依赖明确的全局奖励函数。在异构智能体系统中,单一标量目标难以捕捉多样化行为。本文提出多智能体偏好集成学习(MAGPIE),通过为每个智能体建模特定偏好来应对挑战。每个智能体由专属专家通过偏好信号评估,避免全局评价需求。理论证明,优化这些去中心化偏好可收敛至纳什均衡策略。为将局部偏好整合为一致的全局目标,我们基于偏好数据构建智能体特定奖励模型,并通过单调聚合机制融合。进一步证明,优化该聚合奖励模型等价于训练纳什均衡策略。在基准多智能体任务和序列生产线任务上的大量实验表明,MAGPIE性能接近奖励工程基线,展示了其在难以精确设计奖励场景中的应用潜力。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL) is a powerful framework for solving complex collaborative tasks, but it relies heavily on well-defined global reward functions. Designing such rewards is challenging, especially in systems with heterogeneous agents, where a single scalar objective may fail to capture diverse behaviors. In this paper, we introduce Multi-AGent Preference-Integrated lEarning (MAGPIE), which addresses these challenges through agent-specific preference modeling. Each agent is evaluated by a dedicated expert through preference signals, eliminating the need for global evaluation. We theoretically prove that optimizing these decentralized preferences converges to a Nash equilibrium policy. To integrate local preferences into a coherent global objective, we construct agent-specific reward models from preference data and combine them via a monotonic aggregation mechanism. We further prove that optimizing this aggregate reward model is equivalent to training the Nash equilibrium policy. Extensive experiments on benchmark multi-agent tasks and a sequential production line task show that MAGPIE achieves performance comparable to reward-engineered baselines, demonstrating its potential to facilitate policy learning in scenarios where precise reward engineering is impractical.

多智能体偏好学习纳什均衡去中心化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。