arXiv:2502.16863cs.MAcs.LG2025-02被引 20

用大模型解决多智能体协作中的贡献评估难题,让机器像人一样判断谁做得好。

Leveraging Large Language Models for Effective and Explainable Multi-Agent Credit Assignment

  • 将贡献评估转化为序列优化与归因两个模式识别任务,用大模型做中央评判
  • 在多个基准测试中显著超越现有方法,尤其在带安全约束的太空场景表现优异
  • 生成带逐帧个体奖励标注的数据集,适合研究可解释性与协作分析

从自动驾驶协同到太空中装配,学习协作行为对机器人实现共同目标至关重要。当前普遍采用集中训练、分散执行范式,但带来新挑战:如何评估每个智能体动作对团队成败的贡献。这一信用分配问题长期未解,且人类手动分析常优于现有方法。我们结合大语言模型在模式识别上达到人类水平的能力,将信用分配重构为序列改进与归因两个模式识别任务,提出新型LLM-MCA方法。该方法利用集中式大模型作为奖励评鉴器,基于各智能体在情景中的个性化贡献数值分解环境奖励,并据此更新智能体策略网络。我们还提出扩展方法LLM-TACA,让大模型通过传递中间目标直接指导各智能体策略。两种方法在多种基准测试中均显著优于现有最优水平,包括层级觅食、机器人仓库及新提出的包含碰撞安全约束的Spaceworld基准。作为副产品,我们生成了大规模轨迹数据集,每时刻均标注了各智能体的奖励信息,来自我们的大模型评鉴器。

原文摘要 · Abstract (English)

Recent work, spanning from autonomous vehicle coordination to in-space assembly, has shown the importance of learning collaborative behavior for enabling robots to achieve shared goals. A common approach for learning this cooperative behavior is to utilize the centralized-training decentralized-execution paradigm. However, this approach also introduces a new challenge: how do we evaluate the contributions of each agent's actions to the overall success or failure of the team. This credit assignment problem has remained open, and has been extensively studied in the Multi-Agent Reinforcement Learning literature. In fact, humans manually inspecting agent behavior often generate better credit evaluations than existing methods. We combine this observation with recent works which show Large Language Models demonstrate human-level performance at many pattern recognition tasks. Our key idea is to reformulate credit assignment to the two pattern recognition problems of sequence improvement and attribution, which motivates our novel LLM-MCA method. Our approach utilizes a centralized LLM reward-critic which numerically decomposes the environment reward based on the individualized contribution of each agent in the scenario. We then update the agents' policy networks based on this feedback. We also propose an extension LLM-TACA where our LLM critic performs explicit task assignment by passing an intermediary goal directly to each agent policy in the scenario. Both our methods far outperform the state-of-the-art on a variety of benchmarks, including Level-Based Foraging, Robotic Warehouse, and our new Spaceworld benchmark which incorporates collision-related safety constraints. As an artifact of our methods, we generate large trajectory datasets with each timestep annotated with per-agent reward information, as sampled from our LLM critics.

多智能体大模型信用分配可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。