用多模态比较实现智能体协作中的贡献评分,提升奖励分配鲁棒性。
MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation

- 通过大模型生成智能体间对比,将信用分配转为排序聚合问题。
- 在动态参与和延迟反馈下仍能稳定生成有效贡献分。
- 适合需要精准协作评估的具身多智能体系统研究者。
信用分配是协作式多智能体强化学习中的核心挑战,尤其在具身AI场景中,由于反馈有限且延迟,以及活跃智能体数量动态变化。我们提出MARS-RA框架,将信用分配重构为基于贡献的多智能体成对比较的排序聚合问题,利用大模型生成的比较结果。该方法从绝对估计转向相对估计,增强对噪声和动态参与的鲁棒性,并将比较结果转化为潜在奖励塑造的贡献分数。我们提供了理论支持,证明该框架的收敛性与鲁棒性,表明沙普利值可作为解释参考。在多种复杂任务上的实验结果表明,MARS-RA能有效引导智能体实现协同。
原文摘要 · Abstract (English)
Credit assignment is a fundamental challenge in cooperative multi-agent reinforcement learning, particularly in embodied AI settings characterized by limited and delayed feedback as well as dynamically changing numbers of active agents. We propose MARS-RA, a framework that reformulates credit assignment as a rank aggregation problem using contribution-based pairwise comparisons among agents generated by large multimodal models. This shift from absolute to relative estimation ensures robustness against noise and dynamic agent participation, converting comparison results into contribution scores for potential-based reward shaping. We provide theoretical justification for the convergence and robustness of the proposed framework, and show that Shapley values can be used as an interpretive reference. Experimental results on challenging tasks of different types indicate that MARS-RA can guide agents toward effective cooperation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。