arXiv:2603.24324cs.LGcs.AI2026-03

用大模型自动生成协作智能体奖励,提升复杂任务表现

Large Language Model Guided Incentive Aware Reward Design for Cooperative Multi-Agent Reinforcement Learning

  • 用大模型从环境信息生成可执行奖励程序
  • 在四种厨房场景中任务完成率显著提升,瓶颈区域改善最明显
  • 无需人工调参,适合资源有限的协作强化学习场景

为协作多智能体系统设计有效辅助奖励仍具挑战性,因激励错位可能导致次优协调,尤其在稀疏任务奖励难以支撑协同行为时。本文提出一种自主奖励设计框架,利用大语言模型(LLMs)从环境仪器化信息中合成可执行奖励程序。该方法在形式有效性范围内约束候选程序,并在固定计算预算下使用多智能体近端策略优化(MAPPO)从头训练策略。候选程序仅依据稀疏任务回报进行评估与代际选择。框架在四个具有不同通道拥堵、交接依赖和结构不对称性的Overcooked-AI布局中进行评估。所提奖励设计方法在所有场景中均获得更高任务回报与配送次数,尤其在交互瓶颈主导的环境中增益最为显著。诊断分析显示,合成的塑造成分增强了动作选择的相互依赖性,并提升了密集协作任务中的信号对齐度。结果表明,该基于大模型引导的奖励搜索框架减少了人工工程需求,且在有限预算下生成了兼容协作学习的塑造信号。

原文摘要 · Abstract (English)

Designing effective auxiliary rewards for cooperative multi-agent systems remains challenging, as misaligned incentives can induce suboptimal coordination, particularly when sparse task rewards provide insufficient grounding for coordinated behavior. This study introduces an autonomous reward design framework that uses large language models (LLMs) to synthesize executable reward programs from environment instrumentation. The procedure constrains candidate programs within a formal validity envelope and trains policies from scratch using Multi-Agent Proximal Policy Optimization (MAPPO) under a fixed computational budget. The candidates are then evaluated on the basis of their performance, and selection across generations solely based on the sparse task returns. The framework is evaluated in four Overcooked-AI layouts characterized by varying levels of corridor congestion, handoff dependencies, and structural asymmetries. The proposed reward design approach consistently yields higher task returns and delivery counts, with the most pronounced gains observed in environments dominated by interaction bottlenecks. Diagnostic analysis of the synthesized shaping components reveals stronger interdependence in action selection and improved signal alignment in coordination-intensive tasks. These results demonstrate that the proposed LLM-guided reward search framework mitigates the need for manual engineering while producing shaping signals compatible with cooperative learning under finite budgets.

多智能体大模型强化学习奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。