arXiv:2604.00722cs.CL2026-04被引 4

让大模型代理在协作中自动学会分工,靠语言空间的梯度优化。

LangMARL: Natural Language Multi-Agent Reinforcement Learning

  • 用语言描述每个代理的贡献,解决协作中的责任分配难题。
  • 在语言空间直接优化策略,提升样本效率与稀疏奖励下的收敛性。
  • 适合研究多智能体协作、大模型决策解释性的学者和工程师。

大型语言模型(LLM)代理在动态环境中难以自主演化协作策略,主要因粗粒度全局结果掩盖了局部策略优化所需的因果信号。我们识别出这一瓶颈为多智能体信用分配问题,该问题虽在经典多智能体强化学习(MARL)中长期受关注,但在基于LLM的系统中仍被忽视。基于此,我们提出LangMARL框架,将协作MARL中的信用分配与策略梯度演化引入语言空间。LangMARL引入代理级语言信用分配,开创在语言空间进行梯度演化以改进策略的方法,并从回放轨迹中总结任务相关的因果关系,提供密集反馈,在稀疏奖励下提升收敛性。在多种协作多智能体任务上的实验表明,该方法显著提升样本效率、可解释性与强泛化能力。

原文摘要 · Abstract (English)

Large language model (LLM) agents struggle to autonomously evolve coordination strategies in dynamic environments, largely because coarse global outcomes obscure the causal signals needed for local policy refinement. We identify this bottleneck as a multi-agent credit assignment problem, which has long been studied in classical multi-agent reinforcement learning (MARL) but remains underaddressed in LLM-based systems. Building on this observation, we propose LangMARL, a framework that brings credit assignment and policy gradient evolution from cooperative MARL into the language space. LangMARL introduces agent-level language credit assignment, pioneers gradient evolution in language space for policy improvement, and summarizes task-relevant causal relations from replayed trajectories to provide dense feedback and improve convergence under sparse rewards. Extensive experiments across diverse cooperative multi-agent tasks demonstrate improved sample efficiency, interpretability, and strong generalization.

多智能体语言模型强化学习协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。