arXiv:2601.08237cs.AI2026-01被引 3

用语言描述替代人工设计奖励函数,让大模型实现智能体协作

The End of Reward Engineering: How LLMs Are Redefining Multi-Agent Coordination

  • 用自然语言直接生成奖励机制,避免手动调参
  • 大模型可动态调整奖励策略,减少人工干预
  • 适合需要快速适配人类意图的多智能体系统

奖励工程——通过手工设定奖励函数引导智能体行为——仍是多智能体强化学习中的核心难题,尤其在信用分配模糊、环境非平稳及交互复杂性呈组合增长的背景下。我们提出,大型语言模型(LLMs)的发展正推动从人工数值奖励转向语言化目标定义。已有研究显示,LLMs 可从自然语言描述中直接合成奖励函数(如 EUREKA),并在线动态调整奖励形式,仅需极少人工干预(如 CARD)。同时,基于可验证奖励的强化学习(RLVR)实证表明,语言驱动的监督可作为传统奖励工程的有效替代方案。本文从语义奖励定义、动态奖励适应和人类意图对齐三个维度阐述这一转变,指出计算开销、幻觉鲁棒性及大规模系统扩展性等开放挑战。最终,我们展望一种以共享语义表征为基础而非显式数值信号的协作新范式。

原文摘要 · Abstract (English)

Reward engineering, the manual specification of reward functions to induce desired agent behavior, remains a fundamental challenge in multi-agent reinforcement learning. This difficulty is amplified by credit assignment ambiguity, environmental non-stationarity, and the combinatorial growth of interaction complexity. We argue that recent advances in large language models (LLMs) point toward a shift from hand-crafted numerical rewards to language-based objective specifications. Prior work has shown that LLMs can synthesize reward functions directly from natural language descriptions (e.g., EUREKA) and adapt reward formulations online with minimal human intervention (e.g., CARD). In parallel, the emerging paradigm of Reinforcement Learning from Verifiable Rewards (RLVR) provides empirical evidence that language-mediated supervision can serve as a viable alternative to traditional reward engineering. We conceptualize this transition along three dimensions: semantic reward specification, dynamic reward adaptation, and improved alignment with human intent, while noting open challenges related to computational overhead, robustness to hallucination, and scalability to large multi-agent systems. We conclude by outlining a research direction in which coordination arises from shared semantic representations rather than explicitly engineered numerical signals.

多智能体大模型语言奖励强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。