arXiv:2511.02303cs.AIcs.CL2025-11被引 12

解决多智能体大模型推理中的偷懒问题,提升协作效率

Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation

  • 通过因果影响度量识别并抑制代理间依赖失衡
  • 引入可验证奖励机制,让推理代理能重置错误路径
  • 适合需要深度逻辑推理的复杂任务研究者

基于强化学习和可验证奖励训练的大语言模型在复杂推理任务上表现优异。近期工作将其扩展至多智能体场景:元思考代理提出计划并监控进展,推理代理通过多轮对话执行子任务。然而我们发现一个关键缺陷:懒惰行为——一个代理主导,另一个贡献极少,导致协作失效,系统退化为单代理。本文首先理论分析了懒惰行为在多代理推理中自然产生的原因;随后提出一种稳定高效的因果影响测量方法以缓解该问题;最后,针对协作加深后推理代理易陷入多轮对话、被早期噪声响应困住的问题,提出可验证奖励机制,允许推理代理丢弃噪声输出、整合指令并必要时重启动推理过程。大量实验表明,该框架有效缓解了懒惰行为,充分释放了多代理框架在复杂推理任务中的潜力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) trained with reinforcement learning and verifiable rewards have achieved strong results on complex reasoning tasks. Recent work extends this paradigm to a multi-agent setting, where a meta-thinking agent proposes plans and monitors progress while a reasoning agent executes subtasks through sequential conversational turns. Despite promising performance, we identify a critical limitation: lazy agent behavior, in which one agent dominates while the other contributes little, undermining collaboration and collapsing the setup to an ineffective single agent. In this paper, we first provide a theoretical analysis showing why lazy behavior naturally arises in multi-agent reasoning. We then introduce a stable and efficient method for measuring causal influence, helping mitigate this issue. Finally, as collaboration intensifies, the reasoning agent risks getting lost in multi-turn interactions and trapped by previous noisy responses. To counter this, we propose a verifiable reward mechanism that encourages deliberation by allowing the reasoning agent to discard noisy outputs, consolidate instructions, and restart its reasoning process when necessary. Extensive experiments demonstrate that our framework alleviates lazy agent behavior and unlocks the full potential of multi-agent framework for complex reasoning tasks.

多智能体推理增强强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。