arXiv:2510.18032cs.AIcs.MA2025-10被引 3

让多个大模型通过对话互相学习,提升复杂问题的推理能力。

OPTAGENT: Optimizing Multi-Agent LLM Interactions Through Verbal Reinforcement Learning for Enhanced Reasoning

  • 用语言强化学习动态构建合作结构,优化对话质量。
  • 在数学、科学等任务上表现优于现有方法,准确率显著提升。
  • 适合需要多智能体协作推理的场景,如复杂问题求解。

大型语言模型在数学和科学任务中展现出强大的推理能力。为增强复杂推理,多智能体系统被提出以利用多个语言模型智能体的集体智慧。然而,现有协作结构要么预先设定,要么依赖多数投票或圆桌辩论,可能压制少数但正确的观点。近期方法将多智能体系统建模为图网络,但仅优化单个智能体性能,忽视交互质量。我们假设有效沟通对多智能体推理至关重要,且辩论质量具有关键作用。为此,我们提出$ exttt{OPTAGENT}$,一种基于语言的强化学习算法,可动态构建并优化多智能体协作结构。该方法定义动作空间与反馈机制,评估辩论过程中的沟通鲁棒性与连贯性。最终决策通过所有智能体的多数投票实现。我们在数学推理、创意写作、科学推理和数值排序等多种任务上评估了$ exttt{OPTAGENT}$,结果表明其显著优于单智能体提示方法及当前最优的多智能体框架。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable reasoning capabilities in mathematical and scientific tasks. To enhance complex reasoning, multi-agent systems have been proposed to harness the collective intelligence of LLM agents. However, existing collaboration structures are either predefined or rely on majority voting or round-table debates, which can suppress correct but less dominant agent contributions. Recent approaches model multi-agent systems as graph networks but optimize purely for agent performance, neglecting the quality of interactions. We hypothesize that effective agent communication is crucial for multi-agent reasoning and that debating quality plays a significant role. To address this, we propose $\ours$, a multi-agent verbal reinforcement learning algorithm that dynamically constructs and refines multi-agent collaboration structures. Our method defines action spaces and a feedback mechanism that evaluates communication robustness and coherence throughout the debate. The final decision is achieved through a majority vote over all the agents. We assess $\ours$ on various reasoning tasks, including mathematical reasoning, creative writing, scientific reasoning, and numerical sorting. Results demonstrate that our approach significantly outperforms single-agent prompting methods and state-of-the-art multi-agent frameworks on diverse tasks.

多智能体推理增强强化学习语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。