arXiv:2512.24613cs.AI2025-12中稿 · IEEE ITCA 2025

多智能体协作推理模型提升复杂问题解答准确率与一致性

Group Deliberation Oriented Multi-Agent Conversational Model for Complex Reasoning

  • 三角色分工:生成、验证、整合,分步推进推理
  • 在HotpotQA等数据集上准确率提升14.3%~19.2%
  • 适合需要深度推理与可信结论的复杂任务场景

本文提出一种面向群体讨论的多智能体对话模型,以解决单一大语言模型在复杂推理任务中的局限性。模型采用生成、验证与集成三层角色分工架构:观点生成代理产生多样推理视角,证据验证代理检索外部知识并量化事实支持度,一致性仲裁代理整合逻辑一致的结论。引入自博弈机制扩展多路径推理轨迹,检索增强模块动态补充外部知识。设计融合事实一致性和逻辑连贯性的复合奖励函数,并采用改进的近端策略优化策略进行协同训练。实验结果显示,该模型在HotpotQA、2WikiMultihopQA和MeetingBank上的多跳推理准确率分别提升16.8%、14.3%和19.2%,一致性提升21.5%。相比主流多智能体方法,推理效率更高,为复杂推理任务提供了高效稳定的解决方案。

原文摘要 · Abstract (English)

This paper proposes a group deliberation oriented multi-agent conversational model to address the limitations of single large language models in complex reasoning tasks. The model adopts a three-level role division architecture consisting of generation, verification, and integration. An opinion generation agent produces diverse reasoning perspectives, an evidence verification agent retrieves external knowledge and quantifies factual support, and a consistency arbitration agent integrates logically coherent conclusions. A self-game mechanism is introduced to expand multi-path reasoning trajectories, while a retrieval enhancement module dynamically supplements external knowledge. A composite reward function combining factual consistency and logical coherence is designed, and an improved proximal policy optimization strategy is applied for collaborative training. Experimental results show that the proposed model improves multi-hop reasoning accuracy by 16.8 percent on HotpotQA, 14.3 percent on 2WikiMultihopQA, and 19.2 percent on MeetingBank, while improving consistency by 21.5 percent. The model achieves higher reasoning efficiency than mainstream multi-agent approaches, providing an effective and stable solution for complex reasoning tasks.

多智能体复杂推理对话系统逻辑一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。