arXiv:2601.04742cs.CL2026-01被引 5

让多个智能体用不同工具辩论,更准判断事实真伪。

Tool-MAD: A Multi-Agent Debate Framework for Fact Verification with Diverse Tool Augmentation and Adaptive Retrieval

  • 每个智能体配备不同外部工具,激发多样推理视角。
  • 辩论中动态优化检索,实时获取新证据,提升适应性。
  • 结合可信度与相关性评分,有效识别幻觉内容。

大型语言模型在复杂推理和事实验证任务中常出现幻觉和错误。多智能体辩论(MAD)系统通过让多个模型对话来提升答案准确性,但现有框架多依赖内部知识或静态文档,仍易产生幻觉。尽管MADKE引入外部证据缓解此问题,其一次性检索机制难以应对辩论中的新论点或信息变化。为此,我们提出Tool-MAD:一种基于异构外部工具的多智能体辩论框架。该框架包含三项创新:(1)为各智能体分配不同外部工具(如搜索接口或RAG模块),促进多角度推理;(2)设计自适应查询生成机制,根据辩论进程迭代优化证据检索;(3)将忠实度与回答相关性分数纳入决策过程,使裁判智能体可量化评估回应的逻辑一致性与问题契合度,有效检测幻觉。在四个事实验证基准上的实验表明,Tool-MAD持续优于当前最先进MAD框架,最高准确率提升5.5%。在医学专业领域,该框架对不同工具配置和领域条件均表现出强鲁棒性与适应性,具备广泛现实应用潜力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) suffer from hallucinations and factual inaccuracies, especially in complex reasoning and fact verification tasks. Multi-Agent Debate (MAD) systems aim to improve answer accuracy by enabling multiple LLM agents to engage in dialogue, promoting diverse reasoning and mutual verification. However, existing MAD frameworks primarily rely on internal knowledge or static documents, making them vulnerable to hallucinations. While MADKE introduces external evidence to mitigate this, its one-time retrieval mechanism limits adaptability to new arguments or emerging information during the debate. To address these limitations, We propose Tool-MAD, a multi-agent debate framework that enhances factual verification by assigning each agent a distinct external tool, such as a search API or RAG module. Tool-MAD introduces three key innovations: (1) a multi-agent debate framework where agents leverage heterogeneous external tools, encouraging diverse perspectives, (2) an adaptive query formulation mechanism that iteratively refines evidence retrieval based on the flow of the debate, and (3) the integration of Faithfulness and Answer Relevance scores into the final decision process, allowing the Judge agent to quantitatively assess the coherence and question alignment of each response and effectively detect hallucinations. Experimental results on four fact verification benchmarks demonstrate that Tool-MAD consistently outperforms state-of-the-art MAD frameworks, achieving up to 5.5% accuracy improvement. Furthermore, in medically specialized domains, Tool-MAD exhibits strong robustness and adaptability across various tool configurations and domain conditions, confirming its potential for broader real-world fact-checking applications.

多智能体事实验证幻觉检测工具增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。