arXiv:2601.22361cs.CLcs.AI2026-01被引 2

让AI查证更高效:用记忆模块复用证据,减少重复搜索。

MERMAID: Memory-Enhanced Retrieval and Reasoning with Multi-Agent Iterative Knowledge Grounding for Veracity Assessment

  • 多智能体协同查证,通过记忆库动态保存和复用证据。
  • 在3个事实核查数据集上超越现有方法,效率提升显著。
  • 适合需要高可靠性的自动化查证系统开发者使用。

在线内容真实性评估日益重要。大语言模型(LLMs)在自动事实核查与主张验证方面取得进展。传统流程将复杂主张拆分为子主张,检索外部证据并用LLM推理判断真伪。但现有方法常将证据检索视为静态独立步骤,难以有效管理和重用已获取证据。本文提出MERMAID框架,一种内存增强的多智能体迭代知识定位系统,将检索与推理紧密耦合。该框架整合智能体驱动搜索、结构化知识表示与持久记忆模块,在“推理-行动”式迭代过程中实现动态证据获取与跨主张证据复用。通过将检索到的证据存入证据记忆库,系统减少了重复搜索,提升了验证效率与一致性。我们在三个事实核查基准和两个主张验证数据集上,使用GPT、LLaMA、Qwen等多个大模型进行了评估。结果表明,MERMAID达到当前最优性能,并显著提升搜索效率,验证了检索、推理与记忆协同对可靠真实性评估的有效性。

原文摘要 · Abstract (English)

Assessing the veracity of online content has become increasingly critical. Large language models (LLMs) have recently enabled substantial progress in automated veracity assessment, including automated fact-checking and claim verification systems. Typical veracity assessment pipelines break down complex claims into sub-claims, retrieve external evidence, and then apply LLM reasoning to assess veracity. However, existing methods often treat evidence retrieval as a static, isolated step and do not effectively manage or reuse retrieved evidence across claims. In this work, we propose MERMAID, a memory-enhanced multi-agent veracity assessment framework that tightly couples the retrieval and reasoning processes. MERMAID integrates agent-driven search, structured knowledge representations, and a persistent memory module within a Reason-Action style iterative process, enabling dynamic evidence acquisition and cross-claim evidence reuse. By retaining retrieved evidence in an evidence memory, the framework reduces redundant searches and improves verification efficiency and consistency. We evaluate MERMAID on three fact-checking benchmarks and two claim-verification datasets using multiple LLMs, including GPT, LLaMA, and Qwen families. Experimental results show that MERMAID achieves state-of-the-art performance while improving the search efficiency, demonstrating the effectiveness of synergizing retrieval, reasoning, and memory for reliable veracity assessment.

事实核查多智能体记忆机制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。