arXiv:2505.16576cs.CL2025-05被引 4

用多智能体模拟人类查证过程,提升事实核查准确性。

EMULATE: A Multi-Agent Framework for Determining the Veracity of Atomic Claims by Emulating Human Actions

  • 构建多智能体系统,分工完成检索排序与内容评估
  • 在多个基准上优于现有方法,验证框架有效性
  • 适合对可解释性与人类行为一致性要求高的研究者

确定原子性声明的真实性是当前许多事实核查系统的核心环节。现有方法通常通过搜索引擎检索证据,再将证据集与声明输入大语言模型进行分类,但这一流程偏离了人类实际查证方式。近期工作尝试通过迭代式证据检索改进,仅在必要时收集证据。本文提出新系统EMULATE,采用多智能体框架,每个智能体负责任务中的一小部分,如根据预设标准对搜索结果排序或评估网页内容。在多个基准上的大量实验表明,该方法显著优于先前工作,证明了多智能体框架的有效性。

原文摘要 · Abstract (English)

Determining the veracity of atomic claims is an imperative component of many recently proposed fact-checking systems. Many approaches tackle this problem by first retrieving evidence by querying a search engine and then performing classification by providing the evidence set and atomic claim to a large language model, but this process deviates from what a human would do in order to perform the task. Recent work attempted to address this issue by proposing iterative evidence retrieval, allowing for evidence to be collected several times and only when necessary. Continuing along this line of research, we propose a novel claim verification system, called EMULATE, which is designed to better emulate human actions through the use of a multi-agent framework where each agent performs a small part of the larger task, such as ranking search results according to predefined criteria or evaluating webpage content. Extensive experiments on several benchmarks show clear improvements over prior work, demonstrating the efficacy of our new multi-agent framework.

事实核查多智能体人类行为模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。