构建全流程事实核查评估框架,揭示大模型在真实场景中的推理短板。
Towards Comprehensive Stage-wise Benchmarking of Large Language Models in Fact-Checking
- 用自动化流程模拟从提取到验证的完整事实核查链路。
- 16个主流大模型在端到端任务中表现显著低于单一验证准确率。
- 支持动态生成挑战性问题,适合模型安全评估与优化研究者使用。
大语言模型(LLMs)正被广泛部署于实际事实核查系统,但现有评估主要聚焦于声明验证,忽略了包括声明提取和证据检索在内的完整核查流程。这种局限性导致无法揭示现代大模型在系统性推理、事实盲区和鲁棒性方面的缺陷。为填补这一空白,我们提出 FactArena,一个全自动的竞技场式评估框架,对大模型在全链条事实核查流程中进行综合、分阶段测评。FactArena 包含三个核心组件:(i) 基于 LLM 的标准化核查流程,涵盖声明分解、工具增强的证据检索及基于理由的结论预测;(ii) 基于统一参考指南的竞技场式判断机制,确保异构裁判代理间的无偏一致比较;(iii) 基于竞技场的声明演化模块,可自适应生成更具挑战性且语义可控的声明,以探测大模型在固定种子数据之外的事实鲁棒性。在覆盖七种模型家族的16个前沿大模型上,FactArena 生成了稳定可解释的排名。分析显示,静态声明验证准确率与端到端核查能力存在显著差异,凸显全面评估的必要性。该框架为诊断大模型事实推理能力、指导未来模型研发、推动安全关键场景下大模型的可靠部署提供了可扩展且可信的范式。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in real-world fact-checking systems, yet existing evaluations focus predominantly on claim verification and overlook the broader fact-checking workflow, including claim extraction and evidence retrieval. This narrow focus prevents current benchmarks from revealing systematic reasoning failures, factual blind spots, and robustness limitations of modern LLMs. To bridge this gap, we present FactArena, a fully automated arena-style evaluation framework that conducts comprehensive, stage-wise benchmarking of LLMs across the complete fact-checking pipeline. FactArena integrates three key components: (i) an LLM-driven fact-checking process that standardizes claim decomposition, evidence retrieval via tool-augmented interactions, and justification-based verdict prediction; (ii) an arena-styled judgment mechanism guided by consolidated reference guidelines to ensure unbiased and consistent pairwise comparisons across heterogeneous judge agents; and (iii) an arena-driven claim-evolution module that adaptively generates more challenging and semantically controlled claims to probe LLMs' factual robustness beyond fixed seed data. Across 16 state-of-the-art LLMs spanning seven model families, FactArena produces stable and interpretable rankings. Our analyses further reveal significant discrepancies between static claim-verification accuracy and end-to-end fact-checking competence, highlighting the necessity of holistic evaluation. The proposed framework offers a scalable and trustworthy paradigm for diagnosing LLMs' factual reasoning, guiding future model development, and advancing the reliable deployment of LLMs in safety-critical fact-checking applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。