动态评测仍存信息污染,17%~29%的新证据也可能被提前知晓。
Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

- 检验动态评测中新发布声明是否真正依赖外部证据
- 17.09%至29.30%的后截断声明仍可被模型提前掌握
- 污染导致性能虚高最高达11.34点,影响评估可信度
多模态自动事实核查(MAFC)通过检索和推理外部证据来验证声明。然而,现有静态基准大多包含过时声明,这些声明可仅靠大语言模型(LLM)内部知识验证,导致性能估计虚高且无法反映真实能力。为应对这一问题,新兴动态基准采用在模型知识截止日期之后发布的声明,假设其无污染。本文通过实证研究最先进的静态基准AVerTeC和我们新构建的动态基准ClaimReview2025Q4,揭示了污染风险的普遍存在。实验得出16项发现,关键结果包括:(1) 动态评测虽降低但未消除污染风险,17.09%–29.30%的后截断声明仍可能被污染;(2) 许多新发布声明可通过截止日期前已公开的多条公共知识直接或合成验证;(3) 污染可造成显著性能虚高,使宏平均F1提升高达11.34点,并扭曲系统排名。基于此,我们在严格控制污染条件下重新评估了SOTA LLMs,提出可信评估实践指南。
原文摘要 · Abstract (English)
Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09\%--29.30\% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。