arXiv:2605.26663cs.CLcs.IR2026-05

发现验证模型常误判缺失证据,提出诊断方法提升评测可靠性

Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification

论文配图:Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification
图 1 · 摘自论文原文
  • 设计构造感知诊断协议,识别不同证据构造对模型的影响
  • 实测多种构造下模型表现不一致,跨构造泛化能力差
  • 建议报告评分时附带构造类别,避免误导性结论

证据缺失不等于证据不足,但事实验证基准常使二者表现相似。当前的‘证据不足’(NEI)标签多通过人工构造证据条件定义,而这一选择隐性决定了模型的学习目标。本文提出NEI-CAP——一种构造感知的诊断协议,每个NEI样本均标注其构造来源;该协议可审计捷径线索、通过人工裁定验证难点样本,并测试模型在不同构造间的泛化能力。我们在SciFact上实现该协议,并以FEVER和HoVer作为外部对照。结果表明,不同构造下的模型表现不可靠:编码器型验证器及在捷径构造上训练的指令微调解码器无法识别语义相关的证据不足情况;混合构造训练虽缩小差距但未能消除。固定命题诊断显示,证据构造会影响对支持/反驳标签的信心,不仅影响NEI召回率。因此,建议在报告分数时附带构造家族信息,并提供一份含检查项的基准改进清单。

原文摘要 · Abstract (English)

Evidence absence is not evidence insufficiency, but fact verification benchmarks can make them observationally similar. The Not Enough Information (NEI) label is often operationalized through constructed evidence conditions, and that choice silently determines what a verifier learns. We introduce NEI-CAP, a construction-aware diagnostic protocol for insufficient-evidence evaluation. Each NEI example carries the construction family that produced it; NEI-CAP audits shortcut cues, validates hard cases through human adjudication, and tests whether competence transfers across constructions. We instantiate the protocol on SciFact, with FEVER and HoVer as bounded external controls. Across these settings, NEI competence does not transfer reliably: encoder verifiers and an instruction-tuned decoder trained on shortcut-prone constructions fail to recognize semantically related insufficient evidence, and mixed-construction training narrows but does not close the gap. Fixed-claim diagnostics further show that the evidence condition shifts confidence in the reference Support/Refute label, not only NEI recall, so an aggregate NEI score can hide which problem a model has actually solved. We therefore recommend reporting the construction family alongside the score, and distill the results into a checklist for benchmarks that carry an insufficient-evidence label.

事实验证模型诊断数据构造NEI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。