arXiv:2604.01657cs.CL2026-04

分析9个事实核查数据集,发现多数测试依赖直接证据提取。

What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis

  • 用大模型生成推理轨迹,量化分析验证过程
  • 仅半数数据集需信息整合,数值推理极少出现
  • 不同领域错误类型差异大,科学类易过度谨慎

尽管事实核查取得快速进展,我们仍缺乏对这些基准测试实际考察推理能力的系统理解。通过使用GPT-4o-mini为9个数据集中的2.4万条声明-验证样本生成结构化推理轨迹,我们发现直接证据提取占主导地位,而多句合成与数值推理严重不足。数据集层面分析显示显著偏差:某些数据集几乎仅测试词汇匹配,另一些则在约一半案例中要求信息合成。利用一个10亿参数的推理验证器,我们进一步识别出五种错误类型,并发现错误模式随领域差异显著——通用领域以词汇重叠偏差为主,科学领域表现为过度谨慎,数学领域则主要因算术推理失败。研究结果表明,高基准分数主要反映的是检索加蕴含的能力。本文提出建议,以构建更具有挑战性的评估套件,真正测试验证系统所需的推理能力。

原文摘要 · Abstract (English)

Despite rapid progress in claim verification, we lack a systematic understanding of what reasoning these benchmarks actually exercise. We generate structured reasoning traces for 24K claim-verification examples across 9 datasets using GPT-4o-mini and find that direct evidence extraction dominates, while multi-sentence synthesis and numerical reasoning are severely under-represented. A dataset-level breakdown reveals stark biases: some datasets almost exclusively test lexical matching, while others require information synthesis in roughly half of cases. Using a compact 1B-parameter reasoning verifier, we further characterize five error types and show that error profiles vary dramatically by domain -- general-domain verification is dominated by lexical overlap bias, scientific verification by overcautiousness, and mathematical verification by arithmetic reasoning failures. Our findings suggest that high benchmark scores primarily reflect retrieval-plus-entailment ability. We outline recommendations for building more challenging evaluation suites that better test the reasoning capabilities verification systems need.

事实核查推理分析数据集评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。