arXiv:2602.10380cs.CLcs.AI2026-02

复杂声明分解效果差,根源在于证据对齐和子声明错误类型。

The Alignment Bottleneck in Decomposition-Based Claim Verification

  • 提出新数据集,包含时间限定证据和人工标注的子声明支持片段。
  • 只有细粒度且严格对齐的证据才能提升验证效果,否则性能反而下降。
  • 保守不判断比错误判断更抗噪,适合应对标签噪声场景。

结构化声明分解常被提议用于验证复杂多面声明,但实证结果不一致。我们指出这种不一致性源于两个被忽视的瓶颈:证据对齐问题与子声明错误特征。为此,我们构建了一个真实世界复杂声明的新数据集,包含时间限定证据和人工标注的子声明证据范围。在两种证据对齐设置下评估分解方法:子声明对齐证据(SAE)和重复的声明级证据(SRE)。结果表明,仅当证据细粒度且严格对齐时,分解才能带来显著性能提升;而依赖重复声明级证据的标准设置(SRE)不仅未能改进,反而在多个数据集和领域(PHEMEPlus、MMM-Fact、COVID-Fact)中导致性能下降。此外,在子声明标签存在噪声的情况下,错误性质决定了下游鲁棒性:保守的‘弃权’策略显著减少错误传播,优于激进但错误的预测。研究建议未来框架应优先关注精确证据整合,并校准子声明验证模型的标签偏倚。

原文摘要 · Abstract (English)

Structured claim decomposition is often proposed as a solution for verifying complex, multi-faceted claims, yet empirical results have been inconsistent. We argue that these inconsistencies stem from two overlooked bottlenecks: evidence alignment and sub-claim error profiles. To better understand these factors, we introduce a new dataset of real-world complex claims, featuring temporally bounded evidence and human-annotated sub-claim evidence spans. We evaluate decomposition under two evidence alignment setups: Sub-claim Aligned Evidence (SAE) and Repeated Claim-level Evidence (SRE). Our results reveal that decomposition brings significant performance improvement only when evidence is granular and strictly aligned. By contrast, standard setups that rely on repeated claim-level evidence (SRE) fail to improve and often degrade performance as shown across different datasets and domains (PHEMEPlus, MMM-Fact, COVID-Fact). Furthermore, we demonstrate that in the presence of noisy sub-claim labels, the nature of the error ends up determining downstream robustness. We find that conservative "abstention" significantly reduces error propagation compared to aggressive but incorrect predictions. These findings suggest that future claim decomposition frameworks must prioritize precise evidence synthesis and calibrate the label bias of sub-claim verification models.

声明验证证据对齐子声明鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。