arXiv:2604.10990cs.CLcs.AI2026-04

现有验证模型常靠关键元素通过,而非全面验证,导致错误接受复杂不成立的科学主张。

When Verification Fails: How Compositionally Infeasible Claims Escape Rejection

论文配图:When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
图 1 · 摘自论文原文
  • 构造复合不成立的主张,使关键约束支持但非关键约束矛盾
  • 多数模型在新测试中过度接受此类主张,暴露验证漏洞
  • 不同模型差异源于阈值设定,非推理能力不足,策略无法根本解决

科学主张验证旨在判断主张是否由证据支持,是防止虚假信息、确立科学发现的基础。在闭世界假设(CWA)下,仅当所有主张约束均被正向支持时,才接受该主张。我们发现,现有验证基准无法区分严格遵循此标准的模型与采用简化捷径‘显著约束检查’的模型——后者仅对最显著的约束应用拒绝标准,若其被支持则接受主张。由于现有基准通过扰动单一显著元素构造不可行主张,因此不足以区分严谨验证与简单依赖显著约束的行为。为此,我们构建了复合不可行主张:显著约束被支持,但非显著约束被反驳。在多个模型家族和模态上,原本在现有基准表现饱和的模型均持续过度接受这些主张,证实此类捷径普遍存在。通过上下文干预实验,我们发现不同模型与提示策略在共享的ROC曲线上占据不同位置,表明模型间的差距源于验证阈值差异,而非推理能力;组合推理瓶颈是当前验证行为的结构性缺陷,仅靠策略引导无法克服。

原文摘要 · Abstract (English)

Scientific claim verification, the task of determining whether claims are entailed by scientific evidence, is fundamental to establishing discoveries in evidence while preventing misinformation. This process involves evaluating each asserted constraint against validated evidence. Under the Closed-World Assumption (CWA), a claim is accepted if and only if all asserted constraints are positively supported. We show that existing verification benchmarks cannot distinguish models enforcing this standard from models applying a simpler shortcut called salient-constraint checking, which applies CWA's rejection criterion only to the most salient constraint and accepts when that constraint is supported. Because existing benchmarks construct infeasible claims by perturbing a single salient element they are insufficient at distinguishing between rigorous claim verification and simple salient-constraint reliance. To separate the two, we construct compositionally infeasible claims where the salient constraint is supported but a non-salient constraint is contradicted. Across model families and modalities, models that otherwise saturate existing benchmarks consistently over-accept these claims, confirming the prevalence of such shortcut reasoning. Via model context interventions, we show that different models and prompting strategies occupy distinct positions on a shared ROC curve, indicating that the gap between model families reflects differences in verification threshold rather than underlying reasoning ability, and that the compositional inference bottleneck is a structural property of current verification behavior that strategy guidance alone cannot overcome.

科学验证模型评估推理偏差闭环假设

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。