改进大模型事实核查,让证据真正支持结论
The Warrant Gap: Claim-Conditioned Re-scoring for Fact-Checking
- 根据完整论点重新评分提取的证据片段,提升推理合理性
- 在多个数据集上将错误支持率降低27.6个百分点
- 适合需要高可信度证据链的自动审核系统
基于大语言模型的事实核查系统在标准评测中准确率高,但常输出无证据支持的“支持”判断。现有结构化分解方法因僵化提取规则,丢失了完整论点上下文。本文提出SIFT——基于全论点条件对提取证据进行重评分,并引入WSP(有据支持比例)指标,通过自动自然语言推理检查证据是否蕴含论点。在FEVER、SciFact、5PILS和DP四个数据集上,使用四种开源模型评估。SIFT使原本因简单分解损失高达27.6分的准确率得以恢复,且显著优于直接提示;WSP指标在人类黄金证据上达到AUC 0.92、精度0.98,具备良好校准性。
原文摘要 · Abstract (English)
Fact-checking systems built on LLMs achieve high verdict accuracy on standard benchmarks, yet routinely output Supports labels whose cited evidence does not license the claim. Structured decomposition is the natural way to inspect those warrants, but rigid extraction protocols strip the full-claim context that facets need. We introduce SIFT -- claim-conditioned re-scoring of extracted evidence spans against the full claim -- paired with WSP (Warranted Supports Proportion), an automatic NLI check that the cited warrant entails the claim. We evaluate on FEVER, SciFact, 5PILS, and DP across four open-source backbones. SIFT recovers accuracy on cells where naive decomposition costs up to 27.6 points, while raising WSP above direct prompting; WSP itself calibrates against human gold evidence at AUC 0.92 and precision 0.98.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。