用大模型分阶段验证科学主张与引用是否匹配,更准更快。
DeepSciVerify: Verifying Scientific Claim--Citation Alignment via LLM-Driven Evidence Escalation

- 先看摘要判断主张,不确定时才查全文
- 在SCitance上准确率达86.7,比纯摘要方法高4.5点
- 67%的情况无需查全文,适合需要高效可信的科研写作
大语言模型生成报告时,主张与引用证据常不匹配,影响其在科研等高风险场景的可靠性。我们提出DeepSciVerify,一种两阶段科学主张-引用验证方法:先基于摘要进行推理,对不确定的案例选择性地升级到段落级全文证据分析。该设计利用不同大模型在不确定性下的互补行为——部分模型更保守,部分更果断。在SCitance基准测试中,DeepSciVerify取得86.7的Micro-F1,较强的仅用摘要基线提升4.5点,且67%的实例无需全文检索即可解决。结果表明,选择性证据升级可同时提升验证准确率与效率。
原文摘要 · Abstract (English)
Misalignment between claims and their cited evidence is a common failure mode in reports generated by large language models, limiting their reliability in scientific and other high-stakes settings. We present DeepSciVerify, a two-stage pipeline for scientific claim-citation verification that combines abstract-level reasoning with selective escalation to passage-level evidence. The system first verifies claims using the abstract and defers uncertain cases, retrieving and analyzing full-text passages only when necessary. This design leverages complementary behaviors across LLMs, as some models are more conservative while others are more decisive under uncertainty. On the SCitance benchmark, DeepSciVerify achieves 86.7 Micro-F1, outperforming strong abstract-only baselines by +4.5 points while resolving 67% of instances without full-text retrieval. These results suggest that selective evidence escalation improves both accuracy and efficiency in claim-citation verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。