arXiv:2607.20527cs.AI2026-07

发现科学生成系统引文可信度评估不可靠,提出可验证的防护机制。

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis

  • 设计可校准的黄金基准评估协议,确保验证器可靠性。
  • 实测引文不支持率在3%至18%间波动,取决于验证严格程度。
  • 提供无需依赖模型、可部署的防护层,保障生成结果可信性。

基于大模型的科学合成系统(如OpenScholar和PaperQA2)虽能返回带引用的答案,现有评估方式依赖固定归因模型或人工评分,却未检验该检查本身的可靠性。我们发现其结果极不稳定:相同输出下,引文不支持率从3%到18%不等,仅随验证严格度变化;验证者对需标记的引文分歧显著(负类一致性0.27~0.30),导致无单一可信标记集,跨论文比较无效。为此,我们提出黄金锚定评估协议与可部署防护机制:前者验证验证器性能、测量重归因并校准人类黄金标准下的保证值;后者采用可替换的验证器(召回率0.94,独立测试集),用确定性BM25实现重归因;再通过分拆置信区间层,在有限样本下对未被捕捉的错误引用施加分布无关的边界约束,提供捕获率保证而非结论正确性。该边界在保留数据上成立,并识别出迁移至部署的关键条件——校准-负例难度,给出具体再校准方案,弥补此前置信事实研究的空白。在四个公开的27-35B模型及三个代理流水线、三个公共基准(SciFact、QASA、PubMedQA)上验证,所有关键数值均附置信区间,协议与防护以单GPU开源工具包形式发布。

原文摘要 · Abstract (English)

Agentic LLM systems such as OpenScholar and PaperQA2 read the scientific literature and return cited answers, and both they and their benchmarks already check whether those citations hold, with a fixed attribution model or human graders. Neither audits the reliability of that check itself. We show it is not reliable, and that this matters. On identical agent outputs the measured unsupported-citation rate ranges from about 3% to about 18% depending only on the verifier's strictness, and although verifiers agree on which citations are supported, they disagree on which to flag (negative-specific agreement 0.27 to 0.30), so no single flag set is trustworthy and cross-paper comparison is invalid without a named verifier and protocol. We present a gold-anchored evaluation protocol and a deployable guard that make this behavior measurable and bounded. The protocol validates the verifier, measures re-attribution, and calibrates a guarantee against human gold rather than another model's verdict; the verifier is a swappable instrument chosen on cost (recall 0.94 on the supported class, held out), and re-attribution is a commodity step where a deterministic BM25 matches the best open generator. The guard adds a split-conformal layer placing a distribution-free, finite-sample bound on truly unsupported citations that slip past a chosen flagging rule, a guarantee on catch rate rather than conclusion correctness. The bound holds on held-out gold, and we identify and quantify the condition governing its transfer to deployment, calibration-negative difficulty, with a concrete recalibration recipe, left untested by prior conformal-factuality work. Validated across four open 27-35B models and three agentic pipelines on public benchmarks (SciFact, QASA, PubMedQA), with confidence intervals on every headline number, the protocol and guard ship as an open single-GPU kit.

引文可信大模型评估防护机制科学生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。