arXiv:2608.01409cs.CL2026-08

评估大模型生成生物医学证据的可靠性,发现检索既助力又干扰。

When Retrieval Helps and Distracts: Evaluating Evidence-Generating LLMs for Biomedical Claim Verification

  • 对比多种模型在生成可信证据上的表现,重点分析检索增强的效果。
  • 微调后的LLM在生成证据上最强,但检索对不同数据源效果差异大。
  • 提出新评测指标,揭示检索是否真正提升判断力,适合验证研究者使用。

生物医学事实核查系统不仅需预测主张是否支持、矛盾或未涉及,还应生成忠实、完整且可用于验证的证据。我们在CARE-XAI这一涵盖五个生物医学与健康事实核查来源的统一基准上研究了证据生成任务。比较基础指令型LLM、PubMed检索增强的LLM、微调后的LLM、仅标签输入的LLM以及生物医学编码器分类器在统一评估协议下的表现。生物医学分类器在仅判别结果上仍最强,而微调后的LLM在生成证据方面最优。PubMed检索效果混合:在如PubMedQA和SciFact等与PubMed对齐的数据源上有所帮助,但在更广泛的公共卫生主张上可能产生干扰。我们引入Bio-GRACE,一种基于黄金参考的标准诊断工具,用于衡量检索证据能否恢复参考证据的决策优势。结果显示,检索效用具有数据源依赖性,支持选择性检索,并揭示检索召回率与词法证据重叠不足以衡量生物医学事实核查的有效性。

原文摘要 · Abstract (English)

Biomedical fact-checking systems must do more than predict whether a claim is supported, contradicted, or unaddressed: they should also produce evidence that is faithful, complete, and useful for verification. We study this evidence-generation setting on CARE-XAI, a unified benchmark spanning five biomedical and health fact-checking sources. We compare base instruction LLMs, PubMed retrieval-augmented LLMs, fine-tuned LLMs, label-only LLMs, and biomedical encoder classifiers under a shared evaluation protocol. Biomedical classifiers remain strongest for verdict-only prediction, while fine-tuned LLMs are the strongest evidence-generating systems. PubMed retrieval is mixed: it helps PubMed-aligned sources such as PubMedQA and SciFact, but can distract models on broader public-health claims. We introduce Bio-GRACE, a gold-reference-normalized diagnostic for measuring whether retrieved evidence recovers the decision benefit of reference evidence. Bio-GRACE shows that retrieval utility is source-dependent, motivates selective retrieval, and exposes why retrieval recall and lexical evidence overlap are insufficient for biomedical fact-checking.

大模型生物医学证据生成事实核查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。