arXiv:2605.14115cs.CL2026-05ACL

医学问答中证据冲突时,模型易出错,需改进判断可靠性。

When Evidence Conflicts: Uncertainty and Order Effects in Retrieval-Augmented Biomedical Question Answering

论文配图:When Evidence Conflicts: Uncertainty and Order Effects in Retrieval-Augmented Biomedical Question Answering
图 1 · 摘自论文原文
  • 测试不同证据顺序下模型表现,发现相同内容顺序不同结果差11.4%~25.2%
  • 在矛盾证据条件下,模型准确率下降,超四分之一预测结果翻转
  • 提出冲突感知弃答机制,提升困难场景下选择性准确率7.2~33.4点

生物医学检索增强型大模型常面临不完整、误导或内部矛盾的证据,但现有评估多聚焦于有益上下文下的答案准确性,而非冲突情境下的可靠性。使用HealthContradict数据集,我们在五种受控证据条件下评估六种开源大模型:无检索、仅正确、仅错误,以及两种混合条件(含正确与矛盾文档,顺序相反)。在证据冲突的顺序对比中,同一组文档仅顺序反转,所有模型准确率均下降,11.4%–25.2%的预测结果发生改变。为支持复杂情形下的拒答,我们评估了一种结合模型置信度与证据冲突检测器的冲突感知弃答分数。在最困难的两种条件下,该分数相比仅依赖置信度的方法显著提升选择性准确率,分别在75%、50%、25%覆盖率下取得7.2–33.4分和3.6–14.4分的平均增益。结果表明,矛盾证据既是不确定性问题,也是鲁棒性挑战,亟需在评估与拒答机制中显式考虑证据分歧。

原文摘要 · Abstract (English)

Biomedical retrieval-augmented large language models (LLMs) often face evidence that is incomplete, misleading, or internally contradictory, yet evaluation usually emphasizes answer accuracy under helpful context rather than reliability under conflict. Using HealthContradict, we evaluate six open-weight LLMs under five controlled evidence conditions: no retrieved context, correct-only context, incorrect-only context, and two mixed conditions containing both correct and contradictory documents in opposite orders. In this conflicting-evidence order contrast, where the same two documents are both present and only their order is reversed, accuracy drops for every model and 11.4%--25.2% of predictions flip. To support abstention in these difficult cases, we also evaluate a conflict-aware abstention score that combines model confidence with a detector of evidence conflict. In the two hardest conditions, this score improves selective accuracy over confidence-only, with mean gains of 7.2--33.4 points in incorrect-only (`IC') and 3.6--14.4 points in incorrect-first conflicting (`ICC') conditions across 75%, 50%, and 25% coverage. These results show that conflicting biomedical evidence is both an uncertainty and robustness problem and motivate evaluation and abstention methods that explicitly account for evidence disagreement.

医学问答证据冲突大模型可靠性弃答机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。