arXiv:2608.00817cs.AIcs.HC2026-08

带来源引用的AI助诊能提升医生准确率,但可能让人盲目信任有依据的错误建议。

Large language models improve physician accuracy but lead to false reliance

  • 设计智能检索增强型AI系统CORA,结合临床证据支持决策
  • 医生准确率从70.8%升至82.6%,对新发病例提升更明显
  • 有引用的错误建议遭抵制率从92%降至34.8%,存在认知盲区

检索增强的大语言模型(LLMs)有望提供可溯源的临床支持,但其价值取决于展示的证据是否引导而非扭曲医生判断。我们开发了CORA——一种具备代理能力的检索增强型大语言模型,研究源链接辅助对医生决策的影响。CORA保持了基准性能,并在训练数据截止后发布的病例中取得更大提升。在46名医生的研究中,准确率从无辅助时的70.8%提高到使用CORA后的82.6%。支持性引用能预测正确答案(87.7%对比65.5%),但产生显著不对称:当正确建议带有引用时,采纳率从34%升至76.9%;而当错误建议带有引用时,医生拒绝率从92%降至34.8%。结果表明,源链接式大模型可提升医生准确率,但引入了依赖证据可信度的安全风险。

原文摘要 · Abstract (English)

Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.

大模型医疗临床决策可信推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。