arXiv:2605.22047cs.AI2026-05

测试大模型在临床问诊中的动态取证能力,发现互动提问会降低诊断准确率。

Active Evidence-Seeking and Diagnostic Reasoning in Large Language Models for Clinical Decision Support

论文配图:Active Evidence-Seeking and Diagnostic Reasoning in Large Language Models for Clinical Decision Support
图 1 · 摘自论文原文
  • 设计模拟患者系统,评估模型多轮问诊时的证据获取能力。
  • 多轮互动使诊断准确率下降12.75%,证据质量下降24.36%。
  • 适合关注医疗AI安全性和交互式诊断评估的研究者。

大语言模型在静态医疗病例上表现良好,但临床诊断常需在不确定中迭代获取证据。基于前期交互评估工作,我们引入一种类OSCE的标准化患者模拟器,构建可复现的主动诊断询问基准。在468个病例和15个模型的实验中,多轮证据获取使诊断准确率相比全上下文评估下降12.75%,支持性证据质量下降24.36%;错误分析表明,性能下降与过早确诊和低效提问相关。结果表明,静态全上下文基准可能高估模型在交互式取证场景下的表现,呼吁发展互补的交互评估机制以保障临床决策支持的安全性。

原文摘要 · Abstract (English)

Large language models perform well on static medical examinations, yet clinical diagnosis often requires iterative evidence gathering under uncertainty. Building on prior interactive evaluation efforts, we introduce an OSCE-inspired standardized patient simulator and a controlled, reproducible benchmark for active diagnostic inquiry. Across 468 cases and 15 models in our protocol, we observe that multi-turn evidence seeking reduces diagnostic accuracy by 12.75% and lowers supporting-evidence quality by 24.36% relative to full-context evaluation; error analyses associate these drops with premature diagnostic closure and inefficient questioning. Together, these results suggest that static full-context benchmarks may overestimate performance in interactive evidence-seeking settings, motivating complementary interactive assessment for safer clinical decision support.

临床决策大模型交互评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。