arXiv:2605.08533cs.AI2026-05

让医生与大模型对话,显著提升急诊诊断准确率。

Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care

  • 医生通过多轮对话向大模型提问,结合完整病历逐步推理。
  • 住院医在难题上的诊断正确率从58.9%提升至73.4%。
  • 适合急诊科医生、临床AI研究者参考,尤其关注人机协作。

急诊医学要求在不确定性下快速准确决策。尽管大模型性能不断提升,但其作为实时诊疗助手的证据仍有限。本研究设计了MedSyn系统,让医生在仅见主诉的情况下,逐步向大模型查询完整病历信息。七名医生(三名高年资,四名住院医)完成了52例来自MIMIC-IV的数据集,按难度分层。盲评显示,住院医在难题中的正确率从0.589提升至0.734;标准化完全正确率改善0.092(p=0.071;d=0.47),属中等效应。自动评估指标也证实提升:标准化任意匹配准确率提高0.156(p<0.0001),住院医F1得分提升0.138(p<0.0001)。对话分析发现,资深医生倾向于目标明确的假设驱动提问,而住院医依赖更广泛的问题;跨经验水平的一致性提升0.145(p<0.0001)。交互式大模型支持能有效增强诊断推理。

原文摘要 · Abstract (English)

Clinical decision-making in emergency medicine demands rapid, accurate diagnoses under uncertainty. Despite benchmark progress, evidence for LLMs as interactive aids in live physician workflows remains sparse. MedSyn lets physicians iteratively query an LLM provided with the full clinical record while initially viewing only the chief complaint. Seven physicians (three seniors, four residents) completed baseline and AI-assisted sessions across 52 MIMIC-IV cases stratified by difficulty. Blinded evaluation showed residents' Hard-case correctness rose from 0.589 to 0.734; difficulty-standardised completely-correct rates confirmed a medium effect (Δ = 0.092; p = 0.071; d = 0.47). Automated metrics corroborated these gains: standardised any-match accuracy improved by 0.156 (p < 0.0001), and residents showed the largest F1 gain (Δ = 0.138; p < 0.0001). Dialogue analysis revealed expertise-dependent strategies (seniors asked targeted, hypothesis-driven questions; residents relied on broader queries) and cross-expertise concordance increased (Δ = 0.145; p < 0.0001). Interactive LLM support meaningfully enhances diagnostic reasoning.

急诊医学人机协作大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。