AI助诊系统让医生看得懂、信得过,实测降低认知负担。
Augmenting Clinical Decision-Making with an Interactive and Interpretable AI Copilot: A Real-World User Study with Clinicians in Nephrology and Obstetrics
- 基于电子病历动态预测风险,用可视化与大模型推荐增强可解释性。
- 16名医生实测显示任务完成更快、错误更少,工作负荷显著下降。
- 新手用它理清思路,专家拿它验证逻辑,适配不同临床风格。
临床医生对黑箱AI存在疑虑,限制了其在高风险医疗场景的应用。本文提出AICare,一个交互式且可解释的AI协作者,用于辅助临床决策。该系统通过分析纵向电子健康记录,将动态风险预测与可理解的可视化及大语言模型驱动的诊断建议相结合。我们在肾病学和妇产科领域对16名医生开展了被试内平衡实验,综合评估了任务完成时间、错误率、主观评价(NASA-TLX、SUS、信心评分)及半结构化访谈。结果表明,AICare有效降低了认知负荷。定性分析进一步发现,信任是通过验证建立的:初级医生将其作为认知支架来组织分析,而资深医生则采取对抗性验证来挑战AI逻辑。本研究为设计透明协作型AI系统提供了实践启示,支持不同推理风格,真正实现对临床判断的增强而非替代。
原文摘要 · Abstract (English)
Clinician skepticism toward opaque AI hinders adoption in high-stakes healthcare. We present AICare, an interactive and interpretable AI copilot for collaborative clinical decision-making. By analyzing longitudinal electronic health records, AICare grounds dynamic risk predictions in scrutable visualizations and LLM-driven diagnostic recommendations. Through a within-subjects counterbalanced study with 16 clinicians across nephrology and obstetrics, we comprehensively evaluated AICare using objective measures (task completion time and error rate), subjective assessments (NASA-TLX, SUS, and confidence ratings), and semi-structured interviews. Our findings indicate AICare's reduced cognitive workload. Beyond performance metrics, qualitative analysis reveals that trust is actively constructed through verification, with interaction strategies diverging by expertise: junior clinicians used the system as cognitive scaffolding to structure their analysis, while experts engaged in adversarial verification to challenge the AI's logic. This work offers design implications for creating AI systems that function as transparent partners, accommodating diverse reasoning styles to augment rather than replace clinical judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。