透明AI辅助可显著提升睡眠觉醒事件诊断准确率,且受临床医生欢迎。
Assessing the Real-World Utility of Explainable AI for Arousal Diagnostics: An Application-Grounded User Study
- 用透明白盒AI作为事后质检,比黑箱辅助提升30%诊断准确率。
- 人机协作使评分误差降低,且减少医生间差异,计数结果更可靠。
- 医生偏好透明、适时的AI辅助,利于临床系统落地推广。
人工智能在生物医学信号解析中已达到或超过人类专家水平,但其临床应用不仅需高精度,还需医生能理解何时何地信任算法建议。本研究基于真实应用场景,对8位专业睡眠医学医师开展实验,让他们在三种条件下评估多导睡眠图中的夜间觉醒事件:(i)纯人工评分,(ii)黑箱AI辅助,(iii)透明白盒AI辅助。辅助方式包括从开始介入或作为事后质量控制(QC)。系统评估了不同辅助类型与时机对事件级和以计数为核心的临床性能、耗时及用户体验的影响。结果表明,无论哪种方式,人机协作均显著优于独立专家,且降低评分者间差异。尤其值得注意的是,以靶向质量控制形式部署的透明AI,相比黑箱辅助可实现约30%的事件级性能提升,且后置质检进一步改善计数准确性。尽管白盒与质检模式增加耗时,但起始阶段介入的辅助更高效,多数参与者表示偏好。七名医生明确表示愿接受该系统无需重大修改。综上,策略性地使用透明AI辅助,可在保证准确性的前提下兼顾临床效率,是实现可信、可接受的医疗AI集成的有效路径。
原文摘要 · Abstract (English)
Artificial intelligence (AI) systems increasingly match or surpass human experts in biomedical signal interpretation. However, their effective integration into clinical practice requires more than high predictive accuracy. Clinicians must discern \textit{when} and \textit{why} to trust algorithmic recommendations. This work presents an application-grounded user study with eight professional sleep medicine practitioners, who score nocturnal arousal events in polysomnographic data under three conditions: (i) manual scoring, (ii) black-box (BB) AI assistance, and (iii) transparent white-box (WB) AI assistance. Assistance is provided either from the \textit{start} of scoring or as a post-hoc quality-control (\textit{QC}) review. We systematically evaluate how the type and timing of assistance influence event-level and clinically most relevant count-based performance, time requirements, and user experience. When evaluated against the clinical standard used to train the AI, both AI and human-AI teams significantly outperform unaided experts, with collaboration also reducing inter-rater variability. Notably, transparent AI assistance applied as a targeted QC step yields median event-level performance improvements of approximately 30\% over black-box assistance, and QC timing further enhances count-based outcomes. While WB and QC approaches increase the time required for scoring, start-time assistance is faster and preferred by most participants. Participants overwhelmingly favor transparency, with seven out of eight expressing willingness to adopt the system with minor or no modifications. In summary, strategically timed transparent AI assistance effectively balances accuracy and clinical efficiency, providing a promising pathway toward trustworthy AI integration and user acceptance in clinical workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。