AI辅助脑动脉瘤检测未提升医生准确率,反而增加读片时间。
Assessing workflow impact and clinical utility of AI-assisted brain aneurysm detection: a multi-reader study
- 用2名经验不同的放射科医生对比有无AI辅助的诊断表现。
- 模型在测试集上敏感度74%、假阳性率1.6%,但医生表现无显著提升。
- AI使读片时间平均增加15秒,且医生信心未受影响,适合临床验证研究者参考。
尽管已有大量基于AI的放射学异常检测算法,但其在临床环境中的整合效果极少被评估。本文通过多读者研究,评估了一款脑动脉瘤检测AI模型的临床适用性与实用性,比较了两名经验不同(2年和13年)的放射科医生的表现。使用公开的Time-Of-Flight Magnetic Resonance Angiography数据集(N=460),其中360例用于训练/验证,100例作为未见测试集进行阅片。尽管模型在测试集上达到74%敏感度和1.6%假阳性率,但无论是初级还是资深医生,其敏感度均未显著提升(p=0.59, p=1)。此外,两位医生在AI辅助下的读片时间均显著延长(平均+15秒;p=3×10⁻⁴ 和 p=3×10⁻⁵)。医生的诊断信心在两种条件下保持不变。结果强调了在真实临床环境中对AI算法进行有效性与工作流影响评估的重要性,提醒研究者关注算法的实际应用效果。
原文摘要 · Abstract (English)
Despite the plethora of AI-based algorithms developed for anomaly detection in radiology, subsequent integration into clinical setting is rarely evaluated. In this work, we assess the applicability and utility of an AI-based model for brain aneurysm detection comparing the performance of two readers with different levels of experience (2 and 13 years). We aim to answer the following questions: 1) Do the readers improve their performance when assisted by the AI algorithm? 2) How much does the AI algorithm impact routine clinical workflow? We reuse and enlarge our open-access, Time-Of-Flight Magnetic Resonance Angiography dataset (N=460). We use 360 subjects for training/validating our algorithm and 100 as unseen test set for the reading session. Even though our model reaches state-of-the-art results on the test set (sensitivity=74%, false positive rate=1.6), we show that neither the junior nor the senior reader significantly increase their sensitivity (p=0.59, p=1, respectively). In addition, we find that reading time for both readers is significantly higher in the "AI-assisted" setting than in the "Unassisted" (+15 seconds, on average; p=3x10^(-4) junior, p=3x10^(-5) senior). The confidence reported by the readers is unchanged across the two settings, indicating that the AI assistance does not influence the certainty of the diagnosis. Our findings highlight the importance of clinical validation of AI algorithms in a clinical setting involving radiologists. This study should serve as a reminder to the community to always examine the real-word effectiveness and workflow impact of proposed algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。