用多专家模拟框架实时评估脑出血AI诊断可信度
Automated Real-time Assessment of Intracranial Hemorrhage Detection AI Using an Ensembled Monitoring Model (EMM)
- 构建集成监测模型,通过多视角评估AI预测置信度
- 在2919例数据上实现可信度分类,指导不同处置策略
- 无需访问AI内部结构,适合临床落地的黑箱模型监控
放射科人工智能工具部署后通常缺乏监控。由于无法实时评估每例AI预测的置信度,使用者需自行判断结果可靠性,增加了认知负担,降低效率,可能引发误诊。为此,我们提出集成监测模型(EMM),借鉴临床多专家共识机制,专为黑箱商业AI设计。EMM独立运行,无需访问内部组件或中间输出,仍可提供可靠的置信度评估。以2919例多样化颅内出血检测数据为测试集,证明EMM能有效分类预测置信度,建议不同应对措施,提升AI整体性能,减轻医生负担。同时提供技术实施要点与临床转化最佳实践。
原文摘要 · Abstract (English)
Artificial intelligence (AI) tools for radiology are commonly unmonitored once deployed. The lack of real-time case-by-case assessments of AI prediction confidence requires users to independently distinguish between trustworthy and unreliable AI predictions, which increases cognitive burden, reduces productivity, and potentially leads to misdiagnoses. To address these challenges, we introduce Ensembled Monitoring Model (EMM), a framework inspired by clinical consensus practices using multiple expert reviews. Designed specifically for black-box commercial AI products, EMM operates independently without requiring access to internal AI components or intermediate outputs, while still providing robust confidence measurements. Using intracranial hemorrhage detection as our test case on a large, diverse dataset of 2919 studies, we demonstrate that EMM successfully categorizes confidence in the AI-generated prediction, suggesting different actions and helping improve the overall performance of AI tools to ultimately reduce cognitive burden. Importantly, we provide key technical considerations and best practices for successfully translating EMM into clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。