临床多模态诊断中,不确定预测反而可能让模型更不靠谱。
An Empirical Analysis of Calibration and Selective Prediction in Multimodal Clinical Condition Classification
- 用多模态ICU数据测试模型不确定性判断能力
- 高不确定预测常对应正确结果,低不确定反出错
- 适合关注医疗AI安全性的研究者和开发者
随着人工智能系统向临床应用推进,确保预测可靠性对关键决策至关重要。一种保障策略是选择性预测:模型在不确定时将判断交给人工专家。本文基于多模态ICU数据,实证评估了多标签临床诊断中基于不确定性的选择性预测表现。在多种先进单模态与多模态模型上,我们发现尽管标准指标表现良好,选择性预测仍会显著降低性能。这一失败源于严重的类别依赖性校准偏差:模型对正确预测赋予过高不确定性,对错误预测则过低估计,尤其在少数类临床条件下更为明显。常见综合评估指标会掩盖这些现象,难以有效评估选择性预测行为。结果表明,在多模态临床诊断中,选择性预测存在特定任务下的失效模式,强调需采用校准感知的评估方法,以保障临床AI系统的安全性与鲁棒性。
原文摘要 · Abstract (English)
As artificial intelligence systems move toward clinical deployment, ensuring reliable prediction behavior is fundamental for safety-critical decision-making tasks. One proposed safeguard is selective prediction, where models can defer uncertain predictions to human experts for review. In this work, we empirically evaluate the reliability of uncertainty-based selective prediction in multilabel clinical condition classification using multimodal ICU data. Across a range of state-of-the-art unimodal and multimodal models, we find that selective prediction can substantially degrade performance despite strong standard evaluation metrics. This failure is driven by severe class-dependent miscalibration, whereby models assign high uncertainty to correct predictions and low uncertainty to incorrect ones, particularly for underrepresented clinical conditions. Our results show that commonly used aggregate metrics can obscure these effects, limiting their ability to assess selective prediction behavior in this setting. Taken together, our findings characterize a task-specific failure mode of selective prediction in multimodal clinical condition classification and highlight the need for calibration-aware evaluation to provide strong guarantees of safety and robustness in clinical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。