arXiv:2503.23819cs.LGcs.AI2025-03被引 3

用置信分析量化皮肤病变模型在不同人群的预测不确定性,提升临床AI公平性与可信度。

Conformal uncertainty quantification to evaluate predictive fairness of foundation AI model for skin lesion classes across patient demographics

  • 采用无模型依赖的置信分析法,为每个预测生成不确定性分数。
  • 发现性别、年龄、种族差异导致预测不确定性显著变化,揭示模型公平性短板。
  • 可作为评估大模型临床部署公平性的新指标,适合医疗AI开发者参考。

基于深度学习的医学图像诊断AI系统性能已接近人类专家水平,但其复杂黑箱特性限制了在高风险医疗场景的应用。尤其是近年大规模基础模型(如Google DermFoundation)通过自监督方式在海量数据上训练,虽具备强泛化能力,但其生成的特征嵌入过程不可解释,难以获得临床信任。本文针对皮肤病变分类任务,利用置信分析方法,在多个公开基准数据集上量化视觉变换器(ViT)基础模型在不同患者群体(按性别、年龄、种族划分)中的预测不确定性。该方法不依赖具体模型,可在群体层面保证覆盖率,同时为每个个体提供不确定性评分。训练中采用基于动态F1得分的采样策略缓解类别不平衡问题,并研究此步骤对不确定性量化的影响。结果表明,该方法可有效评估基础模型特征嵌入的鲁棒性,为提升临床AI的可信度与公平性提供新路径。

原文摘要 · Abstract (English)

Deep learning based diagnostic AI systems based on medical images are starting to provide similar performance as human experts. However these data hungry complex systems are inherently black boxes and therefore slow to be adopted for high risk applications like healthcare. This problem of lack of transparency is exacerbated in the case of recent large foundation models, which are trained in a self supervised manner on millions of data points to provide robust generalisation across a range of downstream tasks, but the embeddings generated from them happen through a process that is not interpretable, and hence not easily trustable for clinical applications. To address this timely issue, we deploy conformal analysis to quantify the predictive uncertainty of a vision transformer (ViT) based foundation model across patient demographics with respect to sex, age and ethnicity for the tasks of skin lesion classification using several public benchmark datasets. The significant advantage of this method is that conformal analysis is method independent and it not only provides a coverage guarantee at population level but also provides an uncertainty score for each individual. We used a model-agnostic dynamic F1-score-based sampling during model training, which helped to stabilize the class imbalance and we investigate the effects on uncertainty quantification (UQ) with or without this bias mitigation step. Thus we show how this can be used as a fairness metric to evaluate the robustness of the feature embeddings of the foundation model (Google DermFoundation) and thus advance the trustworthiness and fairness of clinical AI.

AI公平性不确定性量化医学影像基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。