arXiv:2506.23298eess.IVcs.AI2025-06中稿 · MICCAI 2025 main c…被引 6

发现并修复医疗图像分类中多模态大模型的预测偏差与不公平性

Exposing and Mitigating Calibration Biases and Demographic Unfairness in MLLM Few-Shot In-Context Learning for Medical Image Classification

  • 提出分层校准方法CALIN,从群体到子群体逐级估计校准矩阵
  • 在3个医学数据集上提升准确率,同时显著降低不同人群间的预测偏差
  • 适合关注医疗AI公平性与可信度的研究者与临床部署团队

多模态大语言模型(MLLMs)在医学图像分析中具备少样本上下文学习的巨大潜力。然而,将其安全应用于真实临床场景需深入分析预测准确性及其相关校准误差,尤其在不同人口统计学子群体中的表现。本文首次系统研究了MLLMs在少样本上下文学习中用于医学图像分类时的校准偏差与人口统计不公平性。我们提出一种推理阶段校准方法CALIN,通过双层流程——从总体到子群体逐步估计所需校准量(以校准矩阵表示)——在推理时调整预测置信度。在三个医学影像数据集(PAPILA、HAM10000、MIMIC-CXR)上的实验表明,CALIN能有效实现公平的置信度校准,提升整体预测准确率,并呈现最小的公平性-性能权衡。代码已开源。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have enormous potential to perform few-shot in-context learning in the context of medical image analysis. However, safe deployment of these models into real-world clinical practice requires an in-depth analysis of the accuracies of their predictions, and their associated calibration errors, particularly across different demographic subgroups. In this work, we present the first investigation into the calibration biases and demographic unfairness of MLLMs' predictions and confidence scores in few-shot in-context learning for medical image classification. We introduce CALIN, an inference-time calibration method designed to mitigate the associated biases. Specifically, CALIN estimates the amount of calibration needed, represented by calibration matrices, using a bi-level procedure: progressing from the population level to the subgroup level prior to inference. It then applies this estimation to calibrate the predicted confidence scores during inference. Experimental results on three medical imaging datasets: PAPILA for fundus image classification, HAM10000 for skin cancer classification, and MIMIC-CXR for chest X-ray classification demonstrate CALIN's effectiveness at ensuring fair confidence calibration in its prediction, while improving its overall prediction accuracies and exhibiting minimum fairness-utility trade-off. Our codebase can be found at https://github.com/xingbpshen/medical-calibration-fairness-mllm.

医疗AI公平性校准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。