提升医疗视觉语言模型的公平性与鲁棒性,避免对特定患者群体偏见。
Robust Fairness Vision-Language Learning for Medical Image Analysis
- 通过动态坏样本挖掘识别并修正错误图文对。
- 利用Sinkhorn距离使各群体损失分布保持一致,提升公平性。
- 在公平性指标上提升8.6%,适合医疗AI公平性研究者。
视觉语言模型(VLM)在医疗图像分析中具备处理多模态输入、超越传统方法性能的潜力。然而,在实际应用中,公平性与鲁棒性至关重要,以确保模型对所有患者都保持可靠。本文提出一个保障VLM公平性与鲁棒性的框架:训练时通过动态坏样本挖掘算法识别并修正有偏差的图像-文本配对,并利用Sinkhorn距离约束受保护群体的损失分布不偏离整体损失分布。实验表明,该框架在公平性加权AUC上最高提升8.6%。
原文摘要 · Abstract (English)
The advent of Vision-Language Models (VLMs) in medical image analysis has the potential to help process multimodal inputs and increase performance over traditional inference methods. However, when considering the domain in which these models will be implemented, fairness and robustness are important to ensure the model stays true for any patient. In this paper, we introduce a framework for ensuring robustness and fairness of VLM models. This framework modifies the loss function at training by identifying and adjusting faulty image-text pairs through a Dynamic Bad Pair Mining algorithm and also utilizing Sinkhorn distance to ensure the loss distributions of protected groups do not deviate from the total loss. Experimental testing of our framework shows up to a 8.6\% improvement when looking at equity-scaled AUC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。