arXiv:2508.09087cs.CV2025-08被引 1

不依赖人口属性标签,用无监督聚类减少眼科影像模型的偏见

Addressing Bias in VLMs for Glaucoma Detection Without Protected Attribute Supervision

  • 通过图像嵌入聚类自动发现潜在人群子组
  • 用梯度相似性加权,提升表现差的子组预测性能
  • 适合关注医疗AI公平性的研究者和临床应用开发者

视觉语言模型在多模态任务中表现卓越,但在训练中未显式使用人口属性时仍可能产生偏差。本文聚焦视网膜眼底图像的青光眼自动筛查,针对该任务在弱势群体中更易误诊的问题,提出一种无需保护属性监督的去偏方法。基于重加权对比学习框架,首先通过无监督聚类提取图像嵌入中的代理子组;其次计算CLIP式多模态损失与SimCLR式图像对对比损失之间的梯度相似性权重;最后在联合top-k加权目标中,对表现较差的子组进行自适应强化。该方法在Harvard FairVLMed青光眼子集上评估,使用等几率距离(EOD)、等子组AUC(ES AUC)和分组AUC衡量跨推断人口子组的公平性,验证了其降低子组差异的有效性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have achieved remarkable success on multimodal tasks such as image-text retrieval and zero-shot classification, yet they can exhibit demographic biases even when explicit protected attributes are absent during training. In this work, we focus on automated glaucoma screening from retinal fundus images, a critical application given that glaucoma is a leading cause of irreversible blindness and disproportionately affects underserved populations. Building on a reweighting-based contrastive learning framework, we introduce an attribute-agnostic debiasing method that (i) infers proxy subgroups via unsupervised clustering of image-image embeddings, (ii) computes gradient-similarity weights between the CLIP-style multimodal loss and a SimCLR-style image-pair contrastive loss, and (iii) applies these weights in a joint, top-$k$ weighted objective to upweight underperforming clusters. This label-free approach adaptively targets the hardest examples, thereby reducing subgroup disparities. We evaluate our method on the Harvard FairVLMed glaucoma subset, reporting Equalized Odds Distance (EOD), Equalized Subgroup AUC (ES AUC), and Groupwise AUC to demonstrate equitable performance across inferred demographic subgroups.

医疗AI去偏多模态青光眼

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。