arXiv:2508.18886cs.CV2025-08

通过双分支提示学习,实现医学视觉语言模型的跨模态去偏对齐。

Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models

  • 设计双分支结构分离敏感属性与目标属性,实现跨模态解耦对齐。
  • 在8个医疗数据集上,仅用360万参数即达最佳公平性与准确率。
  • 适合关注医疗AI公平性、跨模态对齐的研究者和临床应用开发者。

在医疗诊断中保障不同人口群体的公平性对实现公平医疗至关重要,尤其在成像设备和临床实践差异导致的分布偏移下。视觉语言模型(VLMs)具备强泛化能力,文本提示可编码身份属性,便于显式识别并消除敏感方向。然而,现有去偏方法通常独立处理视觉与文本模态,导致残余的跨模态错位与公平性缺口。为此,我们提出DualFairVL,一种联合去偏与对齐的多模态提示学习框架。该框架采用并行双分支结构,分离敏感属性与目标属性,实现模态间的解耦且对齐表示。通过线性投影构建近似正交的文本锚点,引导交叉注意力生成融合特征;超网络进一步解耦属性信息,生成实例感知的视觉提示,编码双模态公平性与鲁棒性信号。在视觉分支中引入基于原型的正则化,强化敏感特征分离并增强与文本锚点的对齐。在四个模态的八个医疗影像数据集上的实验表明,DualFairVL在分布内与分布外设置下均达到当前最优的公平性与准确率,显著优于全微调及参数高效基线,仅需360万可训练参数。代码将在发表后公开。

原文摘要 · Abstract (English)

Ensuring fairness across demographic groups in medical diagnosis is essential for equitable healthcare, particularly under distribution shifts caused by variations in imaging equipment and clinical practice. Vision-language models (VLMs) exhibit strong generalization, and text prompts encode identity attributes, enabling explicit identification and removal of sensitive directions. However, existing debiasing approaches typically address vision and text modalities independently, leaving residual cross-modal misalignment and fairness gaps. To address this challenge, we propose DualFairVL, a multimodal prompt-learning framework that jointly debiases and aligns cross-modal representations. DualFairVL employs a parallel dual-branch architecture that separates sensitive and target attributes, enabling disentangled yet aligned representations across modalities. Approximately orthogonal text anchors are constructed via linear projections, guiding cross-attention mechanisms to produce fused features. A hypernetwork further disentangles attribute-related information and generates instance-aware visual prompts, which encode dual-modal cues for fairness and robustness. Prototype-based regularization is applied in the visual branch to enforce separation of sensitive features and strengthen alignment with textual anchors. Extensive experiments on eight medical imaging datasets across four modalities show that DualFairVL achieves state-of-the-art fairness and accuracy under both in- and out-of-distribution settings, outperforming full fine-tuning and parameter-efficient baselines with only 3.6M trainable parameters. Code will be released upon publication.

医疗AI公平性多模态提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。