arXiv:2505.01851cs.CV2025-05

提出联邦视觉语言模型公平性优化框架,显著降低群体偏差。

Mitigating Group-Level Fairness Disparities in Federated Visual Language Models

  • 通过反事实正则化调整潜在偏见嵌入,实现跨层公平提示。
  • 在图像表征中消除群体偏差,平均减少45%的公平性差距。
  • 适合关注隐私保护下多模态模型公平性的研究者与开发者。

视觉语言模型(VLMs)在多模态任务中表现卓越,但在联邦学习(FL)环境中维护不同人口群体间的公平性仍面临挑战。本文提出FVL-FP框架,结合联邦学习与公平提示调优技术,以缓解群体公平性问题。核心包含三个创新组件:(1) 跨层人口公平提示(CDFP),通过反事实正则化调整潜在偏见嵌入;(2) 人口子空间正交投影(DSOP),将公平提示文本映射至群体子空间以去除图像表示中的偏见;(3) 公平感知提示融合(FPF),动态平衡客户端贡献,兼顾性能与公平性指标。在四个基准数据集上的实验表明,该方法相比标准联邦学习平均降低45%的人口偏差,同时保持任务性能在当前最优结果的6%以内。FVL-FP有效应对联邦设置下的非独立同分布数据问题,计算开销极小,为隐私保护型多模态系统中的公平性能提供高效解决方案。

原文摘要 · Abstract (English)

Visual language models (VLMs) have shown remarkable capabilities in multimodal tasks but face challenges in maintaining fairness across demographic groups, particularly when deployed in federated learning (FL) environments. This paper addresses the critical issue of group fairness in federated VLMs by introducing FVL-FP, a novel framework that combines FL with fair prompt tuning techniques. We focus on mitigating demographic biases while preserving model performance through three innovative components: (1) Cross-Layer Demographic Fair Prompting (CDFP), which adjusts potentially biased embeddings through counterfactual regularization; (2) Demographic Subspace Orthogonal Projection (DSOP), which removes demographic bias in image representations by mapping fair prompt text to group subspaces; and (3) Fair-aware Prompt Fusion (FPF), which dynamically balances client contributions based on both performance and fairness metrics. Extensive evaluations across four benchmark datasets demonstrate that our approach reduces demographic disparity by an average of 45\% compared to standard FL approaches, while maintaining task performance within 6\% of state-of-the-art results. FVL-FP effectively addresses the challenges of non-IID data distributions in federated settings and introduces minimal computational overhead while providing significant fairness benefits. Our work presents a parameter-efficient solution to the critical challenge of ensuring equitable performance across demographic groups in privacy-preserving multimodal systems.

联邦学习公平性多模态提示调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。