arXiv:2411.12785cs.CV2024-11CVPR被引 18

提出新方法同时消除图像与文本偏见,保持跨模态对齐能力。

Joint Vision-Language Social Bias Removal for CLIP

  • 通过均衡调节图文嵌入的偏见实现联合去偏。
  • 在去除偏见的同时维持模型跨模态对齐性能。
  • 设计新评估协议,更全面衡量去偏效果与泛化能力。

视觉-语言预训练模型如CLIP在下游任务中表现出色,但其固有的社会偏见严重限制了实际应用。现有方法通常通过移除模型嵌入中的偏见属性信息来缓解问题,但我们发现这些方法常伴随跨模态对齐能力的显著下降。分析表明,这是由于图像与文本嵌入去偏不均衡所致。为此,我们提出一种新型多模态去偏框架,先对齐图文偏见,再从双模态中同时移除。该方法在实现多模态去偏的同时,保持了去偏后嵌入的跨模态对齐性。此外,我们提出新的评估协议,可全面量化模型去偏能力和跨模态对齐水平,并评估去偏模型的泛化性能。本工作为未来研究提供了新思路与指导。

原文摘要 · Abstract (English)

Vision-Language (V-L) pre-trained models such as CLIP show prominent capabilities in various downstream tasks. Despite this promise, V-L models are notoriously limited by their inherent social biases. A typical demonstration is that V-L models often produce biased predictions against specific groups of people, significantly undermining their real-world applicability. Existing approaches endeavor to mitigate the social bias problem in V-L models by removing biased attribute information from model embeddings. However, after our revisiting of these methods, we find that their bias removal is frequently accompanied by greatly compromised V-L alignment capabilities. We then reveal that this performance degradation stems from the unbalanced debiasing in image and text embeddings. To address this issue, we propose a novel V-L debiasing framework to align image and text biases followed by removing them from both modalities. By doing so, our method achieves multi-modal bias mitigation while maintaining the V-L alignment in the debiased embeddings. Additionally, we advocate a new evaluation protocol that can 1) holistically quantify the model debiasing and V-L alignment ability, and 2) evaluate the generalization of social bias removal models. We believe this work will offer new insights and guidance for future studies addressing the social bias problem in CLIP.

去偏CLIP多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。