arXiv:2411.08923cs.CVcs.LG2024-11ICLR被引 5

用偏好优化提升对比学习模型的鲁棒性与公平性

Aligning Visual Contrastive learning models via Preference Optimization

  • 引入偏好优化训练对比学习模型,对齐人类偏好
  • 在字体攻击下表现更优,同时保持下游任务准确率
  • 适合需要公平性与抗干扰能力的应用场景

对比学习模型通过嵌入空间中的表示对齐,展现了捕捉语义相似性的强大能力。然而,其性能受限于训练数据的质量及其内在偏差。尽管偏好优化(PO)方法如基于人类反馈的强化学习(RLHF)和直接偏好优化(DPO)已被用于对齐生成模型与人类偏好,但尚未应用于对比学习。本文提出一种新方法,利用不同PO方法训练对比学习模型,以分解复杂概念。该方法系统性地将模型行为对齐至期望偏好,提升目标任务表现。特别地,我们关注增强模型对抗字体攻击和归纳偏见的鲁棒性,这在视觉-语言模型如CLIP中常见。实验表明,使用PO训练的模型优于标准对比学习方法,同时保留处理对抗挑战的能力,并在其他下游任务上保持高精度。该方法适用于需要公平性、鲁棒性及特定偏好对齐的任务。我们在图像字体攻击场景下评估该方法,并探索其解耦性别概念与缓解性别偏见的能力,展示方法的多样性。

原文摘要 · Abstract (English)

Contrastive learning models have demonstrated impressive abilities to capture semantic similarities by aligning representations in the embedding space. However, their performance can be limited by the quality of the training data and its inherent biases. While Preference Optimization (PO) methods such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) have been applied to align generative models with human preferences, their use in contrastive learning has yet to be explored. This paper introduces a novel method for training contrastive learning models using different PO methods to break down complex concepts. Our method systematically aligns model behavior with desired preferences, enhancing performance on the targeted task. In particular, we focus on enhancing model robustness against typographic attacks and inductive biases, commonly seen in contrastive vision-language models like CLIP. Our experiments demonstrate that models trained using PO outperform standard contrastive learning techniques while retaining their ability to handle adversarial challenges and maintain accuracy on other downstream tasks. This makes our method well-suited for tasks requiring fairness, robustness, and alignment with specific preferences. We evaluate our method for tackling typographic attacks on images and explore its ability to disentangle gender concepts and mitigate gender bias, showcasing the versatility of our approach.

对比学习偏好优化鲁棒性公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。