用多样的负样本提升图文模型对齐效果
Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
- 基于CLIP空间和稀疏自编码器提取语义差异因子
- 通过重构难度与语义多样性筛选多负样本,提升监督信息量
- 引入重要性采样策略,高效处理多负样本比较
直接偏好优化(DPO)已从纯文本模型扩展至视觉-语言模型。然而,现有方法依赖过于简化的成对比较,仅通过基础扰动或相似性检索生成单一负图像,无法捕捉多模态偏好的复杂性,导致优化偏差和幻觉问题。为此,我们提出MISP-DPO,首个在多模态DPO中利用语义多样负图像的框架,采用Plackett-Luce模型实现多负样本建模。该方法将提示词与候选图像嵌入到CLIP空间,使用稀疏自编码器揭示可解释的语义偏差因素。负样本依据重构难度、与正样本的语义偏离度及彼此间多样性进行选择,从而提供更广泛且更具信息量的监督信号。为处理多负样本比较,采用Plackett-Luce目标函数并引入重要性采样策略,显著提升训练效率。在五个多样化基准上的实验表明,MISP-DPO在多模态对齐性能上持续优于先前方法,验证了语义感知的多负样本采样在基于偏好学习中的有效性。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has recently been extended from text-only models to vision-language models. However, existing methods rely on oversimplified pairwise comparisons, generating a single negative image via basic perturbations or similarity-based retrieval, which fail to capture the complex nature of multimodal preferences, inducing optimization bias and hallucinations. To address this issue, we propose MISP-DPO, the first framework to incorporate multiple, semantically diverse negative images in multimodal DPO via the Plackett-Luce model. Our method embeds prompts and candidate images in CLIP (Contrastive Language-Image Pretraining) space and applies a sparse autoencoder to uncover semantic deviations into interpretable factors. Negative samples are selected based on reconstruction difficulty, semantic deviation from the positive, and mutual diversity, yielding broader and more informative supervision. To handle multi-negative comparisons, we adopt a Plackett-Luce objective and introduce an importance sampling strategy that improves training efficiency. Experiments across five diverse benchmarks demonstrate that MISP-DPO consistently improves multimodal alignment over prior methods, validating the effectiveness of semantic-aware, multi-negative sampling in preference-based learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。