arXiv:2502.14888cs.CVcs.AI2025-02ACL被引 3

量化并利用视觉语言模型的模态差异,提升下游任务表现。

Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models

论文配图:Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
图 1 · 摘自论文原文
  • 提出模态主导性得分(MDS),区分视觉、语言与跨模态特征。
  • 无需训练即可改进性别分类偏见、生成对抗样本和图像生成控制。
  • 适合关注模型可解释性与轻量级编辑的研究者。

视觉语言模型的成功主要归功于跨模态对齐,但现有对齐算法仍存在模态差距,而这种差距在人类感知中是必要的,如视觉纹理和语言语调等模态特异性现象。为此,我们提出计算模态差距的方法。首先引入模态主导性得分(MDS),将多模态特征分为三类:视觉主导、语言主导和跨模态特征。接着,设计自动可解释性度量,实现大规模评估。最后证明,无需训练的模型编辑能有效提升多个下游任务,包括缓解性别分类偏差、生成跨模态对抗样本,以及在文生图中实现模态特异性控制。结合任务无关的可解释工具,本工作为系统分析和轻量级编辑多模态模型提供了新视角。

原文摘要 · Abstract (English)

The success of vision-language models is primarily attributed to effective alignment across modalities such as vision and language. However, modality gaps persist in existing alignment algorithms and appear necessary for human perception as evidenced by modality-specific phenomena like visual texture and linguistic tone. These observations motivate us to computationally measure and leverage modality gaps to improve downstream tasks. We first introduce the Modality Dominance Score (MDS), which attributes multimodal features to specific modalities by categorizing them into three classes: vision-dominant features, language-dominant features, and cross-modal features. We then propose automatic interpretability metrics to evaluate these modality-specific features in a scalable manner. Finally, we demonstrate that the training-free model editing enhances multiple downstream tasks, including mitigating bias in gender classification, generating cross-modal adversarial examples, and enabling modality-specific control in text-to-image generation. Combined with task-agnostic interpretability tools, our work offers insights for systematic analysis and lightweight editing of multimodal models.

多模态可解释性模型编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。