通过建模视觉-语言关系提升多模态对齐效果
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
- 从多层视觉特征中提取语义关系线索
- 在8个基础模型上实现最优的图文问答与图像描述性能
- 适合需要精准语义对齐的多模态应用研究者
视觉-语言微调已成为构建多模态基础模型的有效范式。尽管文本上下文常揭示图像中的语义关系,但现有微调方法在对齐视觉与语言时通常忽略此类信息,导致性能受限。为此,我们提出一种基于语义与关系的多模态对齐与融合方法。首先,从不同视觉编码器中提取多层次语义特征,以捕捉更丰富的视觉关系线索;其次,学习将视觉特征映射到具有潜在关系的语义组;最后,通过可继承的交叉注意力融合视觉与文本特征,并全局剔除相关性低的视觉-语言特征对以去除冗余关系。我们在8个基础模型及2项下游任务(视觉问答、图像描述)上评估,结果表明该方法优于所有现有方法。
原文摘要 · Abstract (English)
Vision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook this information when aligning vision and language, thus leading to suboptimal performance. Toward solving this problem, we propose a method that can improve multimodal alignment and fusion based on both semantics and relationships.Specifically, we first extract multilevel semantic features from different vision encoder to capture more visual cues of the relationships. Then, we learn to project the vision features to group related semantics, among which are more likely to have relationships. Finally, we fuse the visual features with the textual by using inheritable cross-attention, where we globally remove the redundant visual relationships by discarding visual-language feature pairs with low correlation. We evaluate our proposed method on eight foundation models and two downstream tasks, visual question answering and image captioning, and show that it outperforms all existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。