让大模型更懂图像中关键主体,提升图文检索精准度
Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval

- 通过显式建模视觉显著主体,引导跨模态注意力对齐
- 在MMEB基准上达到当前最佳性能,显著减少语义偏差
- 适合需要精细图文匹配的视觉语言任务研究者
尽管统一多模态检索(UMR)借助大模型取得进展,现有嵌入方法主要依赖样本级对比学习,忽视了主体级语义。这导致复杂查询中难以准确关联语义一致的视觉主体,出现定位偏差——模型无法精准识别文本提及的视觉区域。同时,缺乏对显著视觉主体的显式建模,使大模型过度依赖文本线索,造成视觉模态忽视与知识利用不足。为此,我们提出显著主体感知多模态嵌入(SSA-ME),通过大模型与视觉专家联合识别并强调图像-文本对中的显著视觉概念,并引入显著性引导目标,使跨模态注意力更聚焦于语义有意义区域。此外,特征再生模块基于显著图重新校准视觉特征,实现模态间平衡且语义连贯的融合。大量实验表明,该方法在MMEB基准上达到领先性能,证明主体级建模可显著提升多模态检索效果。定性分析进一步验证了方法的可解释性与有效性。
原文摘要 · Abstract (English)
Despite significant progress in Unified Multimodal Retrieval (UMR) powered by Large Multimodal Models (LMMs), existing embedding methods primarily focus on sample-level objectives via contrastive learning while overlooking the crucial subject-level semantics. This limitation hinders the model's ability to group semantically coherent subjects in complex multimodal queries, manifesting as semantic alignment deviation--where models fail to accurately localize salient text-referred regions in visual content. Moreover, without explicit guidance to model salient visual subjects, LMMs tend to over-rely on textual cues, resulting in visual modality neglect and suboptimal utilization of visual knowledge. To this end, we propose Salient Subject-Aware Multimodal Embedding (SSA-ME), a novel framework designed to enhance fine-grained representation learning through saliency-aware modeling. SSA-ME leverages LMMs and visual experts to identify and emphasize salient visual concepts in image-text pairs, and introduces a saliency-guided objective to better align cross-modal attention with semantically meaningful regions. Additionally, a feature regeneration module recalibrates visual features based on the derived saliency maps, ensuring a balanced and semantically coherent integration across modalities. Extensive experiments show that our method achieves state-of-the-art performance on the MMEB benchmark, demonstrating that incorporating subject-level modeling substantially improves multimodal retrieval. Comprehensive qualitative analyses further illustrate the interpretability and effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。