提出自适应融合全局与细粒度感知的多模态嵌入方法,提升复杂场景理解能力。
Adaptive Global and Fine-Grained Perceptual Fusion for MLLM Embeddings Compatible with Hard Negative Amplification
- 通过提示MLLM生成多维语义嵌入并自适应融合,兼顾全局与细粒度信息。
- 在MMEB和MMVP-VLM上达到当前最优性能,显著提升细粒度理解能力。
- 无需数据编辑即可增强批内难例,适合需要强泛化能力的应用场景。
多模态嵌入作为视觉与语言对齐的桥梁,现有基于CLIP和MLLM的模型仅能捕捉全局语义信息。尽管已有研究关注细粒度理解,但当前MLLM面临复杂场景中同时包含全局与细粒度感知的混合模式,亟需兼容的融合机制。本文提出自适应全局与细粒度感知融合的MLLM嵌入方法(AGFF-Embed),通过提示MLLM生成多个侧重不同语义维度的嵌入,并实现自适应、平滑聚合。此外,结合显式梯度放大技术(EGA),在不修改数据集的前提下实现批内难例增强。在MMEB和MMVP-VLM基准上的评估表明,AGFF-Embed在通用与细粒度理解任务上均达到当前最优表现。
原文摘要 · Abstract (English)
Multimodal embeddings serve as a bridge for aligning vision and language, with the two primary implementations -- CLIP-based and MLLM-based embedding models -- both limited to capturing only global semantic information. Although numerous studies have focused on fine-grained understanding, we observe that complex scenarios currently targeted by MLLM embeddings often involve a hybrid perceptual pattern of both global and fine-grained elements, thus necessitating a compatible fusion mechanism. In this paper, we propose Adaptive Global and Fine-grained perceptual Fusion for MLLM Embeddings (AGFF-Embed), a method that prompts the MLLM to generate multiple embeddings focusing on different dimensions of semantic information, which are then adaptively and smoothly aggregated. Furthermore, we adapt AGFF-Embed with the Explicit Gradient Amplification (EGA) technique to achieve in-batch hard negatives enhancement without requiring fine-grained editing of the dataset. Evaluation on the MMEB and MMVP-VLM benchmarks shows that AGFF-Embed comprehensively achieves state-of-the-art performance in both general and fine-grained understanding compared to other multimodal embedding models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。