arXiv:2602.19091cs.CV2026-02被引 6

用压缩机制统一提升多模态模型的检索与生成能力。

CREM: Compression-Driven Representation Enhancement for Multimodal Retrieval and Comprehension

  • 通过可学习的合唱令牌设计压缩提示,融合多模态语义。
  • 在MMEB上达顶尖检索性能,同时保持强生成能力。
  • 适合需要兼顾检索与生成的多模态应用开发者。

多模态大语言模型(MLLMs)在视觉描述和视觉问答等理解任务中表现优异,但直接应用于基于嵌入的检索任务仍具挑战,因其输出格式与优化目标存在差异。以往方法常采用对比微调适配检索,却牺牲了生成能力。我们认为生成与嵌入任务均依赖于共享的认知机制,即跨模态表示对齐与上下文理解。为此,我们提出CREM(压缩驱动表示增强模型),构建统一框架,在提升多模态表示用于检索的同时保留生成能力。具体而言,引入基于压缩的提示设计,使用可学习的合唱令牌聚合多模态语义,并设计压缩感知注意力的联合训练策略,融合对比与生成目标。大量实验表明,CREM在MMEB上达到最先进检索性能,同时在多个理解基准上保持强大生成表现。研究发现,在提出的压缩驱动范式下,生成监督可进一步提升MLLM的表示质量。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown remarkable success in comprehension tasks such as visual description and visual question answering. However, their direct application to embedding-based tasks like retrieval remains challenging due to the discrepancy between output formats and optimization objectives. Previous approaches often employ contrastive fine-tuning to adapt MLLMs for retrieval, but at the cost of losing their generative capabilities. We argue that both generative and embedding tasks fundamentally rely on shared cognitive mechanisms, specifically cross-modal representation alignment and contextual comprehension. To this end, we propose CREM (Compression-driven Representation Enhanced Model), with a unified framework that enhances multimodal representations for retrieval while preserving generative ability. Specifically, we introduce a compression-based prompt design with learnable chorus tokens to aggregate multimodal semantics and a compression-driven training strategy that integrates contrastive and generative objectives through compression-aware attention. Extensive experiments demonstrate that CREM achieves state-of-the-art retrieval performance on MMEB while maintaining strong generative performance on multiple comprehension benchmarks. Our findings highlight that generative supervision can further improve the representational quality of MLLMs under the proposed compression-driven paradigm.

多模态检索增强生成模型压缩机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。