arXiv:2510.27350cs.CV2025-10被引 23

RzenEmbed统一建模图文视频文档,提升多模态检索效果

RzenEmbed: Towards Comprehensive Multimodal Retrieval

  • 两阶段训练+硬度加权损失,聚焦难样本优化
  • 在MMEB上达新SOTA,视频与文档检索显著领先
  • 适合需要跨模态精准检索的开发者与研究者

多模态大模型的发展推动了基于CLIP的框架,生成强大的通用嵌入用于检索任务。然而,现有方法主要针对自然图像,对视频和视觉文档等关键视觉模态支持有限。为此,我们提出RzenEmbed,一个统一框架,可学习文本、图像、视频和视觉文档等多种模态的嵌入表示。采用新颖的两阶段训练策略:第一阶段聚焦基础文本与多模态检索;第二阶段引入改进的InfoNCE损失,包含两个关键增强:一是硬度加权机制,使模型在每批中优先关注困难样本;二是缓解假负样本影响并降低数据噪声。该策略不仅增强模型判别能力,还提升指令遵循性能。通过可学习温度参数与模型融合进一步提升表现。RzenEmbed在MMEB基准上达到新SOTA,不仅整体得分最优,且在挑战性的视频与视觉文档检索任务中全面超越先前工作。

原文摘要 · Abstract (English)

The rapid advancement of Multimodal Large Language Models (MLLMs) has extended CLIP-based frameworks to produce powerful, universal embeddings for retrieval tasks. However, existing methods primarily focus on natural images, offering limited support for other crucial visual modalities such as videos and visual documents. To bridge this gap, we introduce RzenEmbed, a unified framework to learn embeddings across a diverse set of modalities, including text, images, videos, and visual documents. We employ a novel two-stage training strategy to learn discriminative representations. The first stage focuses on foundational text and multimodal retrieval. In the second stage, we introduce an improved InfoNCE loss, incorporating two key enhancements. Firstly, a hardness-weighted mechanism guides the model to prioritize challenging samples by assigning them higher weights within each batch. Secondly, we implement an approach to mitigate the impact of false negatives and alleviate data noise. This strategy not only enhances the model's discriminative power but also improves its instruction-following capabilities. We further boost performance with learnable temperature parameter and model souping. RzenEmbed sets a new state-of-the-art on the MMEB benchmark. It not only achieves the best overall score but also outperforms all prior work on the challenging video and visual document retrieval tasks. Our models are available in https://huggingface.co/qihoo360/RzenEmbed.

多模态检索嵌入学习视频理解视觉文档

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。