arXiv:2502.12799cs.CLcs.CV2025-02ACL被引 3

提出图文交错检索新任务,解决多图多文混合内容的语义理解难题。

Towards Text-Image Interleaved Retrieval

  • 设计图文交错查询生成管道,构建真实场景下的TIIR基准数据集。
  • 提出的马特罗什卡多模态嵌入器在减少视觉标记数的同时显著提升检索效果。
  • 适合关注多模态检索、大模型高效推理的研究者与应用开发者。

当前多模态信息检索研究主要聚焦于单图像输入,限制了包含多图像与图文交错内容的实际应用场景。本文提出图文交错检索(TIIR)任务,其中查询和文档均为交错的文本-图像序列,要求模型理解交错上下文中的语义以实现有效检索。基于自然生成的wikiHow教程,我们构建了一个TIIR基准数据集,并设计专用流程生成交错查询。为探索该任务,我们适配了多个现成检索器,并通过交错多模态大模型(MLLM)建立密集基线。随后提出新型马特罗什卡多模态嵌入器(MME),通过在不同粒度下压缩视觉标记数量,缓解基于MLLM的TIIR模型中视觉标记过多的问题。实验表明,简单适配现有模型无法持续获得良好效果。我们的MME在显著减少视觉标记数的前提下,相比基线取得显著性能提升。我们提供了详尽分析,并将公开数据集与代码以促进后续研究。

原文摘要 · Abstract (English)

Current multimodal information retrieval studies mainly focus on single-image inputs, which limits real-world applications involving multiple images and text-image interleaved content. In this work, we introduce the text-image interleaved retrieval (TIIR) task, where the query and document are interleaved text-image sequences, and the model is required to understand the semantics from the interleaved context for effective retrieval. We construct a TIIR benchmark based on naturally interleaved wikiHow tutorials, where a specific pipeline is designed to generate interleaved queries. To explore the task, we adapt several off-the-shelf retrievers and build a dense baseline by interleaved multimodal large language model (MLLM). We then propose a novel Matryoshka Multimodal Embedder (MME), which compresses the number of visual tokens at different granularity, to address the challenge of excessive visual tokens in MLLM-based TIIR models. Experiments demonstrate that simple adaption of existing models does not consistently yield effective results. Our MME achieves significant improvements over the baseline by substantially fewer visual tokens. We provide extensive analysis and will release the dataset and code to facilitate future research.

多模态检索图文交错大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。