统一处理文本、图像、音频和视频的检索模型,支持跨模态联合搜索。
Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video
- 用统一模型处理四种模态的检索,支持跨模态与联合模态查询。
- 在文本、图像、视频检索任务上表现优异,提升真实文档场景下的召回率。
- 适合需要多模态信息检索的应用,如智能搜索、数字图书馆。
我们提出Omni-Embed-Nemotron,一种统一的多模态检索嵌入模型,旨在应对现实世界信息需求日益增长的复杂性。尽管检索增强生成(RAG)通过引入外部知识显著提升了语言模型能力,但现有基于文本的检索器依赖于干净、结构化的输入,在处理包含丰富视觉与语义内容的真实文档(如PDF、幻灯片或视频)时表现不佳。近期工作如ColPali表明,使用基于图像的表示保留文档布局可提升检索质量。受Qwen2.5-Omni等最新多模态模型的启发,我们进一步将检索能力扩展至音频和视频模态。Omni-Embed-Nemotron支持单一模型下的跨模态(如文本-视频)与联合模态(如文本-视频+音频)检索。本文描述了其架构、训练设置及评估结果,并验证了其在文本、图像和视频检索中的有效性。
原文摘要 · Abstract (English)
We present Omni-Embed-Nemotron, a unified multimodal retrieval embedding model developed to handle the increasing complexity of real-world information needs. While Retrieval-Augmented Generation (RAG) has significantly advanced language models by incorporating external knowledge, existing text-based retrievers rely on clean, structured input and struggle with the visually and semantically rich content found in real-world documents such as PDFs, slides, or videos. Recent work such as ColPali has shown that preserving document layout using image-based representations can improve retrieval quality. Building on this, and inspired by the capabilities of recent multimodal models such as Qwen2.5-Omni, we extend retrieval beyond text and images to also support audio and video modalities. Omni-Embed-Nemotron enables both cross-modal (e.g., text - video) and joint-modal (e.g., text - video+audio) retrieval using a single model. We describe the architecture, training setup, and evaluation results of Omni-Embed-Nemotron, and demonstrate its effectiveness in text, image, and video retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。