arXiv:2505.02466cs.IR2025-05中稿 · SIGIR 2025被引 33

Tevatron 2.0统一多模态跨语言检索,支持从文本到音视频的通用向量模型。

Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality

  • 构建统一流水线,支持多语言、多模态、不同规模的检索模型训练与评估
  • 推出OmniEmbed,首个融合文本、图像、视频、音频的统一嵌入模型
  • 适用于跨语言和跨模态检索研究,适合需要高效检索系统的研究者

大型语言模型(LLMs)的发展推动了百亿级检索模型的研究,这些模型在多种检索任务和语言间具有强泛化能力。同时,大视觉-语言模型的进步为多模态检索带来新机遇。为此,我们更新了Tevatron工具包,引入统一管道,支持研究人员在不同规模、多语言及多种模态下探索检索模型。本演示论文突出该工具包的关键特性,连接学术界与工业界,支持神经检索器的高效训练、推理与评估。我们展示了一个统一的密集检索模型,在多语言和多模态任务中表现优异,并开展跨模态零样本研究以验证其潜力。此外,我们发布OmniEmbed——据我们所知,首个统一支持文本、图像文档、视频和音频检索的嵌入模型,为未来研究提供基准。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have driven interest in billion-scale retrieval models with strong generalization across retrieval tasks and languages. Additionally, progress in large vision-language models has created new opportunities for multimodal retrieval. In response, we have updated the Tevatron toolkit, introducing a unified pipeline that enables researchers to explore retriever models at different scales, across multiple languages, and with various modalities. This demo paper highlights the toolkit's key features, bridging academia and industry by supporting efficient training, inference, and evaluation of neural retrievers. We showcase a unified dense retriever achieving strong multilingual and multimodal effectiveness, and conduct a cross-modality zero-shot study to demonstrate its research potential. Alongside, we release OmniEmbed, to the best of our knowledge, the first embedding model that unifies text, image document, video, and audio retrieval, serving as a baseline for future research.

检索系统多模态多语言嵌入模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。