arXiv:2506.09409cs.IR2025-06

用统一嵌入模型实现跨模态视频检索,效果领先。

MAGMaR Shared Task System Description: Video Retrieval with OmniEmbed

  • 用OmniEmbed统一建模文本、图像、音频和视频特征
  • 在MultiVENT 2.0上实现多语言视频检索最优性能
  • 适合需要跨模态检索的开发者与研究者

由于视觉、听觉和文本模态的复杂融合,有效的视频检索仍具挑战。本文在MAGMaR共享任务中探索使用Tevatron 2.0工具包中的OmniEmbed多模态嵌入模型进行统一检索。该模型在包含视觉帧、音频轨道和文本描述的MultiVENT 2.0数据集上生成统一嵌入,支持鲁棒的多模态检索。通过联合微调多模态数据,我们在复杂多语言视频检索任务中取得显著提升。本方案在2025年5月20日时,为公开提交中最高分,验证了统一多模态检索方法的实用性。模型检查点已开源。

原文摘要 · Abstract (English)

Effective video retrieval remains challenging due to the complexity of integrating visual, auditory, and textual modalities. In this paper, we explore unified retrieval methods using OmniEmbed, a powerful multimodal embedding model from the Tevatron 2.0 toolkit, in the context of the MAGMaR shared task. Evaluated on the comprehensive MultiVENT 2.0 dataset, OmniEmbed generates unified embeddings for text, images, audio, and video, enabling robust multimodal retrieval. By finetuning OmniEmbed with the combined multimodal data--visual frames, audio tracks, and textual descriptions provided in MultiVENT 2.0, we achieve substantial improvements in complex, multilingual video retrieval tasks. Our submission achieved the highest score on the MAGMaR shared task leaderboard among public submissions as of May 20th, 2025, highlighting the practical effectiveness of our unified multimodal retrieval approach. Model checkpoint in this work is opensourced.

视频检索多模态嵌入模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。