多模态多语言视频检索系统,融合视觉音频文本信息提升搜索精准度。
MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion

- 通过多模态特征提取与加权倒数排名融合,平衡视觉、音频、文本信号。
- 在MultiVENT 2.0和TVR上相较顶尖模型提升nDCG@20达81%。
- 适合需要跨模态精准检索的场景,如学术视频或复杂信息查询。
视频天然包含多种模态:视觉事件、文字叠加、声音与语音,均对检索至关重要。然而,当前主流多模态语言模型(如VAST、LanguageBind)基于视觉-语言模型构建,过度依赖视觉信号。检索基准也强化此偏见,仅关注视觉查询而忽略其他模态。我们提出搜索系统MMMORRF,从视觉与音频模态中提取文本与特征,并采用新颖的模态感知加权倒数排名融合策略。该系统兼具高效性与有效性,在满足用户真实信息需求而非仅视觉描述的视频检索任务中表现优异。我们在MultiVENT 2.0和TVR两个面向特定信息需求的多模态基准上评估,结果显示其nDCG@20相比领先多模态编码器提升81%,相比单模态检索提升37%,充分证明融合多元模态的价值。
原文摘要 · Abstract (English)
Videos inherently contain multiple modalities, including visual events, text overlays, sounds, and speech, all of which are important for retrieval. However, state-of-the-art multimodal language models like VAST and LanguageBind are built on vision-language models (VLMs), and thus overly prioritize visual signals. Retrieval benchmarks further reinforce this bias by focusing on visual queries and neglecting other modalities. We create a search system MMMORRF that extracts text and features from both visual and audio modalities and integrates them with a novel modality-aware weighted reciprocal rank fusion. MMMORRF is both effective and efficient, demonstrating practicality in searching videos based on users' information needs instead of visual descriptive queries. We evaluate MMMORRF on MultiVENT 2.0 and TVR, two multimodal benchmarks designed for more targeted information needs, and find that it improves nDCG@20 by 81% over leading multimodal encoders and 37% over single-modality retrieval, demonstrating the value of integrating diverse modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。