arXiv:2503.05186cs.CV2025-03CVPR被引 18

利用视频帧级描述提升文本-视频检索效果

Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions

  • 通过多模态交互增强视频特征表示
  • 查询感知过滤机制抑制错误信息干扰
  • 双模态匹配与难例损失提升检索精度

近期文本-视频检索研究中,使用视觉语言模型生成的附加描述已展现出显著性能提升。然而,现有方法在捕捉视频中丰富的语义信息(如时间变化)方面仍存在不足,且生成模型引入的错误信息可能导致检索不准。为此,我们提出新框架Narrating the Video(NarVid),充分挖掘帧级描述中的综合信息。NarVid从多个角度利用叙述内容:1)通过叙述与视频的跨模态交互增强特征;2)采用查询感知自适应过滤机制抑制无关或错误信息;3)引入查询-视频与查询-叙述双重相似度匹配得分;4)基于双视角相似度设计难例损失,从多角度学习判别性特征。实验表明,NarVid在多个基准数据集上达到当前最优性能。

原文摘要 · Abstract (English)

In recent text-video retrieval, the use of additional captions from vision-language models has shown promising effects on the performance. However, existing models using additional captions often have struggled to capture the rich semantics, including temporal changes, inherent in the video. In addition, incorrect information caused by generative models can lead to inaccurate retrieval. To address these issues, we propose a new framework, Narrating the Video (NarVid), which strategically leverages the comprehensive information available from frame-level captions, the narration. The proposed NarVid exploits narration in multiple ways: 1) feature enhancement through cross-modal interactions between narration and video, 2) query-aware adaptive filtering to suppress irrelevant or incorrect information, 3) dual-modal matching score by adding query-video similarity and query-narration similarity, and 4) hard-negative loss to learn discriminative features from multiple perspectives using the two similarities from different views. Experimental results demonstrate that NarVid achieves state-of-the-art performance on various benchmark datasets.

文本视频检索多模态学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。