arXiv:2504.01591cs.CV2025-04中稿 · BMVC 2025被引 1

用视觉语言标签提升视频检索精度

Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval

  • 通过基础模型提取模态特有标签,增强跨模态对齐
  • 在六个数据集上均优于或媲美当前最优方法
  • 适合关注多模态对齐与视频检索的研究者

视频检索需对齐视觉内容与自然语言描述。本文提出一种新方法MAC-VR,利用从基础模型自动提取的模态特有标签来增强视频检索。通过在隐空间中对齐不同模态,并学习与对齐由视频及其对应标题特征衍生出的辅助隐概念,使概念间可区分。该方法提升了视觉与文本隐概念的对齐效果。我们在六个多样化数据集上进行了大量实验:MSR-VTT的两个不同划分、DiDeMo、TGIF、Charades和YouCook2。结果一致表明,模态特有标签能有效提升跨模态对齐,在三个数据集上超越当前最先进方法,其余数据集表现相当或更优。

原文摘要 · Abstract (English)

Video retrieval requires aligning visual content with corresponding natural language descriptions. In this paper, we introduce Modality Auxiliary Concepts for Video Retrieval (MAC-VR), a novel approach that leverages modality-specific tags -- automatically extracted from foundation models -- to enhance video retrieval. We propose to align modalities in a latent space, along with learning and aligning auxiliary latent concepts derived from the features of a video and its corresponding caption. We introduce these auxiliary concepts to improve the alignment of visual and textual latent concepts, allowing concepts to be distinguished from one another. We conduct extensive experiments on six diverse datasets: two different splits of MSR-VTT, DiDeMo, TGIF, Charades and YouCook2. The experimental results consistently demonstrate that modality-specific tags improve cross-modal alignment, outperforming current state-of-the-art methods across three datasets and performing comparably or better across others. Project Webpage: https://adrianofragomeni.github.io/MAC-VR/

视频检索跨模态对齐标签增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。