arXiv:2505.23952cs.CV2025-05综述被引 2

通过融合视觉、时间与文本辅助信息,提升文本到视频检索准确率。

Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review

  • 引入物体、时空上下文等多模态辅助信息增强匹配
  • 覆盖81篇论文,在多个数据集上实现领先性能
  • 适合研究跨模态检索与多源信息融合的学者

文本到视频(T2V)检索旨在根据用户文本查询从视频库中找出最相关的视频。传统方法仅依赖视频与文本模态对齐计算相似性进行检索。近年来,研究强调从视频和文本中提取辅助信息以提升检索性能,并缩小两者间的语义鸿沟。这些辅助信息包括视觉属性(如物体)、时空上下文以及语音或重述字幕等文本描述。本综述系统回顾了81篇利用此类辅助信息的T2V检索研究,详细分析其方法,总结基准数据集上的最新成果,并讨论现有数据集及其提供的辅助信息。此外,还提出未来研究的潜在方向,聚焦于如何进一步利用这些信息提升检索效果。

原文摘要 · Abstract (English)

Text-to-Video (T2V) retrieval aims to identify the most relevant item from a gallery of videos based on a user's text query. Traditional methods rely solely on aligning video and text modalities to compute the similarity and retrieve relevant items. However, recent advancements emphasise incorporating auxiliary information extracted from video and text modalities to improve retrieval performance and bridge the semantic gap between these modalities. Auxiliary information can include visual attributes, such as objects; temporal and spatial context; and textual descriptions, such as speech and rephrased captions. This survey comprehensively reviews 81 research papers on Text-to-Video retrieval that utilise such auxiliary information. It provides a detailed analysis of their methodologies; highlights state-of-the-art results on benchmark datasets; and discusses available datasets and their auxiliary information. Additionally, it proposes promising directions for future research, focusing on different ways to further enhance retrieval performance using this information.

文本到视频跨模态检索辅助信息综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。