arXiv:2504.19637cs.CV2025-04被引 1

通过分析视频内冗余与跨样本关联,提升部分相关视频检索效果

Enhanced Partially Relevant Video Retrieval through Inter- and Intra-Sample Analysis with Coherence Prediction

  • 构建伪正样本对增强跨样本语义关联
  • 挖掘冗余片段,区分相关与无关内容,提升表示判别力
  • 通过时序预测强化细粒度语义理解,适合视频检索研究者

部分相关视频检索(PRVR)旨在找到与文本查询部分相关的目标视频。其主要挑战在于文本与视觉模态间的语义不对称性,视频常包含大量与查询无关的内容。现有方法粗略对齐图文对以构建语义空间,忽略了该任务固有的双重特性:跨样本相关性和样本内冗余性。为此,本文提出一种新框架,系统利用这两类特性。首先,交叉相关增强(ICE)模块通过识别语义相似但未配对的文本查询与视频片段,构建伪正样本对,以增强语义空间的鲁棒性;其次,内部冗余挖掘(IRM)模块挖掘冗余视频片段特征,将其与查询相关片段区分开来,促使模型学习更具判别性的表示;最后,引入时序一致性预测(TCP)模块,通过训练模型预测随机打乱视频序列的原始时间顺序,强化细粒度片段级语义的区分能力。大量实验表明,该方法达到当前最优性能。

原文摘要 · Abstract (English)

Partially Relevant Video Retrieval (PRVR) aims to retrieve the target video that is partially relevant to the text query. The primary challenge in PRVR arises from the semantic asymmetry between textual and visual modalities, as videos often contain substantial content irrelevant to the query. Existing methods coarsely align paired videos and text queries to construct the semantic space, neglecting the critical cross-modal dual nature inherent in this task: inter-sample correlation and intra-sample redundancy. To this end, we propose a novel PRVR framework to systematically exploit these two characteristics. Our framework consists of three core modules. First, the Inter Correlation Enhancement (ICE) module captures inter-sample correlation by identifying semantically similar yet unpaired text queries and video moments, combining them to form pseudo-positive pairs for more robust semantic space construction. Second, the Intra Redundancy Mining (IRM) module mitigates intra-sample redundancy by mining redundant moment features and distinguishing them from query-relevant moments, encouraging the model to learn more discriminative representations. Finally, to reinforce these modules, we introduce the Temporal Coherence Prediction (TCP) module, enhancing discrimination of fine-grained moment-level semantics by training the model to predict the original temporal order of randomly shuffled video sequences. Extensive experiments demonstrate the superiority of our method, achieving state-of-the-art results.

视频检索多模态语义对齐时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。