提出鲁棒对齐学习框架,提升不完整视频检索的准确性。
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
- 将查询和视频建模为多元高斯分布,量化数据不确定性
- 通过可学习置信门动态加权关键词,缓解无关内容干扰
- 适配多种模型架构,显著提升部分相关视频检索效果
部分相关视频检索(PRVR)旨在从非裁剪视频中找出与查询部分相关的片段。核心挑战在于应对由数据固有不确定性引发的虚假语义关联:一是查询模糊性,即查询无法完整描述目标视频,常包含无信息量的词元;二是视频部分相关性,大量与查询无关的片段在跨模态对齐中引入上下文噪声。现有方法多关注多尺度片段表示增强和最相关片段检索,但难以抵御具有虚假相似性的干扰视频,导致性能不佳。为此,本文提出鲁棒对齐学习(RAL)框架,显式建模数据不确定性。关键创新包括:1)首次在PRVR中采用概率建模,将视频和查询编码为多元高斯分布,不仅量化不确定性,还支持代理级别匹配以捕捉跨模态对应关系的变异性;2)考虑查询词项的信息异质性,引入可学习置信门动态调整相似度权重。作为即插即用方案,RAL可无缝集成至现有架构。在多种检索骨干网络上的广泛实验验证了其有效性。
原文摘要 · Abstract (English)
Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos partially relevant to a given query. The core challenge lies in learning robust query-video alignment against spurious semantic correlations arising from inherent data uncertainty: 1) query ambiguity, where the query incompletely characterizes the target video and often contains uninformative tokens, and 2) partial video relevance, where abundant query-irrelevant segments introduce contextual noise in cross-modal alignment. Existing methods often focus on enhancing multi-scale clip representations and retrieving the most relevant clip. However, the inherent data uncertainty in PRVR renders them vulnerable to distractor videos with spurious similarities, leading to suboptimal performance. To fill this research gap, we propose Robust Alignment Learning (RAL) framework, which explicitly models the uncertainty in data. Key innovations include: 1) we pioneer probabilistic modeling for PRVR by encoding videos and queries as multivariate Gaussian distributions. This not only quantifies data uncertainty but also enables proxy-level matching to capture the variability in cross-modal correspondences; 2) we consider the heterogeneous informativeness of query words and introduce learnable confidence gates to dynamically weight similarity. As a plug-and-play solution, RAL can be seamlessly integrated into the existing architectures. Extensive experiments across diverse retrieval backbones demonstrate its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。