通过多提议协作与多任务训练,提升弱监督视频片段检索精度。
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval

- 生成多个候选片段并用可学习高斯掩码融合,生成高质量正样本。
- 引入正向与逆向掩码查询重构任务,显著提升模型稳定性和召回率。
- 适合缺乏时序标注但需精准定位视频片段的应用场景。
本研究聚焦于弱监督视频片段检索(VMR),旨在仅使用视频级对应关系而无需时间标注的情况下,从非剪辑视频中找出与给定查询语义相似的片段。以往方法或对视频所有实例预测结果进行聚合,或通过为查询提出重构来间接解决任务,但常产生低质量的时间提议,难以区分同一视频中错位的片段,且因依赖单一辅助任务而缺乏稳定性。为此,我们提出一种新方法——多提议协作与多任务训练(MCMT)。首先生成多个提议,并从中推导出对应的可学习高斯掩码;这些掩码被融合以生成高质量正样本掩码,突出与查询最相关的视频片段。同时,将同视频中的其他片段分类为易负样本,整个视频作为难负样本。训练过程中引入前向和反向掩码查询重构任务,对网络施加更强约束,从而提升检索性能的鲁棒性与稳定性。在两个标准基准上的大量实验验证了该方法的有效性。
原文摘要 · Abstract (English)
This study focuses on weakly-supervised Video Moment Retrieval (VMR), aiming to identify a moment semantically similar to the given query within an untrimmed video using only video-level correspondences, without relying on temporal annotations during training. Previous methods either aggregate predictions for all instances in the video, or indirectly address the task by proposing reconstructions for the query. However, these methods often produce low-quality temporal proposals, struggle with distinguishing misaligned moments in the same video, or lack stability due to a reliance on a single auxiliary task. To address these limitations, we present a novel weakly-supervised method called Multi-proposal Collaboration and Multi-task Training (MCMT). Initially, we generate multiple proposals and derive corresponding learnable Gaussian masks from them. These masks are then combined to create a high-quality positive sample mask, highlighting video clips most relevant to the query. Concurrently, we classify other clips in the same video as the easy negative sample and the entire video as the hard negative sample. During training, we introduce forward and inverse masked query reconstruction tasks to impose more substantial constraints on the network, promoting more robust and stable retrieval performance. Extensive experiments on two standard benchmarks affirm the effectiveness of the proposed method in VMR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。