arXiv:2508.11313cs.CV2025-08IJCAI

通过去噪视频片段提升文本定位准确率

Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval

  • 先用文本条件去噪过滤无关视频片段
  • 在两个数据集上均超越现有最佳方法
  • 适合想提升视频定位精度的研究者

当前文本驱动的视频片段定位(VMR)方法会编码所有视频片段,包括无关内容,破坏多模态对齐并阻碍优化。为此,我们提出「去噪-再检索」范式,显式过滤与文本无关的视频片段,并利用净化后的多模态表示进行精准定位。基于此,我们设计了去噪-再检索网络(DRNet),包含文本条件去噪(TCD)和文本重建反馈(TRF)模块。TCD结合交叉注意力与结构化状态空间块,动态识别噪声片段并生成噪声掩码以净化多模态视频表示。TRF从净化后的视频表示中提炼单一查询嵌入,并与文本嵌入对齐,作为训练时去噪的辅助监督信号。最终,使用文本嵌入在净化后的视频表示上执行条件检索,实现精确的视频片段定位。在Charades-STA和QVHighlights数据集上的实验表明,该方法在所有指标上均优于现有最先进方法。此外,该去噪-再检索范式具有可扩展性,可无缝集成至先进VMR模型以进一步提升性能。

原文摘要 · Abstract (English)

Current text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly filters text-irrelevant clips from videos and then retrieves the target moment using purified multimodal representations. Following this paradigm, we introduce the Denoise-then-Retrieve Network (DRNet), comprising Text-Conditioned Denoising (TCD) and Text-Reconstruction Feedback (TRF) modules. TCD integrates cross-attention and structured state space blocks to dynamically identify noisy clips and produce a noise mask to purify multimodal video representations. TRF further distills a single query embedding from purified video representations and aligns it with the text embedding, serving as auxiliary supervision for denoising during training. Finally, we perform conditional retrieval using text embeddings on purified video representations for accurate VMR. Experiments on Charades-STA and QVHighlights demonstrate that our approach surpasses state-of-the-art methods on all metrics. Furthermore, our denoise-then-retrieve paradigm is adaptable and can be seamlessly integrated into advanced VMR models to boost performance.

视频定位去噪多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。