通过视频内自相似性提升零样本视频片段检索精度
Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval

- 利用视频自身内容的自相似性生成候选片段,避开查询与视频间的模态和语言风格差异
- 在多个ZMR基准上达到当前最优性能,显著提升检索稳定性
- 适合需要跨模态对齐且无标注数据的视频理解场景
零样本视频片段检索旨在克服传统方法依赖大规模带时序标注数据的局限。尽管预训练视觉-语言模型和多模态大语言模型取得进展,现有方法仍严重依赖查询与视频内容的相似性,易受模态和语言风格差异影响,导致候选片段不可靠、检索结果不稳定。为此,我们提出基于自相似性的候选片段生成与评分方法(Self-SiMS),仅从视频内容中提取内在相似关系,避免查询-帧或查询-标题相似性带来的噪声和错配,有效缓解了模态与语言风格差距。此外,引入查询感知的多模态大语言模型推理阶段,进一步增强文本与视频的对齐。大量实验表明,Self-SiMS在多个零样本视频片段检索基准上均达到领先性能。
原文摘要 · Abstract (English)
Zero-shot video moment retrieval aims to overcome the limitations of traditional approaches that require large-scale datasets annotated with text and its relevant temporal spans. Despite advances in pre-trained vision-language models and multimodal large language models, existing ZMR methods still heavily depend on query-to-video content similarity, making them vulnerable to modality and language-style gaps. These gaps lead to unreliable span proposals and unstable moment retrieval results. To address this issue, we propose Self-Similarity-based Moment Proposal and Scoring that instead exploits intrinsic relationships within videos, enabling robust span generation and scoring. By deriving self-similarity only from the video content, we circumvent the noisy and mismatched patterns of query-frame or query-caption similarities, thereby mitigating both modality and language-style gaps. Furthermore, we introduce a query-aware MLLM-based reasoning stage to further sharpen alignment between text and video. Extensive experiments demonstrate that Self-SiMS achieves state-of-the-art performance across ZMR benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。