arXiv:2601.00584cs.CV2026-01AAAI被引 6

解决视频片段检索中语义粒度不匹配问题,提升零样本检索精度。

GranAlign: Granularity-Aware Alignment Framework for Zero-Shot Video Moment Retrieval

  • 通过多粒度查询重写和意图感知描述生成实现跨模态对齐。
  • 在三个主流数据集上均达新最佳,QVHighlights提升3.23% mAP@avg。
  • 无需训练,适合希望快速部署的视觉语言研究者。

零样本视频片段检索(ZVMR)是在无特定任务训练数据的情况下,基于自然语言查询定位未剪辑视频中的时间片段。其主要挑战在于文本查询与视觉内容间语义粒度的不一致。现有方法虽利用高质量预训练知识将视频与语言映射到联合空间,但未能平衡各模态在特定场景下的语义粒度,导致表示虽优却仍存在对齐偏差。本文提出一种无需训练的粒度感知对齐框架GranAlign,引入两项互补技术:基于粒度的查询重写以生成不同层次的语义粒度,以及查询感知的标题生成以将查询意图嵌入视频内容。通过将多层级查询与无查询依赖及有查询依赖的标题配对,有效缓解了语义不匹配问题。实验表明,该方法在三大主流基准(QVHighlights、Charades-STA、ActivityNet-Captions)上均达到新最优,尤其在具有挑战性的QVHighlights数据集上,平均准确率(mAP@avg)提升3.23%。

原文摘要 · Abstract (English)

Zero-shot video moment retrieval (ZVMR) is the task of localizing a temporal moment within an untrimmed video using a natural language query without relying on task-specific training data. The primary challenge in this setting lies in the mismatch in semantic granularity between textual queries and visual content. Previous studies in ZVMR have attempted to achieve alignment by leveraging high-quality pre-trained knowledge that represents video and language in a joint space. However, these approaches failed to balance the semantic granularity between the pre-trained knowledge provided by each modality for a given scene. As a result, despite the high quality of each modality's representations, the mismatch in granularity led to inaccurate retrieval. In this paper, we propose a training-free framework, called Granularity-Aware Alignment (GranAlign), that bridges this gap between coarse and fine semantic representations. Our approach introduces two complementary techniques: granularity-based query rewriting to generate varied semantic granularities, and query-aware caption generation to embed query intent into video content. By pairing multi-level queries with both query-agnostic and query-aware captions, we effectively resolve semantic mismatches. As a result, our method sets a new state-of-the-art across all three major benchmarks (QVHighlights, Charades-STA, ActivityNet-Captions), with a notable 3.23% mAP@avg improvement on the challenging QVHighlights dataset.

视频检索零样本学习多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。