arXiv:2501.02504cs.CVcs.AI2025-01AAAI被引 8

通过视频上下文注意力提升关键词定位准确率

Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection

  • 引入视频上下文聚类与关键词加权机制
  • 在QVHighlights等三组数据集上显著提升性能
  • 适合关注视频片段定位与关键帧检测的研究者

视频片段检索与亮点检测的目标是根据文本查询识别特定视频片段和高光时刻。随着视频内容激增及两类任务的重叠,近期方法尝试同时解决二者。然而,现有方法仍难以充分捕捉整体视频上下文,导致难以判断哪些词最相关。本文提出一种新颖的视频上下文感知关键词注意力模块,通过捕获关键词在整个视频中的变化来克服该局限。为此,我们设计了视频上下文聚类模块,生成整体视频上下文的紧凑表示,增强对关键词动态的理解;同时提出关键词权重检测模块,结合关键词感知对比学习,强化视觉与文本特征间的细粒度对齐。在QVHighlights、TVSum和Charades-STA三个基准上的大量实验表明,所提方法在片段检索与亮点检测任务上显著优于现有方法。代码已公开于:https://github.com/VisualAIKHU/Keyword-DETR

原文摘要 · Abstract (English)

The goal of video moment retrieval and highlight detection is to identify specific segments and highlights based on a given text query. With the rapid growth of video content and the overlap between these tasks, recent works have addressed both simultaneously. However, they still struggle to fully capture the overall video context, making it challenging to determine which words are most relevant. In this paper, we present a novel Video Context-aware Keyword Attention module that overcomes this limitation by capturing keyword variation within the context of the entire video. To achieve this, we introduce a video context clustering module that provides concise representations of the overall video context, thereby enhancing the understanding of keyword dynamics. Furthermore, we propose a keyword weight detection module with keyword-aware contrastive learning that incorporates keyword information to enhance fine-grained alignment between visual and textual features. Extensive experiments on the QVHighlights, TVSum, and Charades-STA benchmarks demonstrate that our proposed method significantly improves performance in moment retrieval and highlight detection tasks compared to existing approaches. Our code is available at: https://github.com/VisualAIKHU/Keyword-DETR

视频理解关键词定位上下文注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。