arXiv:2606.01149cs.CV2026-06被引 1

提出CoSTL框架,同时提升视频片段定位与高光检测精度

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection

论文配图:CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection
图 1 · 摘自论文原文
  • 文本驱动分步提取帧内细节特征,增强细粒度理解
  • 多尺度时序感知模块捕捉全局动态,提升定位准确性
  • 在4个公开数据集上达领先效果,适合视频分析任务

视频片段定位(MR)和高光检测(HD)是视频分析中的关键任务,旨在根据文本查询定位特定时刻并评估片段相关性。现有方法将两者视为相似的视频定位任务,采用相同架构,主要依赖帧级特征进行时序建模,忽视了单帧中与文本查询相关的丰富视觉信息,导致定位不准。为此,我们提出综合时空表征学习框架CoSTL,同时捕捉细粒度图像级信息和时序动态。具体地,CoSTL引入文本驱动的渐进式细粒度图像编码器,通过两步文本引导的知识提取过程学习空间细节表征;此外,多尺度时序感知模块构建全面的时空表征,增强模型对时序变化的处理能力。我们在四个公开基准测试集(QVHighlights、Charades-STA、TACoS、TVSum)上实现最先进的性能。

原文摘要 · Abstract (English)

Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given text query. Recent approaches treat them as similar video grounding tasks and use the same architecture to solve them. These tasks require both fine-grained comprehension at the image level and high-level temporal understanding across the entire video. Existing approaches have primarily focused on temporal modeling using frame-level features, often neglecting the rich visual information related to the text query within individual frames. This oversight leads to inaccurate grounding results. To address this limitation, we propose a Comprehensive Spatial-Temporal Representation Learning Framework (CoSTL), which captures both fine-grained image-level information and temporal dynamics. Specifically, CoSTL incorporates a text-driven progressive fine-grained image encoder, performing a two-step text-driven knowledge extraction process to learn fine-grained spatial representations. Furthermore, a multi-scale temporal perception module captures comprehensive spatial-temporal representations, enhancing the model's ability to process temporal dynamics. We demonstrate state-of-the-art performance on four public benchmarks: QVHighlights, Charades-STA, TACoS, and TVSum.

视频理解时空建模细粒度定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。