arXiv:2603.24934cs.LGcs.AI2026-03中稿 · CVPR被引 3

提升视频时间定位准确率,让模型更专注关键内容。

CVA: Context-aware Video-text Alignment for Video Temporal Grounding

  • 用语义无关片段增强数据,避免背景干扰导致误判。
  • 在时间边界处强化语义一致性,提升对干扰的鲁棒性。
  • 适合需要精准定位视频片段的研究与应用。

我们提出上下文感知视频-文本对齐(CVA),解决视频时间定位中因无关背景干扰导致对齐不准确的问题。框架包含三个核心组件:首先,提出查询感知上下文多样化(QCD)数据增强策略,仅混合语义无关内容,通过基于相似度的替换片段池模拟多样背景,避免查询无关混合引发的‘假阴性’;其次,引入上下文不变边界判别(CBD)损失,一种对比损失,在复杂时间边界处强制保持语义一致,使表征对上下文变化和难负例更鲁棒;第三,设计上下文增强型变压器编码器(CTE),采用分窗自注意力与双向交叉注意力结合可学习查询的分层结构,捕捉多尺度时间上下文。三者协同作用下,CVA在主流视频时间定位基准(QVHighlights、Charades-STA)上达到当前最优性能,尤其在Recall@1(R1)指标上相比现有方法提升约5个百分点,显著缓解了假阴性问题。

原文摘要 · Abstract (English)

We propose Context-aware Video-text Alignment (CVA), a novel framework to address a significant challenge in video temporal grounding: achieving temporally sensitive video-text alignment that remains robust to irrelevant background context. Our framework is built on three key components. First, we propose Query-aware Context Diversification (QCD), a new data augmentation strategy that ensures only semantically unrelated content is mixed in. It builds a video-text similarity-based pool of replacement clips to simulate diverse contexts while preventing the ``false negative" caused by query-agnostic mixing. Second, we introduce the Context-invariant Boundary Discrimination (CBD) loss, a contrastive loss that enforces semantic consistency at challenging temporal boundaries, making their representations robust to contextual shifts and hard negatives. Third, we introduce the Context-enhanced Transformer Encoder (CTE), a hierarchical architecture that combines windowed self-attention and bidirectional cross-attention with learnable queries to capture multi-scale temporal context. Through the synergy of these data-centric and architectural enhancements, CVA achieves state-of-the-art performance on major VTG benchmarks, including QVHighlights and Charades-STA. Notably, our method achieves a significant improvement of approximately 5 points in Recall@1 (R1) scores over state-of-the-art methods, highlighting its effectiveness in mitigating false negatives.

视频定位多模态时序对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。