arXiv:2608.02078cs.CLcs.CV2026-08

解决视频定位中视觉证据与时间戳错位问题,提升定位精度。

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding

论文配图:CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding
图 1 · 摘自论文原文
  • 引入边界特异性视觉证据标记,显式建模关键帧特征。
  • 在多个公开数据集上,平均精度提升3.2%以上。
  • 适合关注视频定位细节与模型可解释性的研究者。

大型视觉语言模型(LVLM)通过强化学习(RL)在视频时间定位(VTG)任务中取得了显著进展。然而,现有方法主要依赖最终预测区间的正确性奖励,未能充分约束边界相关的视觉证据与时间戳预测之间的对齐关系。本文深入分析时间戳预测与底层边界级视觉证据的关系,发现主流基准中普遍存在视觉证据与预测时间戳不匹配的现象。为此,我们提出一种能力感知的视觉边界证据对齐方法(CAVE),通过边界特异性视觉证据奖励增强定位优化,缓解证据与时间戳的错位问题。具体地,CAVE引入边界特异性证据标记,并通过轻量级监督预热初始化其结构化生成与不同边界语义。在强化学习阶段,视觉边界证据对齐奖励促使模型在真实边界内关注特定证据标记,从而促进视觉证据与时间边界的对齐。此外,设计了性能感知门控机制,对定位不佳的样本保留证据指导,定位准确后自动降低监督强度,避免过度约束细粒度边界优化。在多个公开的VTG基准上的实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have achieved substantial performance gains in Video Temporal Grounding (VTG) through reinforcement learning (RL). However, existing methods primarily rely on outcome correctness rewards that evaluate only the final predicted intervals, leaving boundary-related visual evidence and its correspondence with timestamp predictions insufficiently constrained. In this paper, we delve into timestamp prediction and its underlying boundary-level visual evidence, showing prevalent misalignment between visual evidence and predicted timestamps across widely used benchmarks. To address this issue, we propose Competence-Aware Visual Boundary Evidence Alignment (CAVE), which augments localization optimization with boundary-specific visual evidence rewards to mitigate evidence-timestamp misalignment. Specifically, to explicitly represent the boundary-specific visual evidence, CAVE introduces boundary-specific evidence tokens and initializes their structured generation and distinct boundary semantics through a lightweight supervised warm-up. During RL, the visual boundary evidence alignment reward reinforces the visual attention of special evidence tokens within the ground-truth boundaries, thereby promoting alignment between visual evidence and temporal boundaries. Moreover, performance-aware gating for evidence supervision is designed to adaptively retain evidence guidance for poorly localized groups while reducing it once localization becomes sufficiently accurate to avoid over-constraining fine-grained boundary refinement. Extensive experiments on several public VTG benchmarks demonstrate the effectiveness of our method.

视频定位视觉证据强化学习边界对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。