arXiv:2508.07683cs.CVcs.AI2025-08中稿 · ECCV被引 5

用时间锚点约束推理,让视频定位更准更可信。

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding

  • 引入时间锚点机制,强制中间推理紧扣视觉内容
  • 在VTG任务上达到当前最优性能,推理过程更忠实
  • 仅用7B模型自动生成高质量推理数据,无需大模型

视频时间定位(VTG)旨在将自然语言查询与视频中的特定片段对应。现有大型视觉-语言模型虽采用强化学习生成思维链(CoT),但仅依赖结果监督,常导致推理脱离视觉内容并产生幻觉。现有缓解方法依赖外部大模型或奖励模型,计算成本高且模式僵化。为此,我们提出TAR(时间锚点约束推理)框架,引入时间锚点(T-anchor)作为透明可审计的检查点机制,强制思维链逐步细化,持续以视觉证据为依据校准时间预测,显著提升推理过程的忠实性与自主性。此外,我们设计一种自举范式,仅用标准7B模型自动获取高质量思维链数据,摆脱对超大规模模型的依赖。大量实验表明,TAR在保持最先进性能的同时,生成了忠实、自主且逐步优化的推理轨迹。

原文摘要 · Abstract (English)

Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries. While recent Large Vision-Language Models (LVLMs) employ Reinforcement Learning to generate Chains-of-Thought (CoT), they typically rely solely on outcome-based supervision. Consequently, this often leads to hallucinations, where the reasoning process becomes disconnected from the visual content and the final prediction. Existing attempts to mitigate this by relying on external supervision from larger models or separate reward models are computationally expensive and prone to rigid patterns. To address these challenges, we propose TAR (Temporal Anchor-Constrained Reasoning), a framework that introduces the temporal anchor (T-anchor) as a transparent and auditable checkpoint mechanism. T-anchor enforces progressive refinement within the CoT, compelling the model to continuously ground its intermediate thoughts in visual evidence and iteratively calibrate temporal predictions, thereby significantly enhancing the faithfulness and autonomy of the reasoning process and final accuracy. Furthermore, we introduce a bootstrapping paradigm that automatically harvests high-quality CoT data using only a standard 7B model, eliminating the dependency on ultra-large models. Extensive experiments demonstrate that TAR achieves state-of-the-art performance and generates faithful, autonomous, and progressively refined reasoning traces.

视频定位思维链时间锚点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。