arXiv:2602.07801cs.CVcs.AI2026-02被引 4

提出统一框架,让模型精准定位视频关键段并回答问题。

VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos

  • 统一建模视频定位与问答,支持动态剪辑
  • 在长视频上准确率提升,减少幻觉生成
  • 适合需要精确定位的视频理解任务

在长视频理解中,传统均匀采样难以捕捉关键视觉证据,导致性能下降和幻觉增多。为此,近期基于“思考-视频”范式的模型采用定位-裁剪-回答流程,主动识别相关视频片段,在其内进行密集采样并生成答案。然而现有方法效率低、定位能力弱,且流程僵化。本文提出 VideoTemp-o3,一个统一的“思考-视频”框架,联合建模视频定位与问答任务。该框架具备强定位能力,支持按需剪辑,并可修正错误定位。在监督微调阶段,设计统一掩码机制以促进探索并抑制噪声;在强化学习阶段,引入专用奖励机制防止奖励劫持。此外,构建了高质量长视频定位问答数据集及对应基准,用于系统评估不同视频时长下的表现。实验表明,本方法在长视频理解和定位任务上均取得显著性能提升。

原文摘要 · Abstract (English)

In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To address this, recent agentic thinking-with-videos paradigms have emerged, adopting a localize-clip-answer pipeline in which the model actively identifies relevant video segments, performs dense sampling within those clips, and then produces answers. However, existing methods remain inefficient, suffer from weak localization, and adhere to rigid workflows. To solve these issues, we propose VideoTemp-o3, a unified agentic thinking-with-videos framework that jointly models video grounding and question answering. VideoTemp-o3 exhibits strong localization capability, supports on-demand clipping, and can refine inaccurate localizations. Specifically, in the supervised fine-tuning stage, we design a unified masking mechanism that encourages exploration while preventing noise. For reinforcement learning, we introduce dedicated rewards to mitigate reward hacking. Besides, from the data perspective, we develop an effective pipeline to construct high-quality long video grounded QA data, along with a corresponding benchmark for systematic evaluation across various video durations. Experimental results demonstrate that our method achieves remarkable performance on both long video understanding and grounding.

视频理解定位智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。