arXiv:2510.20579cs.CVcs.AI2025-10被引 44

让视频推理过程可追踪,明确标出关键证据出现的时间和位置。

Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence

  • 通过标注时间戳、物体和边界框,显式定位视频中的关键证据。
  • 在V-STAR上比基线模型提升14.4%的mAM和24.2%的mLGM。
  • 适合需要可解释性和高可信度视频理解的场景。

多数视频推理模型仅生成文本推理过程,未标明关键证据出现的时间与位置。尽管OpenAI-o3等模型在图像证据中心推理方面引发关注,但将该能力扩展至视频更具挑战性,需同时处理动态场景中的时空跟踪与定位。本文提出Open-o3-Video,一种非代理框架,通过高亮关键时间点、物体及边界框,将显式的时空证据融入视频推理,使推理过程可追溯且可验证。为此,我们构建了高质量数据集STGR,提供统一的时空监督,填补现有资源空白。进一步采用冷启动强化学习策略,设计特定奖励函数,联合优化答案准确性、时间对齐度与空间精度。在V-STAR基准上,Open-o3-Video实现当前最优性能,相较Qwen2.5-VL基线提升mAM 14.4%、mLGM 24.2%,并在多个视频理解基准上表现一致领先。除准确率外,其生成的可溯源推理轨迹支持置信度感知的测试时扩展,提升答案可靠性。

原文摘要 · Abstract (English)

Most video reasoning models only generate textual reasoning traces without indicating when and where key evidence appears. Recent models such as OpenAI-o3 have sparked wide interest in evidence-centered reasoning for images, yet extending this ability to videos is more challenging due to the need for joint temporal tracking and spatial localization across dynamic scenes. We introduce Open-o3-Video, a non-agent framework that integrates explicit spatio-temporal evidence into video reasoning by highlighting key timestamps, objects, and bounding boxes, making the reasoning process traceable and verifiable. To enable this capability, we first construct high-quality datasets STGR that provide unified spatio-temporal supervision, which is absent in existing resources. We further adopt a cold-start reinforcement learning strategy with specially designed rewards that jointly encourage answer accuracy, temporal alignment, and spatial precision. On the V-STAR benchmark, Open-o3-Video achieves state-of-the-art performance, improving mAM by 14.4% and mLGM by 24.2% over the Qwen2.5-VL baseline, and shows consistent gains across a range of video understanding benchmarks. Beyond accuracy, the grounded reasoning traces produced by Open-o3-Video support confidence-aware test-time scaling, improving answer reliability.

视频推理可解释性时空定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。