arXiv:2607.05089cs.CV2026-07

让视频大模型学会分步找时间证据,提升推理准确性。

TimeThink: Reasoning with Time for Video LLMs

论文配图:TimeThink: Reasoning with Time for Video LLMs
图 1 · 摘自论文原文
  • 将时间线索作为推理基本单元,分步奖励定位过程。
  • 在多个基准上实现开源视频强化学习模型最佳表现。
  • 适合需要精准时序理解的视频分析任务研究者。

视频推理要求模型在长视频序列中识别并验证时间定位的证据。近期视频大语言模型(Video-LLMs)通过与强化学习对齐展现出良好推理能力,但现有方法通常依赖仅监督最终预测结果的回报机制,难以指导模型在中间推理阶段发现相关的时间证据。本文提出TimeThink,一种显式引导视频大模型发现时间证据的强化学习框架。核心思想是将每一步推理对应一个候选时间区间,作为视频推理的基本优化单元。引入逐步时间过程奖励,对时间线索提供局部化信用分配,并设计联合过程-结果优化目标,平衡推理忠实度与任务正确性。为支持可扩展训练,构建了TimeThink-RFT-20K数据集,包含自动标注的时间证据片段。在视频推理、时间定位和通用视频理解等多个基准上的实验表明,TimeThink持续提升时间定位与推理性能,在开源视频强化学习模型中达到领先水平。

原文摘要 · Abstract (English)

Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Video-LLMs) have shown promising reasoning abilities when aligned with reinforcement learning, yet existing approaches typically rely on outcome-based rewards that supervise only the final prediction. Such supervision provides limited guidance on how models should discover the relevant temporal evidence during intermediate reasoning. In this work, we propose TimeThink, a reinforcement learning framework that explicitly guides temporal evidence discovery in Video-LLMs. Our key idea is to treat temporal clue steps as the fundamental optimization primitive of video reasoning, where each reasoning step references a candidate time interval in the video. We introduce a step-wise temporal process reward that provides localized credit assignment for these clues and a joint process--outcome optimization objective that balances reasoning fidelity with task correctness. To enable scalable training, we construct TimeThink-RFT-20K, a dataset with automatically derived temporal evidence segments. Extensive experiments across video reasoning, temporal grounding, and general video understanding benchmarks show that TimeThink consistently improves both temporal localization and reasoning performance, achieving state-of-the-art results among open-source video RL models.

视频推理强化学习时间定位大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。