arXiv:2605.21931cs.CV2026-05

让视频大模型通过时间感知自进化,无需人工标注即可提升推理能力。

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models

论文配图:EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
图 1 · 摘自论文原文
  • 设计时间敏感的提问者与解题者奖励机制,驱动模型从原始视频中自主学习。
  • 在6个基准上超越基线模型,性能接近有监督方法。
  • 适合关注视频理解、自进化系统的研究者和开发者。

近期视频大语言模型(Video-LLMs)通过强化学习展现出强大的视频推理能力,但现有强化学习流程严重依赖人工标注的任务与解答,难以扩展且受限于人类知识。自进化框架虽作为替代方案出现,但主要针对文本和图像等静态模态,无法捕捉视频推理中的核心时间动态。本文提出EvoVid,一种以时间为中心的自进化框架,使视频大模型能直接从原始未标注视频中进行自我提升。具体而言,引入两种互补的时间感知奖励:时间敏感的提问者奖励,通过时间扰动敏感性促进时序依赖问题生成;基于时间定位的解题者奖励,利用视频片段的内在定位实现自动时间监督。在四个基础模型和六个基准上的实验表明,该方法持续优于基线模型及现有自进化框架,性能可媲美有监督方法。结果证明,时间中心的自进化是视频理解与推理的有效且可扩展的新范式。

原文摘要 · Abstract (English)

Recent Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in video reasoning through reinforcement learning (RL). However, existing RL pipelines rely heavily on human-annotated tasks and solutions, making them costly to scale and fundamentally constrained by human expertise. Self-evolving frameworks have recently emerged as a promising alternative through autonomous Questioner-Solver self-play. Unfortunately, these approaches are primarily designed for static modalities such as text and images, fundamentally failing to capture the temporal dynamics that are central to video reasoning. In this work, we propose $\textbf{EvoVid}$, a temporal-centric self-evolving framework that enables Video-LLMs to improve directly from raw, unannotated videos. Specifically, we introduce two complementary temporal-centric rewards: a temporal-aware Questioner reward that encourages temporally dependent question generation through temporal perturbation sensitivity, and a temporal-grounded Solver reward that provides automatic temporal supervision via inherent video segment localization. Extensive experiments across four base models and six benchmarks demonstrate consistent improvements over both base models and existing self-evolving baselines, achieving competitive performance with supervised methods. These results highlight temporal-centric self-evolution as an effective and scalable paradigm for video understanding and reasoning.

视频理解自进化强化学习时间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。