让视频模型自己发现关键片段并提问,提升理解能力。
Video-Zero: Self-Evolution Video Understanding

- 用问答协同进化机制,聚焦视频中关键时间片段。
- 在13个基准上超越多个视频大模型,提升理解准确性。
- 无需人工标注,适合长视频与复杂推理任务研究者。
自进化为减少人工标注依赖的推理模型改进提供了新路径,但将其拓展至视频理解仍面临挑战:视频内容冗长动态,推理所需证据稀疏且时间局部化。直接从完整视频生成难题问答会带来看似困难却弱关联的监督信号,依赖静态线索或语言先验而非时间证据。本文指出,视频自进化核心瓶颈并非难度,而是证据的准确对齐。为此提出Video-Zero——一种无标注的问答器-求解器协同进化框架,聚焦于时间局部化的证据。问题生成器发现有信息量的视频片段并生成基于证据的问题,求解器学习依据支撑证据进行回答和预测对齐。该过程形成证据发现、强监督与对齐学习的闭环迭代。在涵盖时序定位、长视频理解与视频推理的13个基准上,Video-Zero持续提升多个视频视觉语言模型骨干网络性能,验证了以证据为中心的自进化方法的有效性与可迁移性。
原文摘要 · Abstract (English)
Self-evolution offers a promising path for improving reasoning models without relying on intensive human annotation. However, extending this paradigm to video understanding remains underexplored and challenging: videos are long, dynamic, and redundant, while the evidence needed for reasoning is often sparse and temporally localized. Naively generating difficult question-answer pairs from full videos can therefore produce supervision that appears challenging but is weakly grounded, relying on static cues or language priors rather than temporal evidence. In this work, we argue that the key bottleneck of video self-evolution is not difficulty alone, but grounding. We propose Video-Zero, an annotation-free Questioner--Solver co-evolution framework that centers self-evolution on temporally localized evidence. The Questioner discovers informative evidence segments and generates evidence-grounded questions, while the Solver learns to answer and align its predictions with the supporting evidence. This closes an iterative loop of evidence discovery, grounded supervision, and evidence-aligned learning. Across 13 benchmarks spanning temporal grounding, long-video understanding, and video reasoning, Video-Zero consistently improves multiple video VLM backbones, demonstrating the effectiveness and transferability of evidence-centered self-evolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。