arXiv:2602.20159cs.CVcs.AI2026-02被引 23

构建超大规模视频推理数据集,推动视频模型从画质到逻辑理解的升级。

A Very Big Video Reasoning Suite

  • 构建200个任务、百万级视频的标准化推理数据集
  • 发现模型在未见任务上出现早期泛化迹象
  • 提供可复现的人类对齐评估框架,适合研究者验证模型推理能力

视频模型发展长期聚焦视觉质量,忽视了推理能力的探索。视频推理依托时空一致的视觉环境,能自然表达连续性、交互与因果等结构化智能,但系统研究其扩展规律受限于缺乏大规模训练数据。为此,本文提出前所未有的超大规模视频推理数据集VBVR,涵盖200个经体系化分类的推理任务,包含超过一百万段视频片段,规模较现有数据集高出约三个数量级。我们进一步构建了可验证的评估基准VBVR-Bench,通过规则驱动与人类对齐评分机制,摆脱依赖模型打分的局限,实现可复现且可解释的推理能力诊断。基于该套件,我们开展了首个大规模视频推理扩展研究,观察到模型对未见任务存在早期涌现式泛化现象。整体上,VBVR为通用视频推理研究奠定了基础。数据、基准工具包及模型已公开:https://video-reason.com/?v=vbvr。

原文摘要 · Abstract (English)

Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure such as continuity, interaction, and causality. However, systematically studying video reasoning and its scaling behavior is hindered by the lack of large-scale training data. To address this gap, we introduce the Very Big Video Reasoning (VBVR) Dataset, an unprecedentedly large-scale resource spanning 200 curated reasoning tasks following a principled taxonomy and over one million video clips, approximately three orders of magnitude larger than existing datasets. We further present VBVR-Bench, a verifiable evaluation framework that moves beyond model-based judging by incorporating rule-based, human-aligned scorers, enabling reproducible and interpretable diagnosis of video reasoning capabilities. Leveraging the VBVR suite, we conduct one of the first large-scale scaling studies of video reasoning and observe early signs of emergent generalization to unseen reasoning tasks. Together, VBVR lays a foundation for the next stage of research in generalizable video reasoning. The data, benchmark toolkit, and models are publicly available at https://video-reason.com/?v=vbvr .

视频推理大规模数据模型泛化评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。