arXiv:2608.19583cs.CVcs.AI2026-08被引 1

构建视频生成模型视觉推理能力测评基准,揭示当前模型仍不可靠。

VGI-Bench: Probing Visual Intelligence in Video Generation Models

论文配图:VGI-Bench: Probing Visual Intelligence in Video Generation Models
图 1 · 摘自论文原文
  • 设计双层分类体系的27项任务,覆盖视觉推理多维度
  • 最强模型Seedance 2.0仅达51.0%准确率,表现受限于早期假设
  • 揭示模型对输入敏感、缺乏纠错能力,适合评估下一代视频生成模型

近期研究发现,视频生成模型可通过生成帧表现出一定零样本视觉推理能力。然而可靠评估仍具挑战:基准需契合当前视频模型的视觉先验,要求有效演化过程而非仅合理终态,并校准任务难度以保持挑战性又可完成。为此,我们提出VGI-bench,包含27个任务和810个实例,采用两级分类体系(任务领域与技能标签)实现对视觉推理能力的细粒度评估。评估显示,现有生成系统仅能解决部分视觉基础推理任务,但整体仍不可靠,即使最强模型Seedance 2.0在本评估标准下也仅达51.0%准确率。分析进一步揭示输出失败模式、输入条件敏感性、合成微调下的性能迁移边界及内部去噪视角下的有限自修正能力——后续步骤主要优化早期假设,而非纠正推理错误。我们希望VGI-bench能推动下一代视频生成模型的发展。

原文摘要 · Abstract (English)

Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance 2.0, achieving only 51.0% under our evaluation criteria. Our analysis further explore the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors. We hope VGI-bench will help stimulate the development of next-generation video generation models. Website: https://hexuan21.github.io/VGI-Bench/

视频生成视觉推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。