提出新评估框架,检测视频生成模型是否真会推理而非靠结果骗分。
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
- 用多阶段评估机制检查中间生成帧是否合理
- 顶尖模型仅20%能通过全程一致性检验
- 适合研究视频生成与视觉推理的学者
近期视频生成技术展现出一种名为帧链(Chain-of-Frames, CoF)的推理能力,即通过连续生成帧来解决复杂任务。尽管生成视频推理(Generative Video Reasoning, GVR)潜力巨大,现有评估方法多依赖单帧判断,易导致‘结果作弊’——模型虽得出正确结论,但过程错误。为此,我们提出过程感知评估范式,构建涵盖16项任务的基准测试VIPER,覆盖时间、结构、符号、空间、物理及规划推理。同时提出POC@r指标,利用视觉语言模型作为评判者,结合分层评分标准,同时评估中间步骤与最终结果的一致性。实验显示,当前最先进视频模型在[email protected]下表现仅为约20%,存在显著结果作弊现象。我们还探讨了测试时缩放与采样鲁棒性的影响,揭示现有视频生成与真正泛化视觉推理之间仍存在巨大差距。基准数据集已开源:https://github.com/RUCAIBox/VIPER。
原文摘要 · Abstract (English)
Recent breakthroughs in video generation have demonstrated an emerging capability termed Chain-of-Frames (CoF) reasoning, where models resolve complex tasks through the generation of continuous frames. While these models show promise for Generative Video Reasoning (GVR), existing evaluation frameworks often rely on single-frame assessments, which can lead to outcome-hacking, where a model reaches a correct conclusion through an erroneous process. To address this, we propose a process-aware evaluation paradigm. We introduce VIPER, a comprehensive benchmark spanning 16 tasks across temporal, structural, symbolic, spatial, physics, and planning reasoning. Furthermore, we propose Process-outcome Consistency (POC@r), a new metric that utilizes VLM-as-Judge with a hierarchical rubric to evaluate both the validity of the intermediate steps and the final result. Our experiments reveal that state-of-the-art video models achieve [email protected] only about 20% and exhibit a significant outcome-hacking. We further explore the impact of test-time scaling and sampling robustness, highlighting a substantial gap between current video generation and true generalized visual reasoning. Our benchmark are released at https://github.com/RUCAIBox/VIPER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。