测试顶级视频模型能否零样本推理,发现其短时推理尚可,长时因果推理仍不足。
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
- 用新基准MME-CoF评估Veo-3在12个维度的链帧推理能力
- 短时空间一致性与局部动态推理表现良好,长时因果与抽象逻辑仍弱
- 适合用于辅助推理模型,而非独立承担复杂视觉推理任务
近期视频生成模型能产出高保真、时间连贯的视频,表明其可能蕴含大量世界知识。除了真实合成外,它们还展现出视觉感知、建模与操作等新兴行为。但关键问题仍存:视频模型是否已具备在复杂视觉推理场景中作为零样本推理者的能力?本文通过实证研究全面考察这一问题,聚焦主流模型Veo-3。我们从空间、几何、物理、时间及具身逻辑等12个维度系统评估其推理行为,构建了紧凑的MME-CoF基准,支持对链帧(CoF)推理的深入评估。结果表明,当前模型在短时空间连贯性、细粒度定位和局部一致动态方面展现良好推理模式,但在长时因果推理、严格几何约束和抽象逻辑方面仍受限。总体而言,它们尚不足以作为可靠的独立零样本推理者,但作为专用推理模型的补充视觉引擎展现出令人鼓舞的潜力。
原文摘要 · Abstract (English)
Recent video generation models can produce high-fidelity, temporally coherent videos, indicating that they may encode substantial world knowledge. Beyond realistic synthesis, they also exhibit emerging behaviors indicative of visual perception, modeling, and manipulation. Yet, an important question still remains: Are video models ready to serve as zero-shot reasoners in challenging visual reasoning scenarios? In this work, we conduct an empirical study to comprehensively investigate this question, focusing on the leading and popular Veo-3. We evaluate its reasoning behavior across 12 dimensions, including spatial, geometric, physical, temporal, and embodied logic, systematically characterizing both its strengths and failure modes. To standardize this study, we curate the evaluation data into MME-CoF, a compact benchmark that enables in-depth and thorough assessment of Chain-of-Frame (CoF) reasoning. Our findings reveal that while current video models demonstrate promising reasoning patterns on short-horizon spatial coherence, fine-grained grounding, and locally consistent dynamics, they remain limited in long-horizon causal reasoning, strict geometric constraints, and abstract logic. Overall, they are not yet reliable as standalone zero-shot reasoners, but exhibit encouraging signs as complementary visual engines alongside dedicated reasoning models. Project page: https://video-cof.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。