arXiv:2511.16668cs.CV2025-11被引 20

构建视频生成模型的统一推理评估基准,覆盖四大核心能力。

V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models

论文配图:V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
图 1 · 摘自论文原文
  • 从合成与真实图像序列构建多维度可复现任务
  • 六款主流模型在物理推理上表现差异显著
  • 适合研究视频推理与幻觉行为的学者使用

近期生成式视频模型(如Veo-3)展现出惊人的零样本推理能力,亟需系统化、可靠的评估方法。本文提出V-ReasonBench,一个用于评估视频推理的基准测试套件,涵盖结构化问题求解、空间认知、模式推断和物理动态四个关键维度。该基准基于合成与真实世界图像序列,提供可验证答案、可复现、可扩展且无歧义的任务。对六款先进视频模型的评估揭示了各维度间明显差异,尤其在结构化、空间、模式和物理推理方面表现不一。进一步对比强图像模型,分析常见幻觉行为,并研究视频时长对帧链推理的影响。整体而言,V-ReasonBench提供了一个统一、可复现的框架,旨在推动具备更可靠、人类对齐推理能力的模型发展。

原文摘要 · Abstract (English)

Recent progress in generative video models, such as Veo-3, has shown surprising zero-shot reasoning abilities, creating a growing need for systematic and reliable evaluation. We introduce V-ReasonBench, a benchmark designed to assess video reasoning across four key dimensions: structured problem-solving, spatial cognition, pattern-based inference, and physical dynamics. The benchmark is built from both synthetic and real-world image sequences and provides a diverse set of answer-verifiable tasks that are reproducible, scalable, and unambiguous. Evaluations of six state-of-the-art video models reveal clear dimension-wise differences, with strong variation in structured, spatial, pattern-based, and physical reasoning. We further compare video models with strong image models, analyze common hallucination behaviors, and study how video duration affects Chain-of-Frames reasoning. Overall, V-ReasonBench offers a unified and reproducible framework for measuring video reasoning and aims to support the development of models with more reliable, human-aligned reasoning skills.

视频生成推理评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。