构建视频生成推理能力评估基准,发现大模型推理潜力与优化方法。
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models

- 设计分层评测框架,覆盖结构、空间、逻辑与动作规划四类推理
- 商业模型表现优于开源模型,但均存在可提升空间
- 提出无需训练的VideoTPO策略,通过大模型自评优化生成结果
视频生成模型的发展已从追求视觉逼真转向物理合理与逻辑一致。然而,现有评估多聚焦视觉保真度与时间连贯性,难以衡量高级推理能力。为此,我们提出TiViBench,一个面向图像到视频生成模型的分层推理评测基准,涵盖结构推理与搜索、空间与视觉模式推理、符号与逻辑推理、行动规划与任务执行四个维度,包含24种不同场景及三个难度层级。评估显示,商用模型(如Sora 2、Veo 3.1)展现更强推理潜力,而开源模型虽具潜力却受限于训练规模与数据多样性。为释放其潜能,我们引入VideoTPO——一种基于大模型自我分析的测试时优化策略,通过识别生成候选的优劣进行改进,无需额外训练、数据或奖励模型。TiViBench与VideoTPO共同为视频生成推理能力的评估与提升奠定基础。
原文摘要 · Abstract (English)
The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consistency. However, despite recent breakthroughs such as Veo 3's chain-of-frames reasoning, it remains unclear whether these models can exhibit reasoning capabilities similar to large language models (LLMs). Existing benchmarks predominantly evaluate visual fidelity and temporal coherence, failing to capture higher-order reasoning abilities. To bridge this gap, we propose TiViBench, a hierarchical benchmark specifically designed to evaluate the reasoning capabilities of image-to-video (I2V) generation models. TiViBench systematically assesses reasoning across four dimensions: i) Structural Reasoning & Search, ii) Spatial & Visual Pattern Reasoning, iii) Symbolic & Logical Reasoning, and iv) Action Planning & Task Execution, spanning 24 diverse task scenarios across 3 difficulty levels. Through extensive evaluations, we show that commercial models (e.g., Sora 2, Veo 3.1) demonstrate stronger reasoning potential, while open-source models reveal untapped potential that remains hindered by limited training scale and data diversity. To further unlock this potential, we introduce VideoTPO, a simple yet effective test-time strategy inspired by preference optimization. By performing LLM self-analysis on generated candidates to identify strengths and weaknesses, VideoTPO significantly enhances reasoning performance without requiring additional training, data, or reward models. Together, TiViBench and VideoTPO pave the way for evaluating and advancing reasoning in video generation models, setting a foundation for future research in this emerging field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。