arXiv:2602.05986cs.CVcs.AI2026-02被引 7

测试视频生成模型能否理解隐含世界规则,发现普遍缺陷。

RISE-Video: Can Video Generators Decode Implicit World Rules?

  • 构建多维度推理评测基准,聚焦深层认知而非表面画质。
  • 11个顶尖模型在复杂场景下普遍无法正确模拟隐含规则。
  • 适合研究视频生成智能、具身认知的学者与工程师。

尽管生成式视频模型已实现出色的视觉保真度,但其对隐含世界规则的内化与推理能力仍是关键而未被充分探索的领域。为此,我们提出 RISE-Video,首个面向文本-图像到视频(TI2V)合成的推理导向型评估基准,将评价重点从表层美学转向深层认知推理。该基准包含467个经人工精标样本,覆盖八个严格分类,涵盖常识、空间动态及特定主题领域,构成结构化测试平台。我们设计了四维评估协议:推理一致性、时间一致性、物理合理性与视觉质量。为支持可扩展评估,提出基于大视觉语言模型(LMMs)的自动化评测流水线,模拟人类判断。对11个前沿TI2V模型的广泛实验揭示,在隐含约束下模拟复杂场景存在普遍不足,为未来世界模拟生成模型的发展提供关键洞见。

原文摘要 · Abstract (English)

While generative video models have achieved remarkable visual fidelity, their capacity to internalize and reason over implicit world rules remains a critical yet under-explored frontier. To bridge this gap, we present RISE-Video, a pioneering reasoning-oriented benchmark for Text-Image-to-Video (TI2V) synthesis that shifts the evaluative focus from surface-level aesthetics to deep cognitive reasoning. RISE-Video comprises 467 meticulously human-annotated samples spanning eight rigorous categories, providing a structured testbed for probing model intelligence across diverse dimensions, ranging from commonsense and spatial dynamics to specialized subject domains. Our framework introduces a multi-dimensional evaluation protocol consisting of four metrics: \textit{Reasoning Alignment}, \textit{Temporal Consistency}, \textit{Physical Rationality}, and \textit{Visual Quality}. To further support scalable evaluation, we propose an automated pipeline leveraging Large Multimodal Models (LMMs) to emulate human-centric assessment. Extensive experiments on 11 state-of-the-art TI2V models reveal pervasive deficiencies in simulating complex scenarios under implicit constraints, offering critical insights for the advancement of future world-simulating generative models.

视频生成推理能力世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。