将视频生成模型重构为具备状态与动态的通用世界模型,推动从视觉逼真到物理合理的演进。
A Mechanistic View on Video Generation as World Models: State and Dynamics
- 提出状态构建与动态建模双支柱分类框架,区分隐式与显式状态机制。
- 主张用物理持续性与因果推理替代单纯画质评估,提升模型功能性。
- 聚焦记忆增强与潜在因子解耦,适合追求可解释世界模拟的研究者。
大规模视频生成模型展现出涌现的物理一致性,被视为潜在的世界模型。然而,当前“无状态”视频架构与经典以状态为中心的世界模型理论之间仍存在鸿沟。本文通过提出以状态构建与动态建模为核心的新型分类体系,弥合这一差距。状态构建分为隐式范式(上下文管理)与显式范式(潜在压缩),动态建模则从知识融合与架构重构两个角度分析。同时倡导评估范式从视觉保真度转向功能基准,检验物理持续性与因果推理能力。最后指出两大关键前沿:通过数据驱动记忆与压缩保真度提升持续性,以及通过潜在因子解耦与推理优先集成推进因果性。解决这些挑战将使该领域从生成视觉可信视频迈向构建鲁棒、通用的世界模拟器。
原文摘要 · Abstract (English)
Large-scale video generation models have demonstrated emergent physical coherence, positioning them as potential world models. However, a gap remains between contemporary "stateless" video architectures and classic state-centric world model theories. This work bridges this gap by proposing a novel taxonomy centered on two pillars: State Construction and Dynamics Modeling. We categorize state construction into implicit paradigms (context management) and explicit paradigms (latent compression), while dynamics modeling is analyzed through knowledge integration and architectural reformulation. Furthermore, we advocate for a transition in evaluation from visual fidelity to functional benchmarks, testing physical persistence and causal reasoning. We conclude by identifying two critical frontiers: enhancing persistence via data-driven memory and compressed fidelity, and advancing causality through latent factor decoupling and reasoning-prior integration. By addressing these challenges, the field can evolve from generating visually plausible videos to building robust, general-purpose world simulators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。