arXiv:2603.28489eess.IVcs.CV2026-03被引 4

提出视频生成世界模型的三维度高效框架,助力实时仿真应用。

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms

  • 从建模范式、网络结构、推理算法三方面构建高效体系
  • 实证效率提升可支撑自动驾驶等实时交互场景
  • 适合关注视频生成与具身智能融合的研究者

视频生成技术的快速发展使其能够模拟复杂的物理动态和长时序因果关系,具备成为世界模拟器的潜力。然而,理论上的世界模拟能力与时空建模带来的高计算成本之间仍存在显著差距。为此,本文系统性地综述了以效率为核心要求的视频生成框架与技术,提出了涵盖高效建模范式、高效网络架构和高效推理算法的三维新分类体系。研究进一步表明,缩小这一效率差距可直接推动自动驾驶、具身人工智能及游戏模拟等交互式应用的发展。最后,本文指出现有研究前沿,强调效率是视频生成模型演进为通用、实时、鲁棒世界模拟器的根本前提。相关文献汇总见GitHub:https://github.com/Isaachhh/Efficient-VWM-Survey。

原文摘要 · Abstract (English)

The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as potential world simulators. However, a critical gap still remains between the theoretical capacity for world simulation and the heavy computational costs of spatiotemporal modeling. To address this, we comprehensively and systematically review video generation frameworks and techniques that consider efficiency as a crucial requirement for practical world modeling. We introduce a novel taxonomy in three dimensions: efficient modeling paradigms, efficient network architectures, and efficient inference algorithms. We further show that bridging this efficiency gap directly empowers interactive applications such as autonomous driving, embodied AI, and game simulation. Finally, we identify emerging research frontiers in efficient video-based world modeling, arguing that efficiency is a fundamental prerequisite for evolving video generators into general-purpose, real-time, and robust world simulators. A curated GitHub repository of the reviewed literature can be found at https://github.com/Isaachhh/Efficient-VWM-Survey.

视频生成世界模型效率优化具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。