arXiv:2511.08585cs.AIcs.CV2025-11被引 12

AI视频模型正从生成画面转向模拟可交互的虚拟世界。

Simulating the Visual World with Artificial Intelligence: A Roadmap

  • 将视频模型拆解为世界模拟引擎与视觉渲染器
  • 实现物理合理性、长时间一致性和任务规划能力
  • 适合机器人、自动驾驶和游戏等领域研究者参考

视频生成领域正经历转变,从单纯生成视觉吸引人的片段,发展为构建支持交互并保持物理合理性的虚拟环境。这一趋势催生了视频基础模型,它们不仅是视觉生成器,更是隐式世界模型——能够模拟物理动态、智能体与环境的互动以及任务规划。本文系统梳理这一演进过程,将现代视频基础模型视为隐式世界模型与视频渲染器的结合。世界模型编码关于物理规律、交互动态和智能体行为的结构化知识,作为潜在模拟引擎,支持连贯视觉推理、长期时间一致性及目标驱动规划。视频渲染器则将该潜伏模拟转化为真实视觉输出,相当于观察模拟世界的窗口。文章按四代演进路径展开:每一代逐步增强核心能力,最终形成基于视频生成模型的世界模型,具备内在物理合理性、实时多模态交互及跨时空尺度规划能力。针对各代特征,列举代表性工作并分析在机器人、自动驾驶、交互游戏等领域的应用。最后讨论下一代世界模型面临的挑战与设计原则,包括智能体智能在系统构建与评估中的作用。相关工作持续更新见链接。

原文摘要 · Abstract (English)

The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical plausibility. These developments point toward the emergence of video foundation models that function not only as visual generators but also as implicit world models, models that simulate the physical dynamics, agent-environment interactions, and task planning that govern real or imagined worlds. This survey provides a systematic overview of this evolution, conceptualizing modern video foundation models as the combination of two core components: an implicit world model and a video renderer. The world model encodes structured knowledge about the world, including physical laws, interaction dynamics, and agent behavior. It serves as a latent simulation engine that enables coherent visual reasoning, long-term temporal consistency, and goal-driven planning. The video renderer transforms this latent simulation into realistic visual observations, effectively producing videos as a "window" into the simulated world. We trace the progression of video generation through four generations, in which the core capabilities advance step by step, ultimately culminating in a world model, built upon a video generation model, that embodies intrinsic physical plausibility, real-time multimodal interaction, and planning capabilities spanning multiple spatiotemporal scales. For each generation, we define its core characteristics, highlight representative works, and examine their application domains such as robotics, autonomous driving, and interactive gaming. Finally, we discuss open challenges and design principles for next-generation world models, including the role of agent intelligence in shaping and evaluating these systems. An up-to-date list of related works is maintained at this link.

视频生成世界模型虚拟环境多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。