arXiv:2511.09057cs.CVcs.AI2025-11被引 35

PAN能根据自然语言指令模拟长期、可交互的逼真世界动态。

PAN: A World Model for General, Interactable, and Long-Horizon World Simulation

  • 用大语言模型+扩散模型构建可语言控制的虚拟世界模拟器
  • 支持跨领域、长时间、连贯的视频生成,动作条件一致
  • 适合需要长期推理与交互规划的智能体研究者

世界模型使智能体能够想象、预测和推理自身行为对世界演化的影响,从而进行规划与策略制定。现有视频生成模型通常仅支持从提示到完整视频的生成,缺乏因果控制、交互性及长时一致性,难以支撑有目的的推理。而以往世界建模工作多局限于特定领域(如物理、游戏或3D场景),深度与可控性有限,泛化能力弱。本文提出PAN,一种通用、可交互、长时程的世界模型,通过历史数据与自然语言动作条件,预测未来世界状态。PAN采用生成潜变量预测(GLP)架构,结合基于大语言模型的自回归潜变量动态主干,利用海量文本知识实现语言动作条件化;并辅以视频扩散解码器,重建感知细节丰富且时间连贯的视觉观测,实现潜空间推理(想象)与真实世界动态(现实)的统一。在涵盖多元领域的大规模视频-动作对数据上训练,PAN支持开放域、动作条件化的模拟,具备连贯的长期动态。大量实验表明,相比其他视频生成模型与世界模型,PAN在动作条件模拟、长时预测与模拟推理任务中表现优异,向通用世界模型迈进了一步。

原文摘要 · Abstract (English)

A world model enables an intelligent agent to imagine, predict, and reason about how the world evolves in response to its actions, and accordingly to plan and strategize. While recent video generation models produce realistic visual sequences, they typically operate in the prompt-to-full-video manner without causal control, interactivity, or long-horizon consistency required for purposeful reasoning. Existing world modeling efforts, on the other hand, often focus on restricted domains (e.g., physical, game, or 3D-scene dynamics) with limited depth and controllability, and struggle to generalize across diverse environments and interaction formats. In this work, we introduce PAN, a general, interactable, and long-horizon world model that predicts future world states through high-quality video simulation conditioned on history and natural language actions. PAN employs the Generative Latent Prediction (GLP) architecture that combines an autoregressive latent dynamics backbone based on a large language model (LLM), which grounds simulation in extensive text-based knowledge and enables conditioning on language-specified actions, with a video diffusion decoder that reconstructs perceptually detailed and temporally coherent visual observations, to achieve a unification between latent space reasoning (imagination) and realizable world dynamics (reality). Trained on large-scale video-action pairs spanning diverse domains, PAN supports open-domain, action-conditioned simulation with coherent, long-term dynamics. Extensive experiments show that PAN achieves strong performance in action-conditioned world simulation, long-horizon forecasting, and simulative reasoning compared to other video generators and world models, taking a step towards general world models that enable predictive simulation of future world states for reasoning and acting.

世界模型视频生成交互模拟长时预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。