用相机位姿统一控制与3D一致性,让游戏世界更真实可交互。
WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation
- 以相机位姿为统一表示,精确建模用户动作与3D世界的关系。
- 在长时序导航中实现95%以上的位置重访一致性和高视觉质量。
- 适合研究交互式3D生成、虚拟世界构建的开发者与研究员。
近期视频扩散变换器的进步使得可交互的游戏世界模型成为可能,支持用户在长时间内探索生成环境。然而,现有方法在精确动作控制和长期3D一致性方面仍存在挑战。大多数先前工作将用户动作视为抽象条件信号,忽略了动作与3D世界间的根本几何关联——动作引发相对相机运动,并累积形成全局相机位姿。本文提出以相机位姿作为统一几何表示,同时保障即时动作控制与长期3D一致性。首先,定义基于物理的连续动作空间,通过李代数表示用户输入,生成精确的6-DoF相机位姿,并通过相机嵌入器注入生成模型,确保动作对齐。其次,利用全局相机位姿作为空间索引,检索相关历史观测,实现在长时间导航中对位置的几何一致重访。为此,我们构建了一个大规模数据集,包含3,000分钟真实人类游戏数据,附带相机轨迹与文本描述。大量实验表明,本方法在动作可控性、长时序视觉质量和3D空间一致性上显著优于当前最优模型。
原文摘要 · Abstract (English)
Recent advances in video diffusion transformers have enabled interactive gaming world models that allow users to explore generated environments over extended horizons. However, existing approaches struggle with precise action control and long-horizon 3D consistency. Most prior works treat user actions as abstract conditioning signals, overlooking the fundamental geometric coupling between actions and the 3D world, whereby actions induce relative camera motions that accumulate into a global camera pose within a 3D world. In this paper, we establish camera pose as a unifying geometric representation to jointly ground immediate action control and long-term 3D consistency. First, we define a physics-based continuous action space and represent user inputs in the Lie algebra to derive precise 6-DoF camera poses, which are injected into the generative model via a camera embedder to ensure accurate action alignment. Second, we use global camera poses as spatial indices to retrieve relevant past observations, enabling geometrically consistent revisiting of locations during long-horizon navigation. To support this research, we introduce a large-scale dataset comprising 3,000 minutes of authentic human gameplay annotated with camera trajectories and textual descriptions. Extensive experiments show that our approach substantially outperforms state-of-the-art interactive gaming world models in action controllability, long-horizon visual quality, and 3D spatial consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。