Puffin-World让模型自建3D世界,能物理模拟、视觉生成和闭环探索。
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

- 统一建模物理、几何与外观三类世界状态,用全景相机表示支持多任务
- 通过未来帧动态传播实现物理一致的3D世界生成,支持复杂运动轨迹
- 适合做具身智能、多模态交互与自主探索的研究者使用
我们提出Puffin-World,一种统一的多模态架构,无需依赖外部离线模块即可整合物理理解、空间模拟与3D世界生成与重建。为可靠构建和交互3D世界,框架联合建模三种原生世界状态:物理(重力场与纬度)、几何(深度)与外观(图像),并采用统一的全向相机表示,支持多样化任务与灵活运动。在建模基础上,引入物理动态跨帧传播策略。通过将绝对相机属性锚定真实世界,实现物理一致且视觉稳定的生成。进一步在单一生成流程中耦合外观与几何,联合合成每个未来视角并重建其底层几何结构。该统一范式支持多任务协同的闭环应用,如模仿与自校准世界探索。为扩展至复杂场景,我们构建了包含1500万视觉-语言-相机三元组与100万轨迹的Puffin-16M数据集。代码、模型与数据集已开源,以推动该领域研究。
原文摘要 · Abstract (English)
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。