用双状态视频构建3D高斯世界模型,实现真实场景高效重建与模拟。
DSG-World: Learning a 3D Gaussian World Model from Dual State Videos
- 通过双状态观测互补视角,缓解遮挡问题并提升重建完整性。
- 端到端训练,支持高保真渲染与物体级场景操作,无需多阶段处理。
- 适合需要真实世界3D建模与仿真迁移的机器人与视觉研究者。
从有限观测中构建高效且物理一致的世界模型是视觉与机器人领域的长期挑战。现有方法多基于隐式生成模型,训练困难且缺乏3D或物理一致性;而基于单状态的显式3D方法常需分段处理(如分割、背景补全、修复),受遮挡影响大。本文提出DSG-World,一种从双状态视频中端到端构建3D高斯世界模型的新框架。利用同一场景在不同物体配置下的两个扰动观测,双状态提供互补可见性,缓解状态转移中的遮挡问题,实现更稳定完整的重建。方法构建双分割感知的高斯场,并强制双向光度与语义一致性。进一步引入伪中间状态进行对称对齐,并设计协同共剪枝策略优化几何完整性。该模型可直接在显式高斯表示空间中完成真实到仿真迁移,支持高保真渲染与物体级场景操控,无需密集观测或多阶段流程。大量实验表明其在新视角和新状态上具有强泛化能力,验证了方法在真实世界3D重建与仿真中的有效性。
原文摘要 · Abstract (English)
Building an efficient and physically consistent world model from limited observations is a long standing challenge in vision and robotics. Many existing world modeling pipelines are based on implicit generative models, which are hard to train and often lack 3D or physical consistency. On the other hand, explicit 3D methods built from a single state often require multi-stage processing-such as segmentation, background completion, and inpainting-due to occlusions. To address this, we leverage two perturbed observations of the same scene under different object configurations. These dual states offer complementary visibility, alleviating occlusion issues during state transitions and enabling more stable and complete reconstruction. In this paper, we present DSG-World, a novel end-to-end framework that explicitly constructs a 3D Gaussian World model from Dual State observations. Our approach builds dual segmentation-aware Gaussian fields and enforces bidirectional photometric and semantic consistency. We further introduce a pseudo intermediate state for symmetric alignment and design collaborative co-pruning trategies to refine geometric completeness. DSG-World enables efficient real-to-simulation transfer purely in the explicit Gaussian representation space, supporting high-fidelity rendering and object-level scene manipulation without relying on dense observations or multi-stage pipelines. Extensive experiments demonstrate strong generalization to novel views and scene states, highlighting the effectiveness of our approach for real-world 3D reconstruction and simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。