arXiv:2602.00096cs.CVcs.AI2026-02

仅用视频和现成资源构建物理可信的虚拟世界,用于高效训练智能体。

Mirage2Matter: A Physically Grounded Gaussian World Model from Video

  • 通过3D高斯泼溅从多视角视频重建真实场景几何与外观
  • 结合生成模型与标定靶点实现物理真实性与真实尺度对齐
  • 训练的视觉语言动作模型零样本表现媲美甚至超越真实数据

具身智能的可扩展性受限于真实交互数据的稀缺。尽管仿真平台提供替代方案,但现有方法常存在显著的视觉与物理偏差,依赖昂贵传感器、精确机器人校准或深度测量,难以规模化应用。我们提出「Simulate Anything」——一种基于图形的世界建模与仿真框架,仅需多视角环境视频和现成资产即可高效生成高保真具身训练数据。该方法利用3D高斯泼溅(3DGS)从视频中重建真实环境的逼真场景表示,无缝捕捉细微几何与外观。随后,通过生成模型恢复物理真实的表征,并借助精密标定靶点将其集成到仿真环境中,实现重建场景与真实世界的准确尺度对齐。上述组件共同构成统一、可编辑且物理可信的世界模型。在该模拟数据上训练的视觉语言动作(VLA)模型,在下游任务中展现出强劲的零样本性能,达到甚至超过使用真实数据的效果,凸显了基于重建的世界建模在可扩展且实用的具身智能训练中的潜力。

原文摘要 · Abstract (English)

The scalability of embodied intelligence is fundamentally constrained by the scarcity of real-world interaction data. While simulation platforms provide a promising alternative, existing approaches often suffer from a substantial visual and physical gap to real environments and rely on expensive sensors, precise robot calibration, or depth measurements, limiting their practicality at scale. We present Simulate Anything, a graphics-driven world modeling and simulation framework that enables efficient generation of high-fidelity embodied training data using only multi-view environment videos and off-the-shelf assets. Our approach reconstructs real-world environments into a photorealistic scene representation using 3D Gaussian Splatting (3DGS), seamlessly capturing fine-grained geometry and appearance from video. We then leverage generative models to recover a physically realistic representation and integrate it into a simulation environment via a precision calibration target, enabling accurate scale alignment between the reconstructed scene and the real world. Together, these components provide a unified, editable, and physically grounded world model. Vision Language Action (VLA) models trained on our simulated data achieve strong zero-shot performance on downstream tasks, matching or even surpassing results obtained with real-world data, highlighting the potential of reconstruction-driven world modeling for scalable and practical embodied intelligence training.

世界建模3D高斯具身智能视频重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。