分离自运动与环境动态,实现长期稳定的3D场景未来重建
Future Dynamic 3D Reconstruction: Toward 3D World Modeling with Disentangled Ego-Motion

- 将自运动与场景变化解耦,用隐变量建模代理动作
- 在单目观测下实现2秒后场景的几何一致重建
- 利用大模型空间常识,零样本泛化能力强
预测动态环境演化对自主智能体至关重要。尽管生成式世界模型在2D视频合成中实现了高保真度,但其将自运动与环境动态混合在图像平面内,导致长期时间跨度下出现物体变形或消失等物理不一致问题。本文提出FR3D,一种可预测未来动态3D重建的非监督世界建模方法。不同于以往将世界视为图像特征序列的方法,FR3D显式解耦场景3D演化与智能体轨迹,将推断出的自运动作为动作的潜在代理。该解耦机制消除了自运动与世界运动之间的歧义,保障了未来时空的几何一致性。此外,引入教师-学生蒸馏策略,利用现成基础模型的空间“常识”,实现鲁棒的零样本泛化。大量实验表明,FR3D在多个数据集上均能从单目观测中实现2秒后的动态3D重建,性能显著优于现有方法。
原文摘要 · Abstract (English)
Forecasting the evolution of dynamic environments is crucial for autonomous agents. While generative world models have achieved high photorealism in 2D video synthesis by mixing ego-motion and environmental dynamics within the image plane, they exhibit physical inconsistencies, such as morphing or vanishing objects, especially over long time horizons. In this paper, we propose FR3D, a world-modeling approach that predicts a persistent 3D latent representation for future dynamic 3D reconstruction. Unlike prior works that treat the world as a sequence of image-based features, FR3D explicitly decouples the 3D evolution of the scene from the agent's trajectory, treating the inferred ego-motion as a latent proxy for action. This disentanglement resolves ambiguities between self-motion and world-motion, ensuring geometric consistency into the future. Furthermore, we introduce a teacher-student distillation strategy that leverages the spatial "common sense" of off-the-shelf foundation models, leading to robust zero-shot generalization. Extensive experiments demonstrate FR3D's strong performance for future dynamic 3D reconstruction from monocular observations across multiple datasets, even 2 seconds into the future. Project page: https://fr3d-wm.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。