用单目视频预测动态场景的3D结构,实现自监督学习。
Predicting 3D representations for Dynamic Scenes
- 采用以自我为中心的无限三平面表示动态世界
- 4D感知变压器融合视频特征,提升3D重建精度
- 在未见场景中表现优异,具备几何与语义学习能力
我们提出一种新框架,基于单目视频流预测动态辐射场。不同于以往仅关注未来帧预测的方法,本方法进一步生成动态场景的显式3D表示。框架包含两个核心设计:首先,采用以自我为中心的无界三平面来显式表达动态物理世界;其次,开发4D感知变换器,从单目视频中聚合特征以更新三平面。结合两者,可在大规模单目视频上进行自监督训练。模型在NVIDIA动态场景数据集上达到领先性能,证明其在4D物理世界建模上的强大能力。此外,模型对未见场景具有优越泛化性。值得注意的是,我们的方法展现出几何与语义学习的能力。
原文摘要 · Abstract (English)
We present a novel framework for dynamic radiance field prediction given monocular video streams. Unlike previous methods that primarily focus on predicting future frames, our method goes a step further by generating explicit 3D representations of the dynamic scene. The framework builds on two core designs. First, we adopt an ego-centric unbounded triplane to explicitly represent the dynamic physical world. Second, we develop a 4D-aware transformer to aggregate features from monocular videos to update the triplane. Coupling these two designs enables us to train the proposed model with large-scale monocular videos in a self-supervised manner. Our model achieves top results in dynamic radiance field prediction on NVIDIA dynamic scenes, demonstrating its strong performance on 4D physical world modeling. Besides, our model shows a superior generalizability to unseen scenarios. Notably, we find that our approach emerges capabilities for geometry and semantic learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。