arXiv:2603.16736cs.CV2026-03被引 3

让视频生成的3D世界更一致,重建出清晰细节的点云场景。

World Reconstruction From Inconsistent Views

  • 用非刚性对齐将不一致视频帧统一到全局坐标系。
  • 重建点云精度显著优于基线方法,支持可探索3D环境。
  • 适合想从视频生成3D内容的研究者和开发者。

视频扩散模型能生成高质量、多样化的虚拟世界,但单帧间缺乏3D一致性,难以重建3D世界。为此,我们提出新方法:通过非刚性对齐将视频帧统一至全局一致的坐标系,实现高精度点云重建。首先,利用几何基础模型将每帧转化为像素级3D点云,此时表面因不一致而错位;随后,设计专用的非刚性迭代帧到模型ICP算法获得初始对齐,再通过全局优化进一步锐化点云;最后,以该点云为初始化,提出一种新型逆向形变渲染损失,从不一致视图中生成高质量、可探索的3D环境。实验表明,本方法重建的3D场景质量显著优于基线,有效将视频生成模型转化为3D一致的世界生成器。

原文摘要 · Abstract (English)

Video diffusion models generate high-quality and diverse worlds; however, individual frames often lack 3D consistency across the output sequence, which makes the reconstruction of 3D worlds difficult. To this end, we propose a new method that handles these inconsistencies by non-rigidly aligning the video frames into a globally-consistent coordinate frame that produces sharp and detailed pointcloud reconstructions. First, a geometric foundation model lifts each frame into a pixel-wise 3D pointcloud, which contains unaligned surfaces due to these inconsistencies. We then propose a tailored non-rigid iterative frame-to-model ICP to obtain an initial alignment across all frames, followed by a global optimization that further sharpens the pointcloud. Finally, we leverage this pointcloud as initialization for 3D reconstruction and propose a novel inverse deformation rendering loss to create high quality and explorable 3D environments from inconsistent views. We demonstrate that our 3D scenes achieve higher quality than baselines, effectively turning video models into 3D-consistent world generators.

3D重建视频生成点云扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。