arXiv:2504.01016cs.GRcs.AI2025-04ICCV被引 41

用扩散模型从开放世界视频中恢复高保真时序点云,提升3D重建精度

GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors

  • 通过点图变分自编码器构建与视频分布无关的隐空间
  • 训练视频扩散模型生成具有时序一致性的点图序列
  • 在多个数据集上实现最佳3D精度和泛化能力,适合3D重建任务

尽管视频深度估计取得了显著进展,现有方法因仿射不变性预测导致几何保真度不足,限制了其在重建等度量相关下游任务中的应用。我们提出GeometryCrafter,一种新框架,可从开放世界视频中恢复具有时序一致性的高保真点图序列,支持精确的3D/4D重建、相机参数估计及其他基于深度的应用。核心在于一个点图变分自编码器(VAE),学习与视频隐空间分布无关的潜在表示,实现高效的点图编码与解码。基于该VAE,我们训练了一个视频扩散模型,以输入视频为条件建模点图序列分布。在多样数据集上的广泛评估表明,GeometryCrafter在3D精度、时序一致性和泛化能力方面达到当前最优水平。

原文摘要 · Abstract (English)

Despite remarkable advancements in video depth estimation, existing methods exhibit inherent limitations in achieving geometric fidelity through the affine-invariant predictions, limiting their applicability in reconstruction and other metrically grounded downstream tasks. We propose GeometryCrafter, a novel framework that recovers high-fidelity point map sequences with temporal coherence from open-world videos, enabling accurate 3D/4D reconstruction, camera parameter estimation, and other depth-based applications. At the core of our approach lies a point map Variational Autoencoder (VAE) that learns a latent space agnostic to video latent distributions for effective point map encoding and decoding. Leveraging the VAE, we train a video diffusion model to model the distribution of point map sequences conditioned on the input videos. Extensive evaluations on diverse datasets demonstrate that GeometryCrafter achieves state-of-the-art 3D accuracy, temporal consistency, and generalization capability.

3D重建扩散模型点图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。