arXiv:2412.06974cs.CVcs.AI2024-12CVPR被引 159

2秒内完成多视角场景重建,一步到位避免误差累积。

MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds

论文配图:MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds
图 1 · 摘自论文原文
  • 单阶段网络直接处理多视图,用多视角解码块跨视图通信。
  • 在MVS和新视角合成任务上优于先前方法,推理速度提升显著。
  • 适合需要快速高精度重建的机器人、AR/VR应用。

近期稀疏多视角场景重建方法如DUSt3R和MASt3R不再依赖相机标定与位姿估计。然而,它们仅能逐对处理视图以推断像素对齐点云图。当处理超过两视图时,通常需进行大量易出错的成对重建,再经昂贵的全局优化,却常无法修正成对重建误差。为应对更多视图、降低误差并提升推理速度,我们提出快速单阶段前馈网络MV-DUSt3R。其核心是多视图解码块,在考虑一个参考视图的同时,实现任意数量视图间的交叉信息交换。为增强对参考视图选择的鲁棒性,进一步提出MV-DUSt3R+,采用跨参考视图块融合不同参考选择的信息。为进一步支持新视角合成,我们扩展两者并联合训练高斯点云渲染头。在多视图立体重建、多视图位姿估计及新视角合成任务上的实验表明,所提方法显著优于现有技术。代码将公开。

原文摘要 · Abstract (English)

Recent sparse multi-view scene reconstruction advances like DUSt3R and MASt3R no longer require camera calibration and camera pose estimation. However, they only process a pair of views at a time to infer pixel-aligned pointmaps. When dealing with more than two views, a combinatorial number of error prone pairwise reconstructions are usually followed by an expensive global optimization, which often fails to rectify the pairwise reconstruction errors. To handle more views, reduce errors, and improve inference time, we propose the fast single-stage feed-forward network MV-DUSt3R. At its core are multi-view decoder blocks which exchange information across any number of views while considering one reference view. To make our method robust to reference view selection, we further propose MV-DUSt3R+, which employs cross-reference-view blocks to fuse information across different reference view choices. To further enable novel view synthesis, we extend both by adding and jointly training Gaussian splatting heads. Experiments on multi-view stereo reconstruction, multi-view pose estimation, and novel view synthesis confirm that our methods improve significantly upon prior art. Code will be released.

3D重建多视角实时

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。