arXiv:2410.15636cs.CV2024-10被引 1

无需相机位姿即可用任意多张图像快速重建3D模型

LucidFusion: Reconstructing 3D Gaussians with Arbitrary Unposed Images

  • 将3D重建转为无位姿图像到图像的转换,用相对坐标图对齐图像
  • 新方法在无全局3D监督下仍能生成几何一致的3D高斯点云
  • 适合需要灵活输入、快速重建的3D应用开发者使用

近期大型重建模型在单图生成高质量3D物体方面取得显著进展。然而,当前方法通常依赖显式的相机位姿估计或固定视角,限制了其灵活性和实际应用。本文将3D重建重新定义为图像到图像的转换,并引入相对坐标图(RCM),在无需位姿估计的情况下将多张无姿态图像对齐至主视图。尽管RCM简化了流程,但缺乏全局3D监督可能导致输出噪声。为此,我们提出相对坐标高斯(RCG),将每个像素坐标视为高斯中心,并利用可微渲染实现一致的几何与位姿恢复。所提出的LucidFusion框架可处理任意数量的无姿态输入,在数秒内生成鲁棒的3D重建结果,为更灵活、无位姿依赖的3D流水线铺平道路。

原文摘要 · Abstract (English)

Recent large reconstruction models have made notable progress in generating high-quality 3D objects from single images. However, current reconstruction methods often rely on explicit camera pose estimation or fixed viewpoints, restricting their flexibility and practical applicability. We reformulate 3D reconstruction as image-to-image translation and introduce the Relative Coordinate Map (RCM), which aligns multiple unposed images to a main view without pose estimation. While RCM simplifies the process, its lack of global 3D supervision can yield noisy outputs. To address this, we propose Relative Coordinate Gaussians (RCG) as an extension to RCM, which treats each pixel's coordinates as a Gaussian center and employs differentiable rasterization for consistent geometry and pose recovery. Our LucidFusion framework handles an arbitrary number of unposed inputs, producing robust 3D reconstructions within seconds and paving the way for more flexible, pose-free 3D pipelines.

3D重建无位姿高斯渲染图像对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。