arXiv:2411.16877cs.CV2024-11被引 28

无需相机位姿,任意长度图像序列实时重建3D场景。

PreF3R: Pose-Free Feed-Forward 3D Gaussian Splatting from Variable-length Image Sequence

  • 用序列化记忆网络融合多视角图像,直接生成3D高斯场。
  • 每秒处理20帧,实现毫秒级新视角渲染。
  • 适合无标定设备、快速建模的实时应用。

我们提出 PreF3R,一种从任意长度图像序列中进行无位姿前馈式3D重建的方法。与以往方法不同,PreF3R无需相机标定,直接在标准坐标系中重建3D高斯场,从而实现高效的新视角渲染。通过利用 DUSt3R 的成对3D结构重建能力,并引入空间记忆网络处理序列化多视角输入,避免了基于优化的全局对齐。此外,PreF3R 集成了密集高斯参数预测头,支持可微分光栅化下的后续新视角合成。通过联合使用像素级损失和点云回归损失进行监督,显著提升图像真实感与结构精度。给定有序图像序列,PreF3R 可以以 20 FPS 的速度增量式重建3D高斯场,实现真正意义上的实时新视角渲染。实验表明,PreF3R 在无位姿条件下实现了高效的视图合成,且对未见场景具有强泛化能力。

原文摘要 · Abstract (English)

We present PreF3R, Pose-Free Feed-forward 3D Reconstruction from an image sequence of variable length. Unlike previous approaches, PreF3R removes the need for camera calibration and reconstructs the 3D Gaussian field within a canonical coordinate frame directly from a sequence of unposed images, enabling efficient novel-view rendering. We leverage DUSt3R's ability for pair-wise 3D structure reconstruction, and extend it to sequential multi-view input via a spatial memory network, eliminating the need for optimization-based global alignment. Additionally, PreF3R incorporates a dense Gaussian parameter prediction head, which enables subsequent novel-view synthesis with differentiable rasterization. This allows supervising our model with the combination of photometric loss and pointmap regression loss, enhancing both photorealism and structural accuracy. Given a sequence of ordered images, PreF3R incrementally reconstructs the 3D Gaussian field at 20 FPS, therefore enabling real-time novel-view rendering. Empirical experiments demonstrate that PreF3R is an effective solution for the challenging task of pose-free feed-forward novel-view synthesis, while also exhibiting robust generalization to unseen scenes.

3D重建实时渲染无位姿高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。