解决长序列3D重建中的位姿漂移问题,通过光线图引导实现几何与外观协同优化。
NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction

- 用光线图引导的耦合模块,联合优化几何与外观
- 在长序列上显著提升渲染质量与位姿估计精度
- 适合需要稳定快速重建的实时场景应用
无位姿前馈3D高斯点云拼贴(3DGS)已成为快速场景重建的有力范式,但在长图像序列中因累积位姿估计漂移导致性能严重下降,误差传播至几何建模并严重影响渲染保真度。本文重新审视长序列瓶颈,确认位姿漂移是限制重建质量的主要因素。尽管基于SfM的伪真值位姿引入传感器噪声,纯渲染监督常因几何与位姿联合优化纠缠而导致优化不稳和局部极小。为此,提出一种协同无位姿框架,通过光线图引导的耦合模块(RGC)显式关联几何与外观。具体地,将高斯中心锚定于光线图诱导的几何结构,并在统一目标下联合优化RGB重建、光线图一致性与相机正则化,形成双向反馈:更强几何提升渲染,外观监督反向优化几何与位姿。为进一步稳定跨长时间范围的学习,引入双频视角调度策略,结合由易到难的区间扩展与短区间对重放。在域内与跨域数据集上的大量实验表明,渲染与位姿估计均有持续提升,长序列鲁棒性显著增强。消融实验证实核心洞察:显式设计的几何-外观协同是可扩展且抗漂移无位姿前馈3D重建的关键。
原文摘要 · Abstract (English)
Pose-Free Feed-forward 3D Gaussian Splatting (3DGS) has recently emerged as a powerful paradigm for fast scene reconstruction. However, its performance degrades significantly in long image sequences due to cumulative camera pose estimation drift, which propagates errors into geometric modeling and severely limits rendering fidelity. In this work, we revisit the long-sequence bottleneck and identify pose drift as the primary factor restricting reconstruction quality. Furthermore, while SfM-based pseudo ground-truth poses introduce sensor noise, purely rendering-based supervision often leads to optimization instability and local minima due to the entangled optimization of geometry and pose. To address the challenges, we propose a synergistic pose-free framework that explicitly couples geometry and appearance via a Raymap-Guided Coupling Module (RGC). Concretely, we anchor Gaussian centers to raymap-induced geometry and jointly optimize RGB reconstruction, raymap consistency, and camera regularization under a unified objective, yielding a bidirectional feedback loop: stronger geometry improves rendering, and appearance supervision in turn refines geometry and pose. To further stabilize learning across wide temporal ranges, we introduce a Dual-Frequency Viewpoint Scheduling strategy that combines easy-to-hard interval expansion with replay of short-interval pairs. Extensive experiments across in-domain and cross-domain datasets show consistent gains in both rendering and pose estimation, with notably improved robustness on long sequences. Ablation studies validate our central insight: explicitly designed geometry-appearance synergy is the key to scalable and drift-robust pose-free feed-forward 3D reconstruction. Project page: https://xiangyu1sun.github.io/NoDrift3R-project-page/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。