首个多视图点云注册变压器,一步到位预测全局姿态。
FUSER: Feed-Forward MUltiview 3D Registration Transformer and SE(3)$^N$ Diffusion Refinement
- 用统一潜在空间联合处理所有扫描,跳过繁琐的成对匹配。
- 在3DMatch等数据集上精度超越现有方法,计算效率提升显著。
- 适合需要快速高精度三维重建的科研与工业场景。
多视图点云配准传统依赖大量成对匹配构建位姿图,计算开销大且缺乏全局几何约束时易出错。本文提出FUSER,首个前馈式多视图注册变换器,通过统一紧凑的潜在空间联合处理所有扫描,直接预测全局位姿,无需任何成对估计。为保持可扩展性,FUSER利用稀疏3D CNN将每帧扫描编码为低分辨率超点特征,保留绝对平移信息,并通过几何交替注意力模块实现高效的内部与跨扫描推理。特别地,借鉴现成基础模型的2D注意力先验,增强3D特征交互与几何一致性。基于FUSER,进一步提出FUSER-DF,一种在联合SE(3)^N空间进行去噪的扩散精修框架。FUSER作为代理注册模型构建去噪器,推导出先验条件化的SE(3)^N变分下界用于去噪监督。在3DMatch、ScanNet和ArkitScenes上的实验表明,该方法在注册精度和计算效率方面均达到最优表现。
原文摘要 · Abstract (English)
Registration of multiview point clouds conventionally relies on extensive pairwise matching to build a pose graph for global synchronization, which is computationally expensive and inherently ill-posed without holistic geometric constraints. This paper proposes FUSER, the first feed-forward multiview registration transformer that jointly processes all scans in a unified, compact latent space to directly predict global poses without any pairwise estimation. To maintain tractability, FUSER encodes each scan into low-resolution superpoint features via a sparse 3D CNN that preserves absolute translation cues, and performs efficient intra- and inter-scan reasoning through a Geometric Alternating Attention module. Particularly, we transfer 2D attention priors from off-the-shelf foundation models to enhance 3D feature interaction and geometric consistency. Building upon FUSER, we further introduce FUSER-DF, an SE(3)$^N$ diffusion refinement framework to correct FUSER's estimates via denoising in the joint SE(3)$^N$ space. FUSER acts as a surrogate multiview registration model to construct the denoiser, and a prior-conditioned SE(3)$^N$ variational lower bound is derived for denoising supervision. Extensive experiments on 3DMatch, ScanNet and ArkitScenes demonstrate that our approach achieves the superior registration accuracy and outstanding computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。