提出新方法实现50+视角长期3D重建,大幅减少冗余并保持几何一致。
SaLon3R: Structure-aware Long-term Generalizable 3D Reconstruction from Unposed Images
- 用紧凑锚点压缩冗余点,结合注意力机制自适应优化
- 支持超10帧/秒实时重建,冗余降低50%至90%
- 无需相机参数或后处理,适合长视频3D建模
近期3D高斯溅射(3DGS)进展实现了序列输入视图的通用、即时重建。但现有方法常逐像素预测高斯并合并所有视角,导致长时间视频中出现大量冗余和几何不一致。为此,我们提出SaLon3R——一种结构感知的长期3DGS重建框架。据我们所知,SaLon3R是首个在线通用的GS方法,可在超过10 FPS下重建超过50个视角,并实现50%至90%的冗余移除。方法引入紧凑锚点原语,通过可微分的显著性感知高斯量化消除冗余,同时使用3D点变换器从训练数据中学习三维空间结构先验,以精炼锚点属性与显著性,实现区域自适应高斯解码,提升几何保真度。无需已知相机参数或测试时优化,本方法在单次前向传播中有效消除伪影并剪枝冗余3DGS。多个数据集上的实验表明,其在新视角合成与深度估计上均达当前最优性能,展现出卓越的效率、鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Recent advances in 3D Gaussian Splatting (3DGS) have enabled generalizable, on-the-fly reconstruction of sequential input views. However, existing methods often predict per-pixel Gaussians and combine Gaussians from all views as the scene representation, leading to substantial redundancies and geometric inconsistencies in long-duration video sequences. To address this, we propose SaLon3R, a novel framework for Structure-aware, Long-term 3DGS Reconstruction. To our best knowledge, SaLon3R is the first online generalizable GS method capable of reconstructing over 50 views in over 10 FPS, with 50% to 90% redundancy removal. Our method introduces compact anchor primitives to eliminate redundancy through differentiable saliency-aware Gaussian quantization, coupled with a 3D Point Transformer that refines anchor attributes and saliency to resolve cross-frame geometric and photometric inconsistencies. Specifically, we first leverage a 3D reconstruction backbone to predict dense per-pixel Gaussians and a saliency map encoding regional geometric complexity. Redundant Gaussians are compressed into compact anchors by prioritizing high-complexity regions. The 3D Point Transformer then learns spatial structural priors in 3D space from training data to refine anchor attributes and saliency, enabling regionally adaptive Gaussian decoding for geometric fidelity. Without known camera parameters or test-time optimization, our approach effectively resolves artifacts and prunes the redundant 3DGS in a single feed-forward pass. Experiments on multiple datasets demonstrate our state-of-the-art performance on both novel view synthesis and depth estimation, demonstrating superior efficiency, robustness, and generalization ability for long-term generalizable 3D reconstruction. Project Page: https://wrld.github.io/SaLon3R/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。