实时构建带语言对齐的3D场景,15帧/秒下精度超越纯几何方法
LangGS-SLAM: Real-Time Language-Feature Gaussian Splatting SLAM
- 用顶K渲染管线高效生成高维语义特征图
- 通过多标准管理策略减少冗余点云,保持场景完整
- 分频优化几何与语义场,实现实时运行
本文提出一种RGB-D SLAM系统,可在保持低延迟跟踪与建图的同时,重建与语言对齐的密集特征场。首先,引入顶K渲染流水线,一种高吞吐、无语义失真的方法,用于高效渲染高维特征图。为解决由此带来的语义-几何不一致问题并降低内存消耗,进一步设计了多准则地图管理策略,剔除冗余或不一致的高斯点,同时保留场景完整性。最后,提出一种混合场优化框架,根据特征场特性解耦几何与语义场的优化频率,在实时约束下联合优化二者。所提系统在几何保真度上优于纯几何基线,在语义保真度上接近离线方法,且运行速度达15 FPS。结果表明,在线SLAM中实现密集、未压缩的语言对齐特征场既可行又有效,弥合了3D感知与基于语言推理之间的差距。
原文摘要 · Abstract (English)
In this paper, we propose a RGB-D SLAM system that reconstructs a language-aligned dense feature field while sustaining low-latency tracking and mapping. First, we introduce a Top-K Rendering pipeline, a high-throughput and semantic-distortion-free method for efficiently rendering high-dimensional feature maps. To address the resulting semantic-geometric discrepancy and mitigate the memory consumption, we further design a multi-criteria map management strategy that prunes redundant or inconsistent Gaussians while preserving scene integrity. Finally, a hybrid field optimization framework jointly refines the geometric and semantic fields under real-time constraints by decoupling their optimization frequencies according to field characteristics. The proposed system achieves superior geometric fidelity compared to geometric-only baselines and comparable semantic fidelity to offline approaches while operating at 15 FPS. Our results demonstrate that online SLAM with dense, uncompressed language-aligned feature fields is both feasible and effective, bridging the gap between 3D perception and language-based reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。