让3D场景实时带语义,无需逐场景优化
SegSplat: Feed-forward Gaussian Splatting and Open-Set Semantic Segmentation
- 用2D模型特征构建语义记忆库,单次前向传播预测3D高斯的语义标签
- 几何精度媲美顶尖方法,同时实现开集语义分割
- 适合需要快速生成带语义3D环境的机器人与AR应用
我们提出SegSplat,一种新型框架,旨在弥合快速前馈式3D重建与丰富开词汇语义理解之间的差距。通过从多视角2D基础模型特征构建紧凑语义记忆库,并在单次前向传播中为每个3D高斯预测离散语义索引及几何与外观属性,SegSplat高效地为场景赋予可查询的语义。实验表明,SegSplat在几何保真度上达到当前顶尖前馈式3D高斯点云方法水平,同时实现稳健的开集语义分割,关键在于无需任何针对特定场景的语义特征整合优化。该工作标志着迈向实用化、即时生成语义感知3D环境的重要一步,对推进机器人交互、增强现实及其他智能系统具有重要意义。
原文摘要 · Abstract (English)
We have introduced SegSplat, a novel framework designed to bridge the gap between rapid, feed-forward 3D reconstruction and rich, open-vocabulary semantic understanding. By constructing a compact semantic memory bank from multi-view 2D foundation model features and predicting discrete semantic indices alongside geometric and appearance attributes for each 3D Gaussian in a single pass, SegSplat efficiently imbues scenes with queryable semantics. Our experiments demonstrate that SegSplat achieves geometric fidelity comparable to state-of-the-art feed-forward 3D Gaussian Splatting methods while simultaneously enabling robust open-set semantic segmentation, crucially \textit{without} requiring any per-scene optimization for semantic feature integration. This work represents a significant step towards practical, on-the-fly generation of semantically aware 3D environments, vital for advancing robotic interaction, augmented reality, and other intelligent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。