用可微体素提升3D高斯的语义感知能力,实现无需相机位姿的开放词汇场景理解。
Bridging 3D Gaussians and Semantic Occupancy for Comprehensive Open-Vocabulary Scene Understanding from Unposed Images

- 将高斯分布与语义体素场联合优化,通过可微体素上采样提供梯度反馈
- 在ScanNet上实现比基线更高的开放词汇分割和语义占据精度
- 适合需要无标定视角下多模态3D场景理解的研究与应用
从稀疏、无标定图像中实现全面的3D场景理解,要求模型恢复可渲染几何、开放词汇语义和自由/占用3D空间,且不依赖外部相机校准。现有前馈高斯方法虽提升了无位姿重建与语义渲染效果,但其高斯原语主要通过图像空间目标优化,在未观测区域约束较弱。我们提出COVScene,一种无位姿的语义高斯框架,通过可微体素上采样将可渲染高斯原语与密集语义占据场耦合。不同于仅在评估时转换为体素,COVScene在训练计算图内将预测的语义高斯上采样至体素,使体素正则化提供梯度给高斯透明度、几何和语义特征。该框架结合语义感知几何变换器、多任务高斯解码、几何基础蒸馏与占据熵正则化,支持单一表示下的新视角合成、开放词汇语义查询与语义占据预测。在ScanNet和ScanNet++上的实验表明,COVScene保持了有竞争力的渲染质量,提升了开放词汇分割性能,并在无直接体素级监督的情况下,优于自监督基线的语义占据预测效果。
原文摘要 · Abstract (English)
Comprehensive 3D scene understanding from sparse, unposed images requires a model to recover renderable geometry, open-vocabulary semantics, and free/occupied 3D space without relying on external camera calibration. Recent feed-forward Gaussian methods improve pose-free reconstruction and semantic rendering, but their Gaussian primitives are mainly optimized through image-space objectives and remain weakly constrained in unobserved regions. We propose \textit{COVScene}, a pose-free semantic Gaussian framework that couples renderable Gaussian primitives with a dense semantic occupancy field through differentiable volumetric lifting. Instead of converting Gaussians to voxels only at evaluation time, COVScene lifts the predicted semantic Gaussians inside the training computation graph, so volumetric regularization provides gradients to Gaussian opacity, geometry, and semantic features. The framework combines a semantic-aware Geometry Transformer, multi-task Gaussian decoding, geometric foundation distillation, and occupancy entropy regularization to support novel view synthesis, open-vocabulary semantic querying, and semantic occupancy prediction within a single representation. Experiments on ScanNet and ScanNet++ show that COVScene maintains competitive rendering quality, improves open-vocabulary segmentation, and achieves stronger semantic occupancy prediction than the self-supervised baseline without direct voxel-level supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。