提出RoboOcc,提升机器人对场景几何与语义的精细理解
RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots
- 用不透明度引导自编码器缓解高斯点云语义歧义
- 通过几何感知交叉编码器实现精细场景建模,提升全局与局部性能
- 在两个数据集上达到当前最佳,适合机器人感知研究者
3D占用预测使机器人能够获取周围场景的细粒度几何与语义信息,已成为具身感知的关键任务。现有基于3D高斯而非密集体素的方法未能有效利用高斯的几何与不透明度特性,限制了对复杂环境的估计能力及高斯对场景的描述效果。本文提出一种名为RoboOcc的3D占用预测方法,增强机器人对场景几何与语义的理解。该方法采用不透明度引导的自编码器(OSE)缓解重叠高斯点的语义歧义,并引入几何感知交叉编码器(GCE)实现对周围场景的精细几何建模。我们在Occ-ScanNet和EmbodiedOcc-ScanNet数据集上进行大量实验,RoboOcc在局部与全局相机设置下均取得领先性能。消融实验显示,相较现有最佳方法,其在IoU与mIoU指标上分别提升8.47和6.27。
原文摘要 · Abstract (English)
3D occupancy prediction enables the robots to obtain spatial fine-grained geometry and semantics of the surrounding scene, and has become an essential task for embodied perception. Existing methods based on 3D Gaussians instead of dense voxels do not effectively exploit the geometry and opacity properties of Gaussians, which limits the network's estimation of complex environments and also limits the description of the scene by 3D Gaussians. In this paper, we propose a 3D occupancy prediction method which enhances the geometric and semantic scene understanding for robots, dubbed RoboOcc. It utilizes the Opacity-guided Self-Encoder (OSE) to alleviate the semantic ambiguity of overlapping Gaussians and the Geometry-aware Cross-Encoder (GCE) to accomplish the fine-grained geometric modeling of the surrounding scene. We conduct extensive experiments on Occ-ScanNet and EmbodiedOcc-ScanNet datasets, and our RoboOcc achieves state-of the-art performance in both local and global camera settings. Further, in ablation studies of Gaussian parameters, the proposed RoboOcc outperforms the state-of-the-art methods by a large margin of (8.47, 6.27) in IoU and mIoU metric, respectively. The codes will be released soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。