用精修模块提升3D语义场景补全效果,兼容多种现有模型。
Enhancing 3D Semantic Scene Completion with a Refinement Module
- 分两阶段:先粗预测,再用噪声感知与局部几何模块精修。
- 在SemanticKITTI上,平均交并比提升0.4%至0.43%,最高达17.27%。
- 可插即用,适合想提升现有3D补全模型性能的研究者。
我们提出ESSC-RM,一种可无缝集成到现有语义场景补全(SSC)模型中的即插即用增强框架。该框架分两阶段运行:首先由基线SSC网络生成粗粒度体素预测,随后在多尺度监督下,通过基于3D U-Net的预测噪声感知模块(PNAM)和体素级局部几何模块(VLGM)进行精修。在SemanticKITTI数据集上的实验表明,将ESSC-RM集成到CGFormer和MonoScene中,平均交并比(mean IoU)分别从16.87%提升至17.27%,从11.08%提升至11.51%。结果证明,ESSC-RM是一种通用的精修框架,适用于多种SSC模型。
原文摘要 · Abstract (English)
We propose ESSC-RM, a plug-and-play Enhancing framework for Semantic Scene Completion with a Refinement Module, which can be seamlessly integrated into existing SSC models. ESSC-RM operates in two phases: a baseline SSC network first produces a coarse voxel prediction, which is subsequently refined by a 3D U-Net-based Prediction Noise-Aware Module (PNAM) and Voxel-level Local Geometry Module (VLGM) under multiscale supervision. Experiments on SemanticKITTI show that ESSC-RM consistently improves semantic prediction performance. When integrated into CGFormer and MonoScene, the mean IoU increases from 16.87% to 17.27% and from 11.08% to 11.51%, respectively. These results demonstrate that ESSC-RM serves as a general refinement framework applicable to a wide range of SSC models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。