用三轴扫描增强远距离几何,提升摄像头语义场景补全精度
Three Cars Approaching within 100m! Enhancing Distant Geometry by Tri-Axis Voxel Scanning for Camera-based Semantic Scene Completion
- 三轴分步掩码注意力捕捉远近体素关系
- 远距离体素分割mIoU达20.14,优于现有方法
- 适合自动驾驶中长距感知需求
基于摄像头的语义场景补全(SSC)在三维感知领域受到关注。但透视和遮挡导致远距离区域几何信息被低估,对安全导向的自动驾驶系统构成挑战。为此,我们提出ScanSSC模型,包含扫描模块与扫描损失,通过近视角场景上下文增强远距离感知。扫描模块采用轴向掩码注意力,各轴使用由近及远的级联掩码,使远距离体素能捕捉与前序体素的关系。扫描损失沿各轴计算累积逻辑值与对应类别分布间的交叉熵,实现上下文感知信号向远距离体素传播。二者协同作用下,ScanSSC在SemanticKITTI和SSCBench-KITTI-360基准上分别取得44.54和48.29的IoU,以及17.40和20.14的mIoU,达到当前最优性能。
原文摘要 · Abstract (English)
Camera-based Semantic Scene Completion (SSC) is gaining attentions in the 3D perception field. However, properties such as perspective and occlusion lead to the underestimation of the geometry in distant regions, posing a critical issue for safety-focused autonomous driving systems. To tackle this, we propose ScanSSC, a novel camera-based SSC model composed of a Scan Module and Scan Loss, both designed to enhance distant scenes by leveraging context from near-viewpoint scenes. The Scan Module uses axis-wise masked attention, where each axis employing a near-to-far cascade masking that enables distant voxels to capture relationships with preceding voxels. In addition, the Scan Loss computes the cross-entropy along each axis between cumulative logits and corresponding class distributions in a near-to-far direction, thereby propagating rich context-aware signals to distant voxels. Leveraging the synergy between these components, ScanSSC achieves state-of-the-art performance, with IoUs of 44.54 and 48.29, and mIoUs of 17.40 and 20.14 on the SemanticKITTI and SSCBench-KITTI-360 benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。