提出空间感知窗注意力机制,提升自动驾驶中语义占位预测的几何准确性。
SWA-SOP: Spatially-aware Window Attention for Semantic Occupancy Prediction in Autonomous Driving
- 在注意力计算中融入局部空间上下文,增强对稀疏与遮挡区域的感知能力。
- 在基于激光雷达的基准上达到当前最优性能,显著改善场景补全效果。
- 可跨模态通用,在相机方案中同样有效,适合多传感器自动驾驶系统。
自动驾驶感知系统依赖激光雷达和摄像头等传感器感知三维环境。然而,由于遮挡和数据稀疏性,这些传感器常无法获取完整信息。语义占位预测(SOP)通过推断未观测区域的占据状态与语义来解决此问题。现有基于Transformer的SOP方法在注意力计算中缺乏显式空间结构建模,导致几何感知能力有限,尤其在稀疏或遮挡区域表现不佳。为此,我们提出空间感知窗注意力(SWA),一种将局部空间上下文引入注意力的新机制。SWA显著提升了场景补全能力,并在基于激光雷达的SOP基准上取得当前最优结果。我们进一步验证其通用性,将其集成至基于摄像头的SOP流程中,跨模态均获得一致性能提升。
原文摘要 · Abstract (English)
Perception systems in autonomous driving rely on sensors such as LiDAR and cameras to perceive the 3D environment. However, due to occlusions and data sparsity, these sensors often fail to capture complete information. Semantic Occupancy Prediction (SOP) addresses this challenge by inferring both occupancy and semantics of unobserved regions. Existing transformer-based SOP methods lack explicit modeling of spatial structure in attention computation, resulting in limited geometric awareness and poor performance in sparse or occluded areas. To this end, we propose Spatially-aware Window Attention (SWA), a novel mechanism that incorporates local spatial context into attention. SWA significantly improves scene completion and achieves state-of-the-art results on LiDAR-based SOP benchmarks. We further validate its generality by integrating SWA into a camera-based SOP pipeline, where it also yields consistent gains across modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。