Co-Win联合检测与分割激光点云,提升自动驾驶环境感知精度。
Co-Win: Joint Object Detection and Instance Segmentation in LiDAR Point Clouds via Collaborative Window Processing
- 通过分窗并行提取特征,融合点云编码与层级结构。
- 实现高精度实例分割,预测结果与场景上下文一致。
- 适合需要细粒度感知的自动驾驶系统部署。
复杂城市环境中精准感知与场景理解是保障自动驾驶安全高效的关键挑战。本文提出Co-Win,一种新型鸟瞰图(BEV)感知框架,通过点云编码与高效的并行分窗特征提取,应对环境理解中的多模态问题。该方法采用包含专用编码器、基于窗口的主干网络和基于查询的解码头的分层架构,有效捕捉多样化的空间特征与物体间关系。不同于以往将感知视为简单回归任务的方法,本框架引入变分方法与基于掩码的实例分割,实现细粒度场景分解与理解。Co-Win通过渐进式特征提取阶段处理点云数据,确保预测掩码在数据一致性与上下文相关性方面表现优异。此外,该方法生成可解释且多样化的实例预测,有助于提升自动驾驶系统的下游决策与规划能力。
原文摘要 · Abstract (English)
Accurate perception and scene understanding in complex urban environments is a critical challenge for ensuring safe and efficient autonomous navigation. In this paper, we present Co-Win, a novel bird's eye view (BEV) perception framework that integrates point cloud encoding with efficient parallel window-based feature extraction to address the multi-modality inherent in environmental understanding. Our method employs a hierarchical architecture comprising a specialized encoder, a window-based backbone, and a query-based decoder head to effectively capture diverse spatial features and object relationships. Unlike prior approaches that treat perception as a simple regression task, our framework incorporates a variational approach with mask-based instance segmentation, enabling fine-grained scene decomposition and understanding. The Co-Win architecture processes point cloud data through progressive feature extraction stages, ensuring that predicted masks are both data-consistent and contextually relevant. Furthermore, our method produces interpretable and diverse instance predictions, enabling enhanced downstream decision-making and planning in autonomous driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。