用轻量模型实现单目场景的高效3D分割与压缩表示
StixelNExT++: Lightweight Monocular Scene Segmentation and Representation for Collective Perception
- 基于3D Stixel结构,通过聚类小单元提升物体分割精度
- 每帧处理仅需10毫秒,30米内表现媲美主流方法
- 适合车载系统等对实时性要求高的集体感知场景
本文提出StixelNExT++,一种面向单目感知系统的新型场景表示方法。在已有Stixel表示基础上,该方法推断三维Stixel并通过对更小的3D Stixel单元进行聚类,增强物体分割能力。该方法在保持高场景信息压缩率的同时,可灵活适配点云和鸟瞰图表示。其轻量级神经网络在自动生成的基于LiDAR的真值数据上训练,实现实时性能,单帧计算时间低至10毫秒。在Waymo数据集上的实验表明,在30米范围内表现具有竞争力,展示了StixelNExT++在自动驾驶系统集体感知中的潜力。
原文摘要 · Abstract (English)
This paper presents StixelNExT++, a novel approach to scene representation for monocular perception systems. Building on the established Stixel representation, our method infers 3D Stixels and enhances object segmentation by clustering smaller 3D Stixel units. The approach achieves high compression of scene information while remaining adaptable to point cloud and bird's-eye-view representations. Our lightweight neural network, trained on automatically generated LiDAR-based ground truth, achieves real-time performance with computation times as low as 10 ms per frame. Experimental results on the Waymo dataset demonstrate competitive performance within a 30-meter range, highlighting the potential of StixelNExT++ for collective perception in autonomous systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。