用自监督方法让车载传感器感知距离突破250米,适合长途自动驾驶。
Self-Supervised Sparse Sensor Fusion for Long Range Perception
- 基于稀疏表示和多模态时序编码,降低远距离感知的计算开销。
- 在250米距离上检测精度提升26.6%,激光雷达预测误差降低30.5%。
- 无需标注数据,可大规模利用摄像头与激光雷达原始数据预训练。
在城市以外地区,自动驾驶汽车和卡车需在城际高速上以超过100 km/h的速度安全行驶。为保障足够规划与制动空间,感知距离需达到至少250米,约为城市驾驶中50-100米的五倍。扩大感知范围还能将自动驾驶能力从轻型两吨乘用车扩展至大型四十吨卡车,因其惯性大,需更长规划周期。然而,现有感知方法多聚焦短距离,依赖鸟瞰图(BEV)表示,导致内存与计算成本随距离呈二次增长。为此,我们采用稀疏表示,引入高效的3D多模态时序特征编码,并设计新型自监督预训练方案,实现从无标签相机-激光雷达数据中大规模学习。该方法将感知距离扩展至250米,在目标检测上使mAP提升26.6%,在激光雷达预测中将Chamfer Distance降低30.5%,显著优于现有方法。
原文摘要 · Abstract (English)
Outside of urban hubs, autonomous cars and trucks have to master driving on intercity highways. Safe, long-distance highway travel at speeds exceeding 100 km/h demands perception distances of at least 250 m, which is about five times the 50-100m typically addressed in city driving, to allow sufficient planning and braking margins. Increasing the perception ranges also allows to extend autonomy from light two-ton passenger vehicles to large-scale forty-ton trucks, which need a longer planning horizon due to their high inertia. However, most existing perception approaches focus on shorter ranges and rely on Bird's Eye View (BEV) representations, which incur quadratic increases in memory and compute costs as distance grows. To overcome this limitation, we built on top of a sparse representation and introduced an efficient 3D encoding of multi-modal and temporal features, along with a novel self-supervised pre-training scheme that enables large-scale learning from unlabeled camera-LiDAR data. Our approach extends perception distances to 250 meters and achieves an 26.6% improvement in mAP in object detection and a decrease of 30.5% in Chamfer Distance in LiDAR forecasting compared to existing methods, reaching distances up to 250 meters. Project Page: https://light.princeton.edu/lrs4fusion/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。