通过5000万数据训练大模型,提升自动驾驶3D感知能力
STELLAR: Scaling 3D Perception Large Models for Autonomous Driving

- 基于稀疏窗口变换器,融合激光雷达、雷达、摄像头和地图先验
- 在5000万样本上训练超大规模模型,参数达5亿,性能超越现有方法
- 首次系统验证感知模型规模扩展的有效性,适合自动驾驶研发者参考
模型规模扩展在多样化数据的大规模训练中展现出显著成效。然而,由于需融合异构传感器数据并具备复杂的3D空间理解能力,这一范式是否适用于自动驾驶感知系统仍存疑问。为此,我们系统性地研究了规模对这类系统的影响。基于稀疏窗口变换器构建的STELLAR模型,扩展输入模态至激光雷达、雷达、相机与地图先验。在包含5000万驾驶样本的大规模数据集上训练,模型参数规模达5亿。大规模实验揭示了模型性能随模型规模、数据量和计算资源增长的实证规律。所提出的模型在Waymo开放数据集挑战中达到新基准,显著优于此前方法。本工作证明,大规模训练是提升自动驾驶感知模型能力的一条极具前景的路径。
原文摘要 · Abstract (English)
Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradigm would apply to autonomous driving perception systems due to unique challenges, such as fusing heterogeneous sensor data and the need for sophisticated 3D spatial understanding. To bridge this gap, we present a comprehensive study on systematically analyzing the impact of scale on these systems. We develop our STELLAR model based on Sparse Window Transformer, by extending the input modalities to include LiDAR, radar, camera, and map prior. We train the model on a large-scale dataset of 50 million driving examples with up to 500 million parameters. Our large-scale experiments reveal empirical scaling trends that connect model performance to model size, data, and compute. The resulting model establishes a new state-of-the-art on the Waymo Open Dataset challenge, outperforming prior arts by a large margin. Our work demonstrates that large-scale training is a highly promising path for advancing the capabilities of perception models for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。