arXiv:2412.02689cs.RO2024-12被引 10

研究端到端自动驾驶的数据扩展规律,发现数据质量比数量更重要。

Data Scaling Laws for Imitation Learning-Based End-to-End Autonomous Driving

  • 收集400万条驾驶示范,覆盖23类场景,验证数据规模与性能的关系。
  • 数据量增加可提升模型泛化能力,长尾数据少量增加效果显著。
  • 强调数据分布合理性,适合关注自动驾驶模型泛化的研究者。

端到端自动驾驶因可扩展性受到广泛关注,但现有方法受限于真实世界数据规模,难以全面探索其数据扩展规律。为此,我们从多种驾驶场景和行为中收集了大量数据,并对基于模仿学习的端到端自动驾驶范式进行了系统研究。共收集约400万条演示数据,涵盖23种不同场景类型,总计超过3万小时驾驶示范。在1,400个多样化驾驶示范中(开环1,300个,闭环100个)进行严格评估。实验发现:(1) 模型性能随数据量呈幂律增长,但闭环评估中不成立,表明评估方式影响对数据扩展规律的判断,应更关注数据分布而非单纯扩大规模;(2) 长尾场景数据少量增加可显著提升对应场景表现;(3) 合理扩展数据可使模型实现新场景与新动作的组合泛化。结果凸显数据扩展在提升模型跨场景泛化能力中的关键作用,为真实世界安全部署提供保障。项目仓库:https://github.com/ucaszyp/Driving-Scaling-Law

原文摘要 · Abstract (English)

The end-to-end autonomous driving paradigm has recently attracted lots of attention due to its scalability. However, existing methods are constrained by the limited scale of real-world data, which hinders a comprehensive exploration of the scaling laws associated with end-to-end autonomous driving. To address this issue, we collected substantial data from various driving scenarios and behaviors and conducted an extensive study on the scaling laws of existing imitation learning-based end-to-end autonomous driving paradigms. Specifically, approximately 4 million demonstrations from 23 different scenario types were gathered, amounting to over 30,000 hours of driving demonstrations. We performed open-loop evaluations and closed-loop simulation evaluations in 1,400 diverse driving demonstrations (1,300 for open-loop and 100 for closed-loop) under stringent assessment conditions. Through experimental analysis, we discovered that (1) the performance of the driving model exhibits a power-law relationship with the amount of data, but this is not the case in closed-loop evaluation. The inconsistency between the two assessments shifts our focus toward the distribution of data rather than merely expanding its volume. (2) a small increase in the quantity of long-tailed data can significantly improve the performance for the corresponding scenarios; (3) appropriate scaling of data enables the model to achieve combinatorial generalization in novel scenes and actions. Our results highlight the critical role of data scaling in improving the generalizability of models across diverse autonomous driving scenarios, assuring safe deployment in the real world.. Project repository: https://github.com/ucaszyp/Driving-Scaling-Law

自动驾驶数据扩展模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。