用数据选择框架MOSAIC,让自动驾驶模型少用80%数据仍更优。
Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems

- 按数据领域划分,用缩放规律预测各数据对指标影响。
- 迭代添加提升最大指标的样本,用更少数据达更高性能。
- 适合想降数据成本又保效果的自动驾驶研发团队。
面向物理AI应用的大规模深度学习模型依赖多样化的数据收集。这些模型及其训练数据需满足部署于真实环境所需的多种评估标准。现有数据选择策略未考虑数据点对不同指标影响的不确定性。本文提出混合优化的缩放感知迭代采集框架(MOSAIC),其核心为:(i) 将数据集划分为不同领域;(ii) 针对每个领域拟合神经缩放规律至评估指标;(iii) 通过迭代添加使指标提升最大的数据域样本,优化数据混合比例。我们将MOSAIC应用于端到端(E2E)自动驾驶规划模型,使用扩展预测驾驶员模型评分(EPDMS)作为综合驾驶规则合规性评估指标。实验表明,MOSAIC在相同性能下可比基准方法减少高达80%的数据量。
原文摘要 · Abstract (English)
Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, the training data, must address different evaluation criteria necessary for the models to be deployable in real-world environments. Data selection policies can guide the development of the training set, but current frameworks do not account for the ambiguity in how data points affect different metrics. In this work, we propose Mixture Optimization via Scaling-Aware Iterative Collection (MOSAIC), a general data selection framework that operates by: (i) partitioning the dataset into domains; (ii) fitting neural scaling laws from each data domain to the evaluation metrics; and (iii) optimizing a data mixture by iteratively adding data from domains that maximize the change in metrics. We apply MOSAIC to autonomous driving (AD), where an End-to-End (E2E) planner model is evaluated on the Extended Predictive Driver Model Score (EPDMS), an aggregate of driving rule compliance metrics. Here, MOSAIC outperforms a diverse set of baselines on EPDMS with up to 80\% less data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。