arXiv:2605.21372cs.CVcs.AI2026-05

自动优化真实与合成数据混合,提升自动驾驶模型性能。

Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training

论文配图:Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training
图 1 · 摘自论文原文
  • 构建闭环系统,动态调整数据组合以最大化模型表现。
  • 在有限预算下用更少合成数据实现更好效果,超越基线方法。
  • 适合关注数据高效训练与真实-合成协同的自动驾驶研究者。

数据规模是现代深度学习的基础,随着自动驾驶向端到端学习演进,其重要性愈发凸显。真实驾驶数据标注成本高且存在场景偏差,因此利用近乎无限的合成数据进行真实-合成协同训练成为有前景的方向。然而,盲目引入所有可用合成数据效率低下,易引发分布偏移,且在实际训练预算下优化数据混合仍是一个关键但未充分探索的问题。本文认为,训练数据混合需在场景类型和数量上具备明确指导。为此,我们提出一种将数据混合近似为动态优化过程的方法,通过闭环评估反馈迭代调整数据组合,并设计AutoScale——一个全自动闭环数据引擎,统一场景表征、数据混合优化与检索、模型训练与评估。具体而言,提出图正则化自编码器(Graph-RAE)进行驾驶场景表征,引入聚类感知梯度上升(Cluster-GA)进行聚类级重要性估计与重加权,并基于聚类引导向量检索选择高价值样本。在NavSim上的实验表明,AutoScale在受限预算下优于原始协同训练及跨域基线,使用更少合成样本获得更高性能。

原文摘要 · Abstract (English)

Data scaling is fundamental to modern deep learning, and grows increasingly critical as autonomous driving shifts to end-to-end learning. Real-world driving data is expensive to annotate and scene-biased, making real-synthetic co-training with near-infinite synthetic data a promising direction. However, naively incorporating all available synthetic data is inefficient and leads to distribution shifts, and optimizing data mixture under practical training budgets remains a critical yet under-explored problem. In this sense, we claim that the mixture of training data requires clear guidance in terms of scene types and quantities. Particularly in this work, we conceptualize the data mixture approximately as a dynamic optimization process that iteratively adjusts the training data mixture to maximize model performance, guided by closed-loop evaluation feedback, and propose AutoScale, a fully automated closed-loop data engine unifying scene representation, data mixture optimization and retrieval, as well as model training and evaluation. Specifically, we propose Graph Regularized AutoEncoder (Graph-RAE) for driving scene representations, introduce Cluster-aware Gradient Ascent (Cluster-GA) for cluster-wise importance estimation and reweighting, and perform cluster-guided vector retrieval to select high-value samples. Experiments on NavSim demonstrate that AutoScale outperforms vanilla co-training and cross-domain baselines, achieving better performance with fewer synthetic samples under constrained budgets.

自动驾驶数据混合闭环训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。