通过轨迹熵最大化,40%数据量减少仍保持自动驾驶模型性能
Are All Data Necessary? Efficient Data Pruning for Large-scale Autonomous Driving Dataset via Trajectory Entropy Maximization
- 用轨迹分布熵衡量数据价值,迭代筛选高信息量样本
- 在NuPlan数据集上压缩40%数据,闭环性能几乎不变
- 理论可证减少分布偏移,适合大规模自动驾驶数据处理
收集大规模自然驾驶数据对训练鲁棒的自动驾驶规划器至关重要。然而,真实世界数据集常包含大量重复和低价值样本,导致存储成本过高且对策略学习收益有限。为此,我们提出一种基于信息论的数据剪枝方法,能在不损害模型性能的前提下显著减少训练数据量。该方法通过评估驾驶数据的轨迹分布信息熵,以模型无关的方式迭代选取能保留原始数据统计特性的高价值样本。理论上,最大化轨迹熵可有效约束剪枝子集与原始数据分布之间的Kullback-Leibler散度,从而保障泛化能力。在NuPlan基准和大规模模仿学习框架上的综合实验表明,该方法可将数据集规模减少高达40%,同时维持闭环性能。本工作为自动驾驶系统提供了轻量级、理论坚实的数据管理与高效策略学习方案。
原文摘要 · Abstract (English)
Collecting large-scale naturalistic driving data is essential for training robust autonomous driving planners. However, real-world datasets often contain a substantial amount of repetitive and low-value samples, which lead to excessive storage costs and bring limited benefits to policy learning. To address this issue, we propose an information-theoretic data pruning method that effectively reduces the training data volume without compromising model performance. Our approach evaluates the trajectory distribution information entropy of driving data and iteratively selects high-value samples that preserve the statistical characteristics of the original dataset in a model-agnostic manner. From a theoretical perspective, we show that maximizing trajectory entropy effectively constrains the Kullback-Leibler divergence between the pruned subset and the original data distribution, thereby maintaining generalization ability. Comprehensive experiments on the NuPlan benchmark with a large-scale imitation learning framework demonstrate that the proposed method can reduce the dataset size by up to 40% while maintaining closed-loop performance. This work provides a lightweight and theoretically grounded approach for scalable data management and efficient policy learning in autonomous driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。