通过智能筛选长尾数据,让自动驾驶规划更安全可靠。
Mining the Long Tail: A Comparative Study of Data-Centric Criticality Metrics for Robust Offline Reinforcement Learning in Autonomous Motion Planning
- 用模型不确定性等六种方法筛选关键数据样本
- 碰撞率从16.0%降至5.5%,安全提升近三倍
- 按时间步或完整场景加权,各有侧重适用场景
离线强化学习(Offline RL)为从大规模真实驾驶日志中训练自动驾驶规划策略提供了新范式。然而,日志中常见场景远多于罕见的‘长尾’事件,若采用均匀采样会导致策略脆弱且不安全。本文系统性地比较了六种数据筛选策略,分为启发式、不确定性与行为三类,分别在时间步和完整场景两个尺度上评估。使用基于注意力的先进架构训练七组目标条件化保守Q学习(CQL)智能体,在高保真Waymax仿真器中测试。结果表明,所有数据筛选方法均显著优于基线。尤其以模型不确定性为信号的数据驱动筛选,使碰撞率从16.0%降至5.5%,安全性提升近三倍。此外,时间步级加权利于反应式安全,场景级加权则改善长期规划。本研究为离线强化学习中的数据筛选提供全面框架,强调非均匀采样对构建安全可靠自主智能体至关重要。
原文摘要 · Abstract (English)
Offline Reinforcement Learning (RL) presents a promising paradigm for training autonomous vehicle (AV) planning policies from large-scale, real-world driving logs. However, the extreme data imbalance in these logs, where mundane scenarios vastly outnumber rare "long-tail" events, leads to brittle and unsafe policies when using standard uniform data sampling. In this work, we address this challenge through a systematic, large-scale comparative study of data curation strategies designed to focus the learning process on information-rich samples. We investigate six distinct criticality weighting schemes which are categorized into three families: heuristic-based, uncertainty-based, and behavior-based. These are evaluated at two temporal scales, the individual timestep and the complete scenario. We train seven goal-conditioned Conservative Q-Learning (CQL) agents with a state-of-the-art, attention-based architecture and evaluate them in the high-fidelity Waymax simulator. Our results demonstrate that all data curation methods significantly outperform the baseline. Notably, data-driven curation using model uncertainty as a signal achieves the most significant safety improvements, reducing the collision rate by nearly three-fold (from 16.0% to 5.5%). Furthermore, we identify a clear trade-off where timestep-level weighting excels at reactive safety while scenario-level weighting improves long-horizon planning. Our work provides a comprehensive framework for data curation in Offline RL and underscores that intelligent, non-uniform sampling is a critical component for building safe and reliable autonomous agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。