通过意图增强提升驾驶策略多样性,突破单一示范导致的模式坍缩。
Driving Intents Amplify Planning-Oriented Reinforcement Learning
- 用无分类器引导的离散意图条件化流匹配,扩展动作分布
- 多意图GRPO使偏好强化学习保持多样性能,最佳样本达RFS 9.14
- 适合追求高多样性、高鲁棒性端到端自动驾驶系统的研究者
在每场景仅有一个示范轨迹的情况下训练连续动作策略会遭遇模式坍缩:样本聚集于示范行为,无法表示语义上不同的替代方案。基于偏好的评估下,这限制了最佳采样表现——即使使用最优选择也无法恢复采样分布中不存在的内容。本文提出DIAL框架,一种两阶段驾驶意图增强的强化学习方法,用于对齐偏好的连续动作驾驶策略。第一阶段,通过无分类器引导(CFG)将流匹配动作头以离散意图标签为条件,扩展采样分布并打破单示范模式坍缩。第二阶段,将该扩展分布引入多意图GRPO偏好强化学习,确保每个偏好组内涵盖所有意图类别,防止微调重新坍缩至当前偏好模式。在包含八种规则推导意图的端到端驾驶任务中,基于WOD-E2E数据集评估显示:竞争性视觉到动作(VA)与视觉-语言-动作(VLA)监督微调基线在最佳采样128次时仍低于人类示范;最强先前方法RAP在最佳采样64次下最高仅为RFS 8.5;而意图CFG采样将天花板提升至最佳采样128次时的RFS 9.14,首次超越先前最佳(RAP 8.5)和人类示范(8.13);多意图GRPO使保留测试集RFS从7.681提升至8.211,而所有单意图基线均表现更低且随训练下降。结果表明,连续动作策略在演示基础上进行偏好强化学习的瓶颈不仅在于如何更新策略,更在于如何扩展并保持被优化的采样分布。
原文摘要 · Abstract (English)
Continuous-action policies trained on a single demonstrated trajectory per scene suffer from mode collapse: samples cluster around the demonstrated maneuver and the policy cannot represent semantically distinct alternatives. Under preference-based evaluation, this caps best-of-N performance -- even oracle selection cannot recover what the sampling distribution does not contain. We introduce DIAL, a two-stage Driving-Intent-Amplified reinforcement Learning framework for preference-aligned continuous-action driving policies. In the first stage, DIAL conditions the flow-matching action head on a discrete intent label with classifier-free guidance (CFG), which expands the sampling distribution along distinct maneuver modes and breaks single-demonstration mode collapse. In the second stage, DIAL carries this expanded distribution into preference RL through multi-intent GRPO, which spans all intent classes within every preference group and prevents fine-tuning from re-collapsing around the currently preferred mode. Instantiated for end-to-end driving with eight rule-derived intents and evaluated on WOD-E2E: competitive Vision-to-Action (VA) and Vision-Language-Action (VLA) Supervised Finetuning (SFT) baselines plateau below the human-driven demonstration at best-of-128, with the strongest prior (RAP) capping at Rater Feedback Score (RFS) 8.5 even with best-of-64; intent-CFG sampling lifts this ceiling to RFS 9.14 at best-of-128, surpassing both the prior best (RAP 8.5) and the human-driven demonstration (8.13) for the first time; and multi-intent GRPO improves held-out RFS from 7.681 to 8.211, while every single-intent baseline peaks lower and degrades by training end. These results suggest that the bottleneck of preference RL on continuous-action policies trained from demonstrations is not only how to update the policy, but to expand and preserve the sampling distribution being optimized.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。