用闭环高效微调提升扩散模型轨迹生成的多样性和适应性
PlannerRFT: Reinforcing Diffusion Planners through Closed-Loop and Sample-Efficient Fine-Tuning
- 双分支优化同时改进轨迹分布与去噪引导
- 在nuPlan上实现10倍加速,支持大规模并行学习
- 可生成多模式、场景自适应轨迹,适合自动驾驶规划
基于扩散模型的规划器在自动驾驶中展现出生成类人轨迹的潜力。近期方法通过生成-评估循环中的奖励优化进行强化微调,以提升鲁棒性,但难以生成多模态、场景自适应轨迹,导致微调时信息性奖励利用效率低。为此,我们提出PlannerRFT,一种样本高效的扩散规划器强化微调框架。PlannerRFT采用双分支优化,在不改变原有推理流程的前提下,同时优化轨迹分布并自适应引导去噪过程向更有希望的探索方向。为支持规模化并行学习,我们开发了nuMax,其滚动速度比原生nuPlan快10倍。大量实验表明,PlannerRFT实现了当前最优性能,且在学习过程中涌现出显著不同的行为。
原文摘要 · Abstract (English)
Diffusion-based planners have emerged as a promising approach for human-like trajectory generation in autonomous driving. Recent works incorporate reinforcement fine-tuning to enhance the robustness of diffusion planners through reward-oriented optimization in a generation-evaluation loop. However, they struggle to generate multi-modal, scenario-adaptive trajectories, hindering the exploitation efficiency of informative rewards during fine-tuning. To resolve this, we propose PlannerRFT, a sample-efficient reinforcement fine-tuning framework for diffusion-based planners. PlannerRFT adopts a dual-branch optimization that simultaneously refines the trajectory distribution and adaptively guides the denoising process toward more promising exploration, without altering the original inference pipeline. To support parallel learning at scale, we develop nuMax, an optimized simulator that achieves 10 times faster rollout compared to native nuPlan. Extensive experiments shows that PlannerRFT yields state-of-the-art performance with distinct behaviors emerging during the learning process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。