arXiv:2510.14000cs.RO2025-10被引 3

用强化学习先验优化扩散模型,提升狭窄空间停车规划成功率

A Diffusion-Refined Planner with Reinforcement Learning Priors for Confined-Space Parking

  • 用预训练强化学习策略提供动作先验,指导扩散模型训练
  • 在受限空间中规划成功率显著提升,推理步数减少
  • 适合自动驾驶停车系统研发人员参考

随着停车需求增长,对能在狭窄空间中可靠运行的自动化停车规划方法的需求日益迫切。在受限且复杂的环境中,高精度操作是实现高成功率规划的关键,但现有方法常依赖显式动作建模,难以准确捕捉最优动作分布。本文提出DRIP:一种基于强化学习(RL)先验动作分布的扩散精炼规划器。该方法利用预训练的强化学习策略提供先验动作分布,用于正则化扩散模型的训练过程。推理阶段,去噪过程将粗略先验细化为更精确的动作分布。通过在训练中引导去噪轨迹沿强化学习先验分布演化,扩散模型获得良好初始化,从而实现更精准的动作建模、更高的规划成功率,并减少推理步数。我们在不同空间约束程度的停车场景中评估了该方法。实验结果表明,该方法在狭窄空间停车环境中显著提升了规划性能,同时在常规场景中保持强泛化能力。

原文摘要 · Abstract (English)

The growing demand for parking has increased the need for automated parking planning methods that can operate reliably in confined spaces. In restricted and complex environments, high-precision maneuvers are required to achieve a high success rate in planning, yet existing approaches often rely on explicit action modeling, which faces challenges when accurately modeling the optimal action distribution. In this paper, we propose DRIP, a diffusion-refined planner anchored in reinforcement learning (RL) prior action distribution, in which an RL-pretrained policy provides prior action distributions to regularize the diffusion training process. During the inference phase the denoising process refines these coarse priors into more precise action distributions. By steering the denoising trajectory through the reinforcement learning prior distribution during training, the diffusion model inherits a well-informed initialization, resulting in more accurate action modeling, a higher planning success rate, and reduced inference steps. We evaluate our approach across parking scenarios with varying degrees of spatial constraints. Experimental results demonstrate that our method significantly improves planning performance in confined-space parking environments while maintaining strong generalization in common scenarios.

自动驾驶规划算法扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。