用纠错机制提升自动驾驶泊车训练效率,3D高保真仿真中实现稳定表现。
ParkingWorld: End-to-End Autonomous Parking Reinforcement Learning from Corrective Experience in 3DGS Simulation

- 构建分层回放缓冲区,融合人类纠正与失败轨迹,实现高效学习。
- 在3DGS仿真与实车平台测试中,泊车成功率显著提升,安全性能更优。
- 适合需要高可靠性自主泊车的智能驾驶研发团队使用。
自动驾驶泊车需在狭窄、杂乱且高度受限的环境中进行精确低速操作,车辆必须在避开静态障碍物和复杂几何边界的同时完成泊入。与依赖海量高质量专家示范的模仿学习不同,传统强化学习方法常面临训练开销大、探索效率低,甚至在复杂场景下无法学到有效泊车策略的问题。为此,本文提出一种闭环纠错的样本高效强化学习框架(CIL-SERL),完全在逼真的3D高斯泼溅(3DGS)泊车仿真器中训练,该仿真器可实现真实场景的高保真数字重建。受学习实践中纠错笔记的启发,设计了多层级回放缓冲区机制,将标准RL轨迹、人工纠正干预、失败探索轨迹及基于回滚的修正片段分别存储于相互关联的记忆区域,支持结构化采样与针对性学习。所提框架在3DGS仿真环境与物理车辆平台上进行了系统评估。大量实验结果表明,该方法在多样化场景中显著提升了泊车成功率、运行效率与安全性,验证了基于CIL-SERL的端到端自动驾驶泊车方案的有效性与实用性。
原文摘要 · Abstract (English)
Autonomous parking demands precise low-speed maneuvering within narrow, cluttered, and highly constrained environments, where vehicles must navigate tight spaces while avoiding static obstacles and complex geometric boundaries. Unlike imitation learning, which typically requires massive volumes of high-quality expert demonstrations to converge to a stable policy and often suffers from limited generalization to unseen scenarios, traditional reinforcement learning (RL) methods face persistent challenges including excessive training overhead, inefficient exploration, and even failure to learn viable parking strategies in challenging settings. To address these limitations, this paper presents a correction-in-the-loop sample-efficient reinforcement learning (CIL-SERL) framework for end-to-end autonomous parking, which is entirely trained in a photorealistic 3D Gaussian Splatting (3DGS) parking simulator that enables high-fidelity digital reconstruction of real-world scenes. Inspired by error-correction notebooks used in learning practice, we design a novel multi-level replay buffer mechanism. These buffers hierarchically organize and store standard RL rollouts, human corrective interventions, failed exploration trajectories, and rollback-based correction segments in separate yet interconnected memory regions, facilitating structured sampling and targeted learning during training. The proposed framework is systematically evaluated in both the 3DGS simulation environment and a physical vehicle platform. Extensive experimental results demonstrate that our method achieves substantial improvements in parking success rate, operational efficiency, and safety performance across diverse scenarios, validating the effectiveness and practical applicability of the proposed CIL-SERL-based end-to-end autonomous parking solution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。