用大模型提升自动泊车的语义理解与可解释性
ParkingTransformer: LLM-Enhanced End-to-End Trajectory Planning for Autonomous Parking

- 结合大模型与多视角感知,直接从原始数据生成轨迹
- 真实场景成功率88.7%,模拟器得分61.32,表现优于传统方法
- 适合关注自动驾驶可解释性与端到端规划的研究者
端到端自动泊车已成为自动驾驶关键任务。现有方法存在黑箱特性,缺乏高层语义理解与可解释性,难以实现从道路到目标车位的无缝长距离自动泊车。为此,我们提出ParkingTransformer框架,利用多视角感知与大语言模型(LLM)的场景理解能力。通过将轨迹查询与LLM隐式状态特征结合,直接与历史信息和原始传感器数据交互输出规划轨迹,无需密集鸟瞰图(BEV)表示。为弥补LLM空间推理不足,引入3D位置编码显式注入空间几何意识。此外,设计固定窗口流式机制处理历史信息,显著提升长期时序处理效率与推理速度。采用粗到精解码策略逐步提升轨迹精度。在CARLA模拟器与真实车辆平台进行大量闭环实验,结果表明该方法在CARLA中取得61.32分驾驶得分,真实场景平均成功率达88.70%,验证了算法的可行性和有效性。
原文摘要 · Abstract (English)
End-to-end autonomous parking has emerged as a critical task within the realm of autonomous driving. However, existing methods suffer from black-box characteristics, lacking high-level semantic understanding and interpretability, which impedes the realization of seamless long-distance autonomous parking from the road to the target spot. To address these limitations, we propose ParkingTransformer, a novel framework that leverages multi-view perception and the scene understanding capability of Large Language Models (LLMs). By combining trajectory queries with LLMs implicit state features, our method interacts directly with historical information and raw sensor data to output planning trajectories, eliminating the need for dense Bird's-View (BEV) representations. To compensate for the inadequate spatial reasoning ability of LLMs, we introduce 3D positional encoding to explicitly inject spatial geometric awareness. Furthermore, a fixed-window streaming mechanism is designed for historical information processing, significantly improving long-term temporal processing efficiency and inference speed. Additionally, a coarse-to-fine decoding strategy is employed to progressively enhance trajectory precision. Extensive closed-loop experiments are conducted on the CARLA simulator and real-world vehicle platforms. The results demonstrate that our method achieves a driving score of 61.32 in CARLA simulator and an average success rate of 88.70% in real-world experiments, validating the feasibility and effectiveness of the proposed algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。