基于视觉的端到端自动泊车模型,误差降低50%。
TransParking: A Dual-Decoder Transformer Framework with Soft Localization for End-to-End Automatic Parking
- 双解码器Transformer架构,引入软定位机制。
- 在复杂环境下轨迹预测误差比当前最优算法降低约50%。
- 适合追求高精度端到端自动驾驶泊车的开发者参考。
近年来,全可微分的端到端自动驾驶系统成为智能交通领域的研究热点。其中,自动泊车尤为重要,旨在实现复杂环境下的精准停车。本文提出一种纯视觉的Transformer模型,用于端到端自动泊车,通过专家轨迹进行训练。给定摄像头采集的数据作为输入,该模型直接输出未来轨迹坐标。实验结果表明,与同类型当前最优端到端轨迹预测算法相比,本模型的各项误差降低了约50%。因此,该方法为全可微分自动泊车提供了一种有效解决方案。
原文摘要 · Abstract (English)
In recent years, fully differentiable end-to-end autonomous driving systems have become a research hotspot in the field of intelligent transportation. Among various research directions, automatic parking is particularly critical as it aims to enable precise vehicle parking in complex environments. In this paper, we present a purely vision-based transformer model for end-to-end automatic parking, trained using expert trajectories. Given camera-captured data as input, the proposed model directly outputs future trajectory coordinates. Experimental results demonstrate that the various errors of our model have decreased by approximately 50% in comparison with the current state-of-the-art end-to-end trajectory prediction algorithm of the same type. Our approach thus provides an effective solution for fully differentiable automatic parking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。