无需地图的零样本自动驾驶泊车,靠视觉语言模型实现自然语言导航。
VLN-AVP: Zero-Shot Vision-Language Navigation with Hybrid Long-Short-Term Memory for Autonomous Valet Parking

- 融合鸟瞰图与视觉语言模型,用混合记忆追踪环境语义。
- 仿真中成功率超现有方法25%以上,实车测试也表现领先。
- 首个地下停车场的零样本导航基准,适合开放场景智能驾驶研究。
现有自动驾驶泊车方法通常依赖预建地图,严重限制了其在未见环境和开放词汇目标下的可扩展性。受视觉语言导航任务中视觉语言模型应用的启发,我们提出 VLN-AVP,一种用于自动驾驶泊车的零样本导航框架。通过结合鸟瞰图模型的精确空间感知能力与视觉语言模型的通用智能,该框架实现了:1)消除对预建地图的依赖;2)解析泊车场景中的语义环境上下文;3)根据自然语言指令实现直观导航。具体而言,我们引入一种混合记忆系统:短期感知记忆用于追踪语义视觉线索,以弥补现有方法中视觉语言模型单帧推理的局限性;长期拓扑记忆则促进从过往经验中稳定学习策略。为填补现有基准的空白,我们还提出了 VLN-AVP 数据集与基准。该数据集包含10个高保真泊车场景和超过1000个导航任务,是目前规模最大的地下车库场景数据集,也是首个针对地下停车场的视觉语言导航基准。大量实验表明,在仿真环境中,该方法相比现有视觉语言导航方法成功率提升超过25%,相比其他自动驾驶方法提升超过15%。此外,在真实车辆实验中也取得了领先的成功率,验证了其实际可行性。
原文摘要 · Abstract (English)
Existing methods in Autonomous Valet Parking (AVP) typically rely on pre-built maps, which severely restricts their scalability to unseen environments and open-vocabulary targets. Inspired by the application of Vision-Language Models (VLMs) in Vision-Language Navigation (VLN) tasks, we propose VLN-AVP, a zero-shot navigation framework for AVP tasks. By combining the precise spatial perception of a Bird's-Eye-View (BEV) model with the general intelligence of VLMs, our framework 1) eliminates the dependency on pre-built maps, 2) interprets semantic environmental contexts in parking scenarios, and 3) enables intuitive navigation following natural language instructions. Specifically, we introduce a hybrid memory system: a short-term perception memory tracks semantic visual cues to address the limitations of VLM's single-frame reasoning in existing methods, while a long-term topological memory facilitates stable policy learning from past experiences. To bridge the gap in existing benchmarks, we also present the VLN-AVP dataset and benchmark. Featuring 10 high-fidelity parking scenes and over 1,000 navigation episodes, it has the largest number of garage scenes to date and is the first VLN benchmark for underground parking. Extensive experiments demonstrate that in simulation, our method achieves an over 25% improvement in success rate compared to VLN methods and an over 15% improvement compared to other autonomous driving methods. Furthermore, it attains a leading success rate in real-world vehicle experiments, proving its practical feasibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。