arXiv:2512.06865cs.CV2025-12被引 6

用地图图像增强自动驾驶感知,解决视野受限问题。

Spatial Retrieval Augmented Autonomous Driving

论文配图:Spatial Retrieval Augmented Autonomous Driving
图 1 · 摘自论文原文
  • 引入离线地图图像作为额外输入,提升环境记忆能力。
  • 在5项核心任务中,部分任务性能显著提升。
  • 无需额外传感器,可直接接入现有自动驾驶系统。

现有自动驾驶系统依赖车载传感器(如摄像头、激光雷达、惯性单元等)进行环境感知,但受限于实时感知范围,在视线受阻、遮挡或极端天气(如黑暗、雨天)下表现不佳。人类驾驶员却能在能见度低时回忆道路结构。为赋予模型类似‘回忆’能力,我们提出空间检索范式,将离线获取的地理图像作为额外输入。这些图像可从离线缓存(如谷歌地图或自动驾驶数据集)中获取,无需额外传感器,可直接集成到现有自动驾驶任务中。实验中,我们基于谷歌地图API扩展了nuScenes数据集,添加地理图像并对其与自车轨迹对齐。建立了涵盖物体检测、在线建图、占用预测、端到端规划和生成式世界建模的五个核心任务基准。大量实验表明,新增模态可有效提升部分任务性能。相关数据集构建代码、数据及基准将开源,推动该新范式的进一步研究。

原文摘要 · Abstract (English)

Existing autonomous driving systems rely on onboard sensors (cameras, LiDAR, IMU, etc) for environmental perception. However, this paradigm is limited by the drive-time perception horizon and often fails under limited view scope, occlusion or extreme conditions such as darkness and rain. In contrast, human drivers are able to recall road structure even under poor visibility. To endow models with this ``recall" ability, we propose the spatial retrieval paradigm, introducing offline retrieved geographic images as an additional input. These images are easy to obtain from offline caches (e.g, Google Maps or stored autonomous driving datasets) without requiring additional sensors, making it a plug-and-play extension for existing AD tasks. For experiments, we first extend the nuScenes dataset with geographic images retrieved via Google Maps APIs and align the new data with ego-vehicle trajectories. We establish baselines across five core autonomous driving tasks: object detection, online mapping, occupancy prediction, end-to-end planning, and generative world modeling. Extensive experiments show that the extended modality could enhance the performance of certain tasks. We will open-source dataset curation code, data, and benchmarks for further study of this new autonomous driving paradigm.

自动驾驶空间检索地图融合多模态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。