通过时空上下文建模,提升自动驾驶中跨视角视觉定位的精度与鲁棒性。
Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving
- 引入循环跨帧模块,利用历史帧信息增强当前帧特征
- 在CVIS上将定位误差从3.80米降至1.57米,R@1米提升至40.22%
- 适用于真实道路场景,零样本部署下误差仅2.84米
连续可靠的定位对自动驾驶至关重要。跨视角视觉定位通过匹配地面图像与卫星地图,为依赖全球导航卫星系统(GNSS)和高精地图的系统提供互补定位线索。现有方法多独立处理每帧图像,忽视了时序信息,在动态遮挡、光照变化和重复纹理条件下性能受限。本文提出一种增强时序上下文的跨视角序列视觉定位框架。通过递归跨帧模块聚合前序状态的历史上下文,增强当前帧的粗粒度地面特征,进而提升卫星候选区域分类能力;同时利用分层细粒度特征实现精确局部偏移估计。在CVIS数据集上,平均定位误差由3.80米降至1.57米,R@1米从8.14%提升至40.22%。直接迁移至KITTI-CVL后,平均误差为2.61米,领域微调后进一步降至2.27米。零样本实地测试中,平均误差为2.84米,R@5米达96.86%。结果表明,时序上下文增强显著提升了跨视角定位精度,并支持在公开基准与真实道路环境中的稳健部署。
原文摘要 · Abstract (English)
Continuous and reliable localization is essential for autonomous driving. Cross-view visual localization matches ground images with satellite maps, providing complementary localization cues for pipelines that depend on Global Navigation Satellite System (GNSS) signals and high-definition (HD) maps. Most existing cross-view visual localization methods process each frame independently, leaving temporal information underused and limiting accuracy under dynamic occlusion, illumination variation, and repetitive textures. This study proposes a temporal-context-enhanced framework for cross-view sequence visual localization. The proposed recurrent cross-frame module aggregates historical context from the previous state to enhance the coarse ground feature of each current frame. These enhanced features facilitate satellite candidate-region classification, while hierarchical fine-grained features enable precise local offset estimation. On the CVIS dataset, the proposed method reduces mean localization error from 3.80 m to 1.57 m and increases R@1 m from 8.14% to 40.22%. Direct transfer to KITTI-CVL achieves a mean error of 2.61 m, with target-domain fine-tuning further reducing the mean error to 2.27 m. Zero-shot field experiments on a real-world vehicle achieve a mean error of 2.84 m and R@5 m of 96.86%. These results demonstrate that temporal context enhancement significantly improves cross-view localization accuracy and supports robust deployment on public benchmarks and real-world roads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。