arXiv:2607.15491cs.CV2026-07中稿 · ECCV

用路线描述和视频联合定位,提升导航精度。

Trajectory-aware Cross-view Geo-localization with Sequential Observations

论文配图:Trajectory-aware Cross-view Geo-localization with Sequential Observations
图 1 · 摘自论文原文
  • 融合视频与路线文本,实现跨视角地理定位
  • 在39000组数据上,定位准确率显著超越现有方法
  • 适合自动驾驶、智能导航等场景使用

跨视图地理定位旨在将地面观测图像与带地理标签的卫星影像匹配。近期方法表明,视频片段等序列查询比单张图像蕴含更丰富的时空线索,但忽略了另一种互补的序列模态——路线描述:它以更高抽象层次刻画同一轨迹,且常为唯一可用输入(如用户指导自动驾驶车辆至接驳点)。为此,我们构建了包含约39,000个视频-文本-卫星三元组的SeqGeo-VL数据集,并提出TrajLoc统一框架,可同时处理视频片段与路线描述。通过融合密集视觉语义与抽象语言语义,使两类模态相互增强匹配效果。我们进一步设计轻量级模块TrajMod,基于轨迹几何条件化查询嵌入,生成空间感知表示。实验表明,TrajLoc在视频与文本地理定位任务上均显著优于当前最优方法。

原文摘要 · Abstract (English)

Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods show that sequential queries such as video clips yield richer spatiotemporal cues than single images, yet they overlook a complementary sequential modality: route descriptions -- which capture the same trajectory at a higher level of abstraction and are often the only input available (e.g., a user directing an autonomous vehicle to a pickup point). To bridge this gap, we introduce SeqGeo-VL, a dataset of $\sim$39K video-text-satellite triplets, and TrajLoc, a unified framework capable of processing both video clips and route descriptions. By leveraging both dense visual and abstract linguistic semantics, TrajLoc enables these modalities to mutually reinforce cross-view matching. We further propose TrajMod, a lightweight module that conditions query embeddings on trajectory geometry, yielding spatially-aware representations. Experiments show that TrajLoc achieves substantial gains over state-of-the-art methods on both video and text geo-localization. The project page is available at https://humblegamer.github.io/trajloc/.

地理定位多模态自动驾驶序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。