arXiv:2608.16658cs.CVcs.AI2026-08中稿 · The 37th British M…

提出可渐进式定位的跨视角视频地理定位框架,支持动态观测与中断恢复。

X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization

论文配图:X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization
图 1 · 摘自论文原文
  • 分粒度对齐机制同步全局匹配与帧级细节匹配
  • 单帧下比最优方法提升4.7 Recall@1 Recall@10
  • 支持随机起点、长距离定位,贴近真实部署场景

跨视角视频地理定位(CVG)旨在通过检索对应的带地理标签航拍图像来定位地面视角视频。然而,现有方法依赖固定长度输入和事后优化,难以适应部分或动态观测下的在线定位需求。本文提出渐进式跨视角视频地理定位(PCVG),作为面向部署的扩展与评估协议,支持不同时间预算、前缀推理、随机起点评估及带中断的长距离定位。为此,我们提出X²Localizer框架,通过预算相关非对称目标,联合监督全局前缀-航拍图检索与帧-航拍块的聚合匹配。进一步引入滑动窗口重定位(SWRL)策略,动态刷新候选区域以实现失败恢复与长距离部署,无需全序列重处理。大量实验表明,X²Localizer在保持完整视频性能的同时,仅小幅提升+0.1 Recall@1和+0.3 Recall@10;但在单帧场景下,较先前最优方法显著提升+4.7 Recall@1和+11.5 Recall@10。结合SWRL,在随机起点与长距离场景中实现了鲁棒的渐进定位,缩小了基准评估与真实部署间的差距。

原文摘要 · Abstract (English)

Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. However, CVG approaches rely on fixed-length inputs and post-hoc refinement, hindering online-oriented localization under partial or dynamic observations. In this work, we formulate Progressive Cross-view Video Geo-localization (PCVG) as a deployment-oriented extension and evaluation protocol of CVG, enabling localization under varying temporal budgets, prefix-based inference, random-start evaluation, and long-range localization with interruptions. To explore PCVG, we introduce X$^2$Localizer, a cross-grained alignment framework that jointly supervises global prefix-to-aerial retrieval and token-aggregated frame--aerial-tile matching with a budget-dependent asymmetric objective. Furthermore, we introduce a Sliding-Window Re-Localization (SWRL) strategy that dynamically refreshes candidate regions for failure recovery and long-range deployment without full-sequence reprocessing. Extensive experiments show that X$^2$Localizer preserves conventional full-video performance, with marginal gains of +0.1 Recall@1 and +0.3 Recall@10, while substantially improving early localization. In the challenging single-frame setting, X$^2$Localizer improves coarse retrieval by +4.7 Recall@1 and +11.5 Recall@10 over the previous state-of-the-art method. With SWRL, our approach further enables robust progressive localization under random-start and long-distance scenarios, narrowing the gap between benchmark evaluation and real-world deployment.

地理定位视频检索渐进推理多粒度对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。