用双曲空间嵌入地理层级,实现高效精准的全球图像定位。
HierLoc: Hyperbolic Entity Embeddings for Hierarchical Visual Geolocation
- 将地理实体嵌入双曲空间,构建层次化定位框架
- 在OSV5M上误差降低19.5%,子区域准确率提升43%
- 仅需24万实体嵌入,比传统方法节省超95%存储
视觉地理定位任务因全球尺度、视觉模糊性及地理固有的层次结构而具有挑战性。现有方法或依赖大规模检索(需存储超500万图像嵌入),或采用网格分类器(忽略地理连续性),或使用生成模型(空间扩散但细节不足)。本文提出以地理实体为中心的定位范式,将图像直接对齐至国家、区域、次区域和城市等层级实体,通过融合大地测量距离的加权双曲对比学习实现。该层次设计支持可解释预测,且在OSV5M基准上仅需24万实体嵌入(远低于传统500万以上图像嵌入),实现新最佳性能:平均测地线误差降低19.5%,细粒度子区域准确率提升43%。结果表明,几何感知的层次嵌入为全球图像定位提供了可扩展且概念新颖的替代方案。
原文摘要 · Abstract (English)
Visual geolocalization, the task of predicting where an image was taken, remains challenging due to global scale, visual ambiguity, and the inherently hierarchical structure of geography. Existing paradigms rely on either large-scale retrieval, which requires storing a large number of image embeddings, grid-based classifiers that ignore geographic continuity, or generative models that diffuse over space but struggle with fine detail. We introduce an entity-centric formulation of geolocation that replaces image-to-image retrieval with a compact hierarchy of geographic entities embedded in Hyperbolic space. Images are aligned directly to country, region, subregion, and city entities through Geo-Weighted Hyperbolic contrastive learning by directly incorporating haversine distance into the contrastive objective. This hierarchical design enables interpretable predictions and efficient inference with 240k entity embeddings instead of over 5 million image embeddings on the OSV5M benchmark, on which our method establishes a new state-of-the-art performance. Compared to the current methods in the literature, it reduces mean geodesic error by 19.5\%, while improving the fine-grained subregion accuracy by 43%. These results demonstrate that geometry-aware hierarchical embeddings provide a scalable and conceptually new alternative for global image geolocation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。