arXiv:2606.08046cs.AIcs.CV2026-06被引 1

用开放地图数据学习全球地理表征,无需卫星图像

OSMGraphCLIP: Learning Global Location Representations from OpenStreetMap Graphs

论文配图:OSMGraphCLIP: Learning Global Location Representations from OpenStreetMap Graphs
图 1 · 摘自论文原文
  • 将地图要素建模为异构图,通过多尺度编码器捕捉局部与整体结构
  • 在多个领域任务上表现媲美甚至超越基于卫星的模型,尤其在社会经济和公共卫生任务中优势明显
  • 仅依赖地图拓扑即可准确恢复生物群落边界、城市梯度等地理规律,适合地理信息研究者

我们提出OSMGraphCLIP,一种类CLIP的地理空间表征模型,从免费的OpenStreetMap(OSM)数据中学习全局位置嵌入。该模型将地理环境表示为包含道路、建筑、用地区域和兴趣点等类型特征的异构图,保留其拓扑与语义关系。多尺度图编码器同时捕获细粒度局部结构与宏观景观组成,并通过对比对齐目标监督球谐函数位置编码器。我们在气候、生态、社会经济指标、公共健康、土地覆盖、生物多样性及野火预测等多个下游回归与分类任务上评估该模型,结果表明仅使用结构化OSM数据即可支持跨领域的强健全球位置表征。OSMGraphCLIP在多数基准上达到或超过基于卫星的基线性能,尤其在社会经济与公共健康任务中优势显著——因OSM对建成环境的显式语义标注能直接反映人类活动模式,而卫星像素仅能间接捕捉。在生态与环境任务中,尽管未使用任何地球观测数据,模型仍保持与影像方法相当的竞争力。定性分析证实,学习到的嵌入可有效组织地理空间,仅凭地图拓扑即恢复生物群落边界、城市梯度与热带-温带差异。

原文摘要 · Abstract (English)

We present OSMGraphCLIP, a CLIP-style geospatial representation model that learns global location embeddings from freely available OpenStreetMap (OSM) data. OSMGraphCLIP represents geographic environments as heterogeneous graphs of typed OSM features, preserving the topological and semantic relationships among roads, buildings, land-use regions, and points of interest. A multi-scale graph encoder captures both fine-grained local structure and broader landscape composition, and supervises a spherical-harmonics location encoder through a contrastive alignment objective. We evaluate OSMGraphCLIP across a diverse suite of downstream geospatial regression and classification tasks spanning climate, ecology, socioeconomic indicators, public health, land cover, biodiversity, and wildfire forecasting, and show that structured OSM data alone supports strong global location representations across domains. OSMGraphCLIP matches or exceeds satellite-based baselines on the majority of benchmarks, with the most pronounced advantage on socioeconomic and public-health tasks, where OSM's explicit semantic annotation of the built environment encodes patterns of human activity that satellite pixels can only capture indirectly. On ecological and environmental tasks, the model remains closely competitive with imagery-based methods despite using no Earth observation data. Qualitative analysis confirms that the learned embeddings organize geographic space coherently, recovering biome boundaries, urban gradients, and tropical--temperate distinctions from map topology alone.

地理表征开放地图图神经网络无卫星

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。