让地球遥感模型理解城市功能,用兴趣点对齐人机语义。
Beyond AlphaEarth: Toward Human-Centered Geospatial Foundation Models via POI-Guided Contrastive Learning
- 用兴趣点引导跨模态对齐,融合遥感数据与城市语义。
- 在伦敦和新加坡测试中,任务表现提升4.5%至21.9%。
- 支持自然语言查询地理空间位置,适合城市研究者使用。
近期的地理空间基础模型(GFMs)生成了覆盖全球的地表丰富表示,捕捉了物理与环境模式。其中,AlphaEarth Foundation(AE)从多源地球观测(EO)数据生成10米分辨率嵌入,涵盖多样环境与光谱特征。但此类基于EO的表示主要编码物理与光谱模式,缺乏人类活动或城市语义信息,难以捕捉城市功能维度,也限制了自然语言交互与解释能力。本文提出AETHER(AlphaEarth-POI Enriched Representation Learning),一个轻量级框架,通过兴趣点(POIs)引导的多模态对齐,将AE与以人为中心的城市分析对齐。通过强制执行跨模态AE-POI对齐与模态内多尺度一致性,AETHER融合功能城市语义与EO驱动表示,并将嵌入空间锚定于自然语言。结果表示支持城市制图任务与自然语言条件下的空间检索。在伦敦与新加坡的四个下游任务中,性能相对提升4.5%至21.9%。此外,对齐的嵌入空间可实现自然语言驱动的空间定位。该工作推动地理空间表示学习向以人为本、语言可访问的方向发展。
原文摘要 · Abstract (English)
Recent geospatial foundation models (GFMs) produce spatially extensive representations of the Earth's surface that capture rich physical and environmental patterns. Among them, the AlphaEarth Foundation (AE) represents a major step, generating 10 m embeddings from multi-source Earth Observation (EO) data that include diverse environmental and spectral characteristics. However, such EO-driven representations primarily encode physical and spectral patterns rather than human activities or urban semantics, limiting their ability to capture the functional dimensions of cities and making the learned representations difficult to interpret or query using natural language. We introduce AETHER (AlphaEarth-POI Enriched Representation Learning), a lightweight framework that aligns AlphaEarth with human-centered urban analysis through multimodal alignment guided by Points of Interest (POIs). By enforcing both cross-modal AE-POI alignment and intra-modal multi-scale consistency, AETHER integrates functional urban semantics with EO-driven representations and grounds the embedding space in natural language. The resulting representations support both urban mapping tasks and natural language-conditioned spatial retrieval. Experiments across four downstream tasks in Greater London and Singapore demonstrate consistent state-of-the-art performance, with relative improvements ranging from 4.5% to 21.9%. Furthermore, the aligned embedding space enables spatial localization through natural language queries. By aligning EO-based foundation models with human-centered semantics, AETHER improves the interpretability of geospatial representations and advances geospatial representation learning toward human-centered, language-accessible geospatial foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。