arXiv:2608.29553cs.LG2026-08中稿 · ACM SIGSPATIAL 202…

用行为与语义信息增强地理模型,提升城市人类活动预测能力

BEACON: Behavioral and Semantic Enrichment of AlphaEarth Embeddings through Tri-Modal Contrastive Learning

论文配图:BEACON: Behavioral and Semantic Enrichment of AlphaEarth Embeddings through Tri-Modal Contrastive Learning
图 1 · 摘自论文原文
  • 三模态对比学习融合影像、文本和人流数据
  • 对肥胖率等社会指标预测提升最高达43%
  • 适合做城市规划与公共健康分析的研究者

地理空间基础模型如AlphaEarth能生成紧凑且全局一致的地球表面表征,有效迁移至多种下游任务。然而,由于主要基于地球观测影像训练,其嵌入主要捕捉物理与光谱特征,对人类活动和城市功能的编码较弱。为解决此问题,我们提出BEACON,一种三模态对比学习框架,对齐三种互补的城市空间视图:来自AE嵌入的物理表征、来自兴趣点(POI)文本的语义表征,以及来自小时级POI访问记录的人类行为表征,同时保持部署时仅使用图像输入。以休斯顿都会区为例,我们在九个下游任务上评估该框架,包括七个回归和两个分类任务,对比六种基线方法(原始坐标、Space2Vec、SatCLIP、TESSERA、Clay和AlphaEarth),采用冻结的线性与MLP探测器,在五个种子下进行测试。在线性探测下,BEACON相较AlphaEarth在肥胖患病率预测上相对R²提升高达43%,心理状态较差预测提升34%,家庭收入中位数预测提升22%,同时在物理与环境变量预测上仍具竞争力。结果表明,通过融入语义与行为信号,可显著扩展地理空间基础模型在人类中心城市分析中的应用价值。

原文摘要 · Abstract (English)

Geospatial foundation models such as the AlphaEarth Foundation produce compact and globally consistent representations of the Earth's surface that transfer effectively to a wide range of downstream tasks. However, because these models are trained primarily on Earth-observation imagery, their embeddings mainly capture physical and spectral characteristics while encoding human activity and urban function only weakly. To address this limitation, we propose BEACON, a tri-modal contrastive learning framework that aligns three complementary views of urban space: physical representations from AE embeddings, semantic representations from point-of-interest (POI) text, and human behavioral representations from hourly POI visitation, while keeping the deployed representation image-only. Using the Houston Metropolitan Area as a case study area, we evaluated the performance of the BEACON framework on nine downstream tasks, including seven regression and two classification tasks against six baselines (raw coordinates, Space2Vec, SatCLIP, TESSERA, Clay and AlphaEarth), using frozen linear and MLP probes over five seeds. Under a linear probe, BEACON improves relative R^2 over AlphaEarth by up to 43% for obesity prevalence, 34% for poor mental health, and 22% for median household income, while remaining competitive in the prediction of physical and environmental variables. These findings highlight the value of augmenting geospatial foundation models with semantic and behavioral signals, extending their applicability from physical Earth observation to human-centered urban analytics.

地理模型三模态学习城市分析对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。