让遥感模型读懂城市经济,提升跨城跨层级预测能力
AEF-Econ: Toward Plug-and-Play Socioeconomic Foundation Embeddings from AlphaEarth for Urban Remote Sensing
- 用多源数据融合与自适应重建机制增强遥感模型的社会经济理解力
- 跨城预测R2从0.301提升至0.848,跨层级达0.693,显著突破原有瓶颈
- 适合做城市规划、区域经济分析的科研人员和政策制定者使用
AlphaEarth基础模型(AEF)通过多模态自监督学习统一全球遥感嵌入,但其预训练聚焦物理地表信号,难以直接用于社会经济任务。本文整合中国36个城市8年间的七类异构数据——AEF嵌入、人口、夜间灯光、遥感指数、兴趣点(POIs)、城市形态及跨语言文本,构建了包含16个标签的CHN-Econ社会经济基准。在五个维度(融合架构、自监督目标、文本集成、嵌入维度、归一化)上开展31组受控实验。仅以线性探测器使用时,AEF在跨区域和跨层级评估中分别仅达R2 0.301和0.160;经五轴消融优化后,提升至0.832和0.671,但发现低维语义流常被高维流压制。为此提出容量自适应重建(CAR),以各流独立解码器与损失替代共享重建,缓解流间容量竞争。CAR进一步将跨区域和跨层级R2提升至0.848和0.693,并恢复负R2崩溃标签至稳定范围。基于CAR,推断36城8年共1440万像素,发布包含128d与64d压缩版本的AEF-Econ。自诊断与案例研究显示,该模型在无监督下捕捉跨城层级结构与城内空间组织,为补充原版物理嵌入提供社会经济遥感基础嵌入。
原文摘要 · Abstract (English)
AlphaEarth Foundations (AEF) unify global remote sensing foundation embeddings through multimodal self-supervised learning, but their pretraining focuses on physical land-surface signals, limiting plug-and-play use in socioeconomic tasks. We integrate seven heterogeneous data streams across 36 Chinese cities over eight years - AEF embeddings, population, nighttime lights, remote sensing indices, points of interest (POIs), urban morphology, and cross-lingual text - and construct CHN-Econ, a socioeconomic benchmark with 16 labels in three categories. We conduct 31 controlled experiments along five axes: fusion architecture, self-supervised objective, text integration, embedding dimensionality, and normalization. Used alone as a linear probe, AEF achieves R2 values of only 0.301 for cross-region and 0.160 for cross-tier evaluation. The five-axis ablated backbone improves these scores to 0.832 and 0.671, respectively, but reveals that low-dimensional semantic streams are consistently suppressed by high-dimensional streams under shared reconstruction. To address this bottleneck, we propose Capacity-Adaptive Reconstruction (CAR), replacing shared reconstruction with per-stream decoders and stream-level losses to mitigate inter-stream capacity competition. CAR further raises cross-region and cross-tier R2 to 0.848 and 0.693, and restores collapsed labels from negative R2 to a stable range. Using CAR, we infer 14.4 million pixels across 36 cities and eight years and release AEF-Econ, including 128d and 64d compressed versions. Self-diagnostics and case studies show that AEF-Econ captures cross-city hierarchies and intra-urban spatial organization under unsupervised settings, providing a socioeconomic remote sensing foundation embedding complementary to AEF physical embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。