将卫星模型嵌入向量转化为可解释的环境智能,让大模型读懂地球表面数据。
Physically Interpretable AlphaEarth Foundation Model Embeddings Enable LLM-Based Land Surface Intelligence
- 用线性、非线性和注意力方法解析64维嵌入,发现每维对应具体地表特征。
- 12个环境变量预测精度超R²=0.90,温度和高程接近R²=0.97,跨时空稳定。
- 构建基于FAISS的问答系统,自然语言查询可生成有依据的环境评估报告。
卫星基础模型生成密集嵌入,但其物理可解释性尚不明确,限制其在环境决策系统中的应用。基于美国本土2017至2023年间1210万样本,我们首次对谷歌AlphaEarth的64维嵌入进行系统性可解释性分析,覆盖气候、植被、水文、温度和地形等26个环境变量。结合线性、非线性和注意力方法,结果显示各嵌入维度映射到特定地表属性,整体嵌入空间可高保真重建多数环境变量(12个变量的R² > 0.90;温度与高程达R² ≈ 0.97)。最强维度-变量关系在三种方法中一致,且在空间块交叉验证下保持稳定(平均ΔR² = 0.017),跨七年时间也高度一致(年均相关系数r = 0.963)。基于此,我们构建了地表智能系统,通过FAISS索引1210万向量嵌入库,实现自然语言查询到卫星数据支持的评估生成。使用四台LLM轮换生成、系统和评判角色的评估中,360次问答循环的加权得分μ = 3.74 ± 0.77(1–5分制),其中依据性(μ = 3.93)和连贯性(μ = 4.25)表现最优。结果表明,卫星基础模型嵌入是具有物理结构的表示,可转化为环境与地理空间智能的实用工具。
原文摘要 · Abstract (English)
Satellite foundation models produce dense embeddings whose physical interpretability remains poorly understood, limiting their integration into environmental decision systems. Using 12.1 million samples across the Continental United States (2017--2023), we first present a comprehensive interpretability analysis of Google AlphaEarth's 64-dimensional embeddings against 26 environmental variables spanning climate, vegetation, hydrology, temperature, and terrain. Combining linear, nonlinear, and attention-based methods, we show that individual embedding dimensions map onto specific land surface properties, while the full embedding space reconstructs most environmental variables with high fidelity (12 of 26 variables exceed $R^2 > 0.90$; temperature and elevation approach $R^2 = 0.97$). The strongest dimension-variable relationships converge across all three analytical methods and remain robust under spatial block cross-validation (mean $ΔR^2 = 0.017$) and temporally stable across all seven study years (mean inter-year correlation $r = 0.963$). Building on these validated interpretations, we then developed a Land Surface Intelligence system that implements retrieval-augmented generation over a FAISS-indexed embedding database of 12.1 million vectors, translating natural language environmental queries into satellite-grounded assessments. An LLM-as-Judge evaluation across 360 query--response cycles, using four LLMs in rotating generator, system, and judge roles, achieved weighted scores of $μ= 3.74 \pm 0.77$ (scale 1--5), with grounding ($μ= 3.93$) and coherence ($μ= 4.25$) as the strongest criteria. Our results demonstrate that satellite foundation model embeddings are physically structured representations that can be operationalized for environmental and geospatial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。