揭示地球观测嵌入的几何特性,提升智能环境推理能力
Characterizing AlphaEarth Embedding Geometry for Agentic Environmental Reasoning

- 分析64维嵌入的非欧几何结构,发现有效维度仅13.3
- 基于局部几何的检索比向量运算更准确,相关性达R²=0.32
- 构建九工具智能体系统,适配高阶推理模型使用
地球观测基础模型将地表信息编码为密集嵌入向量,但其表示的几何结构及其对下游推理的影响尚未被充分探索。本文分析谷歌AlphaEarth在1210万张美国大陆样本(2017–2023)上的64维嵌入的流形几何特性,并构建一个利用该几何理解的智能体系统以实现环境推理。该流形为非欧几里得结构:有效维度为13.3(参与度比),局部内在维度约10;切空间旋转显著,84%位置夹角超过60°,局部-全局对齐度(平均|cosθ|=0.17)接近随机基线0.125。监督线性探测显示概念方向随流形变化,使用PCA或探测方向进行组合向量运算精度较低。而检索则生成物理一致结果,局部几何可预测检索一致性(R²=0.32)。基于此,我们提出包含九个专用工具的智能体系统,通过FAISS索引嵌入数据库分解环境查询为推理链。五条件消融实验(120个查询,三难度层级)表明,嵌入检索主导响应质量(μ=3.79±0.90对比3.03±0.77参数化模型;量表1–5),在多步比较任务中表现最优(μ=4.28±0.43)。跨模型基准测试显示,几何工具使Sonnet 4.5得分下降0.12,但提升Opus 4.6得分0.07,且Opus展现更高几何对齐度(3.38 vs. 2.64),表明几何表征的价值随消费模型推理能力增强而上升。
原文摘要 · Abstract (English)
Earth observation foundation models encode land surface information into dense embedding vectors, yet the geometric structure of these representations and its implications for downstream reasoning remain underexplored. We characterize the manifold geometry of Google AlphaEarth's 64-dimensional embeddings across 12.1 million Continental United States samples (2017--2023) and develop an agentic system that leverages this geometric understanding for environmental reasoning. The manifold is non-Euclidean: effective dimensionality is 13.3 (participation ratio) from 64 raw dimensions, with local intrinsic dimensionality of approximately 10. Tangent spaces rotate substantially, with 84\% of locations exceeding 60\textdegree{} and local-global alignment (mean$|\cosθ| = 0.17$) approaching the random baseline of 0.125. Supervised linear probes indicate that concept directions rotate across the manifold, and compositional vector arithmetic using both PCA-derived and probe-derived directions yields poor precision. Retrieval instead produces physically coherent results, with local geometry predicting retrieval coherence ($R^2 = 0.32$). Building on this characterization, we introduce an agentic system with nine specialized tools that decomposes environmental queries into reasoning chains over a FAISS-indexed embedding database. A five-condition ablation (120 queries, three complexity tiers) shows that embedding retrieval dominates response quality ($μ= 3.79 \pm 0.90$ vs.\ $3.03 \pm 0.77$ parametric-only; scale 1--5), with peak performance on multi-step comparisons ($μ= 4.28 \pm 0.43$). A cross-model benchmark show that geometric tools reduce Sonnet 4.5's score by 0.12 points but improve Opus 4.6's by 0.07, with Opus achieving higher geometric grounding (3.38 vs.\ 2.64), suggesting that the value of geometric characterization scales with the reasoning capability of the consuming model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。