arXiv:2607.18504cs.LGcs.AI2026-07

对比两款地理空间大模型,发现架构设计比模型本身更影响性能。

Now We Know? A Systematic Comparison of TerraMind and THOR

论文配图:Now We Know? A Systematic Comparison of TerraMind and THOR
图 1 · 摘自论文原文
  • 通过控制变量法分析两种模型在不同维度的差异
  • patch大小和解码器类型解释了大部分性能差异
  • 适合研究地理空间模型设计原理与评估方法的学者

地理空间基础模型(GFMs)的评测日益依赖综合得分排名,但这类排名掩盖了模型间差异的根本原因:架构、解码器容量还是特定任务因素?本研究通过受控比较欧洲航天局Φ-lab开发的两款设计哲学迥异的GFMs——THOR与TerraMind,揭示真实差异。THOR采用计算自适应架构,支持可变图像块尺寸,并在原始分辨率下统一处理哨兵-1、-2、-3数据;TerraMind是一种多模态生成式模型,通过双尺度标记/像素目标预训练,实现任意模态间的跨模态生成(思考模态),可在推理时推断缺失传感器。我们不在单一排行榜上评分,而是系统分析两个模型在图像块大小、解码器复杂度、微调策略、输入模态和模型规模等维度上的表现,覆盖包括气候灾害响应、甲烷泄漏检测、雪量监测和海冰制图在内的十项任务。结果表明,架构选择(特别是图像块大小与解码器类型)对性能方差的影响超过模型身份本身,两者代表互补策略(TerraMind侧重预训练尺度,THOR侧重推理时标记化),且正确解读结果需结合数据集特征。最终呈现的不是单个优胜者,而是一组假设与可泛化的诊断消融方法,适用于未来超越THOR与TerraMind的GFMs。

原文摘要 · Abstract (English)

Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a use-case-specific artefact? This study addresses that gap through a controlled comparison of two GFMs developed under European Space Agency's $Φ$-lab with contrasting design philosophies: THOR, which introduces a compute-adaptive architecture supporting variable patch sizes and unifies Sentinel-1, -2, and -3 data at their native resolutions; and TerraMind, a multimodal generative GFM pretrained with a dual-scale token/pixel objective that enables any-to-any cross-modal generation (Thinking-in-Modalities) to infer missing sensors at inference time. Rather than reporting a single leaderboard, we investigate the axes along which the two architectures actually differ - patch size, decoder complexity, finetuning regime, input modality, and model scale - across ten use cases spanning segmentation and regression in diverse domains, including climate disaster response, methane leak detection, snow monitoring, or sea ice mapping. We find that architectural design choices - patch size and decoder type in particular - explain more performance variance than model identity itself, that the two models embody complementary investment strategies (pretraining-time scale for TerraMind versus inference-time tokenisation for THOR), and that correctly interpreting results requires dataset-level characterisation. The resulting picture is not a single winner but a set of hypotheses and a diagnostic ablation methodology that we expect to generalise to future GFMs beyond THOR and TerraMind.

地理模型多模态架构分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。