对比两款地理空间大模型,发现架构设计比模型本身更影响性能。
Now We Know? A Systematic Comparison of TerraMind and THOR

- 通过控制变量法分析两种模型在不同维度的差异
- patch大小和解码器类型解释了大部分性能差异
- 适合研究地理空间模型设计原理与评估方法的学者
地理空间基础模型(GFMs)的评测日益依赖综合得分排名,但这类排名掩盖了模型间差异的根本原因:架构、解码器容量还是特定任务因素?本研究通过受控比较欧洲航天局Φ-lab开发的两款设计哲学迥异的GFMs——THOR与TerraMind,揭示真实差异。THOR采用计算自适应架构,支持可变图像块尺寸,并在原始分辨率下统一处理哨兵-1、-2、-3数据;TerraMind是一种多模态生成式模型,通过双尺度标记/像素目标预训练,实现任意模态间的跨模态生成(思考模态),可在推理时推断缺失传感器。我们不在单一排行榜上评分,而是系统分析两个模型在图像块大小、解码器复杂度、微调策略、输入模态和模型规模等维度上的表现,覆盖包括气候灾害响应、甲烷泄漏检测、雪量监测和海冰制图在内的十项任务。结果表明,架构选择(特别是图像块大小与解码器类型)对性能方差的影响超过模型身份本身,两者代表互补策略(TerraMind侧重预训练尺度,THOR侧重推理时标记化),且正确解读结果需结合数据集特征。最终呈现的不是单个优胜者,而是一组假设与可泛化的诊断消融方法,适用于未来超越THOR与TerraMind的GFMs。
原文摘要 · Abstract (English)
Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a use-case-specific artefact? This study addresses that gap through a controlled comparison of two GFMs developed under European Space Agency's $Φ$-lab with contrasting design philosophies: THOR, which introduces a compute-adaptive architecture supporting variable patch sizes and unifies Sentinel-1, -2, and -3 data at their native resolutions; and TerraMind, a multimodal generative GFM pretrained with a dual-scale token/pixel objective that enables any-to-any cross-modal generation (Thinking-in-Modalities) to infer missing sensors at inference time. Rather than reporting a single leaderboard, we investigate the axes along which the two architectures actually differ - patch size, decoder complexity, finetuning regime, input modality, and model scale - across ten use cases spanning segmentation and regression in diverse domains, including climate disaster response, methane leak detection, snow monitoring, or sea ice mapping. We find that architectural design choices - patch size and decoder type in particular - explain more performance variance than model identity itself, that the two models embody complementary investment strategies (pretraining-time scale for TerraMind versus inference-time tokenisation for THOR), and that correctly interpreting results requires dataset-level characterisation. The resulting picture is not a single winner but a set of hypotheses and a diagnostic ablation methodology that we expect to generalise to future GFMs beyond THOR and TerraMind.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。