arXiv:2504.11171cs.CVcs.AI2025-04ICCV被引 126

首个地球观测多模态生成模型,支持任意模态互转。

TerraMind: Large-Scale Generative Multimodality for Earth Observation

论文配图:TerraMind: Large-Scale Generative Multimodality for Earth Observation
图 1 · 摘自论文原文
  • 双尺度预训练:融合像素级与标记级信息,捕捉空间细节与上下文关系。
  • 零样本/少样本应用广泛,基准测试表现超越现有方法。
  • 支持生成人工数据提升输出质量,适合遥感与地理研究者使用。

我们提出TerraMind,首个用于地球观测(EO)的任意模态间生成式多模态基础模型。不同于其他多模态模型,TerraMind在双尺度表示上进行预训练,结合跨模态的标记级与像素级数据。标记级编码高层语义信息以学习跨模态关系,像素级则利用细粒度表征捕捉关键空间特征。TerraMind在涵盖九种地理空间模态的全球大规模数据集上进行预训练。本文证明:(i) TerraMind的双尺度早期融合策略可实现多种地球观测的零样本与少样本应用;(ii) 引入“模态思维”(Thinking-in-Modalities, TiM)——在微调与推理阶段生成额外人工数据以提升输出质量;(iii) 在社区标准基准如PANGAEA上达到超越现有水平的表现。预训练数据集、模型权重及代码已开源,采用宽松许可协议。

原文摘要 · Abstract (English)

We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO). Unlike other multimodal models, TerraMind is pretrained on dual-scale representations combining both token-level and pixel-level data across modalities. On a token level, TerraMind encodes high-level contextual information to learn cross-modal relationships, while on a pixel level, TerraMind leverages fine-grained representations to capture critical spatial nuances. We pretrained TerraMind on nine geospatial modalities of a global, large-scale dataset. In this paper, we demonstrate that (i) TerraMind's dual-scale early fusion approach unlocks a range of zero-shot and few-shot applications for Earth observation, (ii) TerraMind introduces "Thinking-in-Modalities" (TiM) -- the capability of generating additional artificial data during finetuning and inference to improve the model output -- and (iii) TerraMind achieves beyond state-of-the-art performance in community-standard benchmarks for EO like PANGAEA. The pretraining dataset, the model weights, and our code are open-sourced under a permissive license.

地球观测多模态生成遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。