arXiv:2512.23903cs.CV2025-12被引 1

在千万亿像素级遥感数据上训练大模型,发现数据仍是瓶颈。

Scaling Remote Sensing Foundation Models: Data Domain Tradeoffs at the Peta-Scale

  • 用超大规模卫星影像数据训练视觉变换器,探索模型扩展规律。
  • 即使在千万亿像素规模下,性能仍受数据限制而非参数量制约。
  • 为遥感领域大模型研发提供数据与算力配置参考,适合遥感研究者。

我们研究人工智能的扩展特性,旨在建立在超过当前技术水平数个数量级的高分辨率光电(EO)数据集上训练基础模型的实用技术。现代多模态机器学习应用,如图像描述、搜索和推理的生成式AI系统,依赖于非文本模态的稳健且领域特化的编码器。在互联网规模数据丰富的自然图像领域,成熟的扩展定律可优化模型容量、训练计算量与数据规模的协同扩展。然而,在遥感等高价值领域,这些关系尚不明确。利用超过一千万亿像素的商业卫星光学数据及MITRE联邦AI沙盒,我们训练了逐步增大的视觉变换器(ViT)主干网络,报告了在千万亿规模下观察到的成功与失败模式,并分析了跨额外遥感模态弥合领域差距的启示。我们发现,即使在此规模下,性能仍处于数据受限而非参数量受限的范畴。这些实际洞察旨在指导数据采集策略、计算预算和优化调度,推动未来前沿规模遥感基础模型的发展。

原文摘要 · Abstract (English)

We explore the scaling behaviors of artificial intelligence to establish practical techniques for training foundation models on high-resolution electro-optical (EO) datasets that exceed the current state-of-the-art scale by orders of magnitude. Modern multimodal machine learning (ML) applications, such as generative artificial intelligence (GenAI) systems for image captioning, search, and reasoning, depend on robust, domain-specialized encoders for non-text modalities. In natural image domains where internet-scale data is plentiful, well-established scaling laws help optimize the joint scaling of model capacity, training compute, and dataset size. Unfortunately, these relationships are much less well understood in high-value domains like remote sensing (RS). Using over a quadrillion pixels of commercial satellite EO data and MITRE's Federal AI Sandbox, we train progressively larger vision transformer (ViT) backbones, report successes and failure modes observed at peta-scale, and analyze implications for bridging domain gaps across additional RS modalities. We observe that even at this scale, performance is consistent with a data-limited regime rather than a model parameter-limited one. These practical insights are intended to inform data collection strategies, compute budgets, and optimization schedules that advance the future development of frontier scale RS foundation models.

遥感大模型数据瓶颈视觉变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。