arXiv:2605.08633cs.DCcs.CV2026-05

用历史数据训练生成压缩模型,实现地球观测数据100到10000倍压缩。

Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction

论文配图:Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction
图 1 · 摘自论文原文
  • 基于历史观测数据学习先验,构建可生成的超高压缩模型。
  • 在超算上实现1.54 EFLOP/s持续训练,最高达2.16 EFLOP/s。
  • 适合需要海量地球观测数据处理的科研与遥感应用。

地球观测正成为科学领域最大的数据生产活动之一,但现有流程仍将压缩视为存储和传输工具,而非数据新用法。本文提出一种生成式压缩框架,从历史地球观测档案中学习,并支持下游任务中100至10,000倍的数据压缩。与通用视觉数据不同,地球观测反复记录同一演化中的行星,使历史先验学习成为极端压缩的基础。为实现该范式,我们在LineShine Armv9 CPU超算上以全栈协同优化方式训练大规模生成压缩模型,涵盖模型设计、内核、内存层次、运行时与并行性。实现端到端训练中1.54 EFLOP/s持续性能,峰值达2.16 EFLOP/s。本工作表明,基于历史先验的生成压缩可将地球观测数据转化为可按需适应任务的采集、传输、存储与科研基础。

原文摘要 · Abstract (English)

Earth observation is becoming one of the largest data-producing activities in science, yet current pipelines still treat compression as a storage and transmission tool rather than a new way to use data. We present a generative compression framework that learns from historical Earth observation archives and enables on-demand 100x to 10,000x data reduction across downstream tasks. Unlike general visual data, Earth observation repeatedly measures the same evolving planet, making historical-prior learning feasible for extreme compression. To realize this paradigm, we train large generative compression models at exascale on the LineShine Armv9 CPU supercomputer, with co-optimization across model design, kernels, memory hierarchy, runtime, and parallelism. Our implementation sustains 1.54 EFLOP/s and peaks at 2.16 EFLOP/s in end-to-end training. This work shows that historical-prior generative compression can turn Earth observation data into an active, task-adaptive foundation for acquisition, delivery, storage, and scientific use.

生成压缩地球观测超算训练数据压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。