发现遥感图像多维度冗余,可大幅降低计算成本。
Hide and Seek: Investigating Redundancy in Earth Observation Imagery
- 分析遥感数据在光谱、时间、空间、语义上的多重冗余性。
- 利用冗余信息实现98.5%性能,计算量减少约4倍。
- 成果适用于多种任务与传感器,具普适性价值。
随着地球观测(EO)数据的快速增长和计算机视觉技术的进步,面向地球观测的机器学习模型规模迅速扩大。然而,这一进展可能忽略了遥感数据与其他领域本质不同的基本特性。本文提出,遥感数据具有显著的多维冗余(光谱、时间、空间、语义),其影响远超现有文献所反映的程度。通过系统性地考察该现象在关键变量维度上的存在性、一致性与实际影响,研究证实:遥感数据中的冗余不仅广泛存在,且可被有效利用——在训练和推理阶段,仅需约1/4的计算量(约4×更少的GFLOPs),即可达到基线性能的约98.5%。这些增益在不同任务、地理区域、传感器、地面采样距离及模型架构间均保持一致,表明多维冗余是遥感数据的结构性特征,而非特定实验设计的产物。本研究为构建更高效、可扩展、易访问的大规模遥感模型奠定了基础。
原文摘要 · Abstract (English)
The growing availability of Earth Observation (EO) data and recent advances in Computer Vision have driven rapid progress in machine learning for EO, producing domain-specific models at ever-increasing scales. Yet this progress risks overlooking fundamental properties of EO data that distinguish it from other domains. We argue that EO data exhibit a multidimensional redundancy (spectral, temporal, spatial, and semantic) which has a more pronounced impact on the domain and its applications than what current literature reflects. To validate this hypothesis, we conduct a systematic domain-specific investigation examining the existence, consistency, and practical implications of this phenomenon across key dimensions of EO variability. Our findings confirm that redundancy in EO data is both substantial and pervasive: exploiting it yields comparable performance ($\approx98.5\%$ of baseline) at a fraction of the computational cost ($\approx4\times$ fewer GFLOPs), at both training and inference. Crucially, these gains are consistent across tasks, geospatial locations, sensors, ground sampling distances, and architectural designs; suggesting that multi-faceted redundancy is a structural property of EO data rather than an artifact of specific experimental choices. These results lay the groundwork for more efficient, scalable, and accessible large-scale EO models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。