用蒸馏方法实现轻量高效的位置编码,支持多源遥感数据统一建模。
SLED: Scalable Location Encoding via Distillation

- 通过地理坐标作为绑定模态,用蒸馏训练位置编码器。
- 小批量(128)即可训练,计算成本仅为现有方法的几分之一。
- 支持单模与多模融合,无需时空对齐,适合遥感研究者使用。
海量地理空间数据为学习高质量地球表征提供了机遇,但地球观测(EO)数据规模庞大、模态多样且传感器类型各异,带来显著挑战。位置编码器能高效将EO压缩为位置专属嵌入,但当前最优方法依赖计算昂贵的CLIP式框架,需16K–32K大批次,存在误负样本问题,且难以扩展至多模态。我们提出基于蒸馏的可扩展位置编码器SLED,以地理坐标为绑定模态,可利用任意地理空间数据模态预训练位置编码器。该框架轻量、模块化,灵活融合多模态,无需样本时空对齐。SLED仅需128小批量即可训练,预训练耗时与算力成本仅为现有模型的几分之一。我们在Sentinel-1、Sentinel-2和Landsat影像上预训练单模与多模SLED模型,结果表明其在19项人类相关基准任务中表现持平或优于现有方法,并验证了多模态预训练的优势。
原文摘要 · Abstract (English)
The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer size of the Earth Observations (EO), differing modalities, and different sensor types pose significant challenges in doing so. Location encoders have emerged as an efficient way of compressing EOs into location-specific embeddings. However, current state-of-the-art location encoders rely on computationally expensive CLIP-style frameworks that require large batch sizes in the 16K--32K range, suffer from false negative samples, and scale poorly with additional modalities. We introduce the Scalable Location Encoder via Distillation (SLED), a distillation-based location encoder that uses geospatial location as a binding modality to pretrain location encoders with any modality of geospatial data. The resulting location encoder framework is lightweight, modular, and can flexibly incorporate multiple modes, while eliminating the need for spatiotemporal coregistration of samples. SLED is performant with batch sizes as small as 128, enabling pretraining at a fraction of the runtime and compute costs of current state-of-the-art models. We demonstrate our approach by pretraining unimodal and multimodal SLED models on Sentinel-1, Sentinel-2, and Landsat imagery. We show that both unimodal and multimodal SLED models keep pace with or outperform existing approaches on a diverse set of 19 human-centric benchmark tasks and explore the benefits of using additional modes in pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。