arXiv:2506.13542cs.CV2025-06被引 1

将遥感图像拆解为带上下文的数值,实现跨模态通用

Atomizer: Generalizing to new modalities by breaking satellite images down to a set of scalars

  • 把遥感图像转为带元信息的数值集合,统一处理不同模态
  • 在不重新训练的情况下,对多种分辨率和空间尺寸表现稳定
  • 适合需要灵活适配新卫星数据的研究者和应用

地球观测卫星数量激增,导致遥感数据在空间、光谱和时间配置上日益多样。现有模型多依赖固定输入格式和模态专用编码器,新配置出现时需重新训练,难以跨模态泛化。我们提出Atomizer,将遥感图像分解为一组标量,每个标量对应像素的光谱波段值,并附加获取时间、空间分辨率、波长和带宽等上下文元信息,形成原子化表示。单一编码器可处理任意模态,无需插值或重采样。Atomizer采用基于傅里叶特征和非均匀径向基函数的结构化标记化,通过交叉注意力将标记映射至潜在空间。在模态隔离评估中,其性能优于标准模型,并在不同分辨率和空间尺寸下保持稳健表现。

原文摘要 · Abstract (English)

The growing number of Earth observation satellites has led to increasingly diverse remote sensing data, with varying spatial, spectral, and temporal configurations. Most existing models rely on fixed input formats and modality-specific encoders, which require retraining when new configurations are introduced, limiting their ability to generalize across modalities. We introduce Atomizer, a flexible architecture that represents remote sensing images as sets of scalars, each corresponding to a spectral band value of a pixel. Each scalar is enriched with contextual metadata (acquisition time, spatial resolution, wavelength, and bandwidth), producing an atomic representation that allows a single encoder to process arbitrary modalities without interpolation or resampling. Atomizer uses structured tokenization with Fourier features and non-uniform radial basis functions to encode content and context, and maps tokens into a latent space via cross-attention. Under modality-disjoint evaluations, Atomizer outperforms standard models and demonstrates robust performance across varying resolutions and spatial sizes.

遥感跨模态原子表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。