arXiv:2512.17224cs.CV2025-12AAAI被引 2

通用遥感模型可适配任意波段与分辨率,解决卫星图像跨传感器泛化难题。

Any-Optical-Model: A Universal Foundation Model for Optical Remote Sensing

  • 设计谱无关编码器,独立建模各波段特征,支持任意波段组合
  • 多尺度自适应嵌入机制,覆盖亚米至百米级空间分辨率
  • 适合跨传感器、缺波段等复杂场景,尤其适合实际遥感部署

光学卫星因波段布局和地面采样距离多样,为生态系统监测到应急响应等任务提供关键数据。但不同传感器间波段配置与空间分辨率差异大,现有遥感基础模型(RSFMs)通常预训练于固定波段与分辨率,难以应对缺波段、跨传感器融合及未知分辨率等现实挑战,限制了泛化能力与实用部署。为此,我们提出通用光学模型(Any Optical Model, AOM),专为任意波段组合、传感器类型与分辨率尺度设计。为在波段缺失或新增时仍保留光谱特征,AOM引入谱无关令牌化器,为每通道分配独立波段嵌入,显式编码光谱身份。为有效捕捉从亚米级到百米级的纹理与上下文模式,设计多尺度自适应补丁嵌入机制,动态调节感受野。此外,结合多尺度语义对齐机制与通道级自监督掩码重建预训练策略,联合建模光谱-空间关系,维持跨分辨率下的全局语义一致性。在超过10个公开数据集(含Sentinel-2、Landsat、HLS)上的大量实验表明,AOM在缺波段、跨传感器、跨分辨率等挑战性设置下持续达到当前最优(SOTA)性能。

原文摘要 · Abstract (English)

Optical satellites, with their diverse band layouts and ground sampling distances, supply indispensable evidence for tasks ranging from ecosystem surveillance to emergency response. However, significant discrepancies in band composition and spatial resolution across different optical sensors present major challenges for existing Remote Sensing Foundation Models (RSFMs). These models are typically pretrained on fixed band configurations and resolutions, making them vulnerable to real world scenarios involving missing bands, cross sensor fusion, and unseen spatial scales, thereby limiting their generalization and practical deployment. To address these limitations, we propose Any Optical Model (AOM), a universal RSFM explicitly designed to accommodate arbitrary band compositions, sensor types, and resolution scales. To preserve distinctive spectral characteristics even when bands are missing or newly introduced, AOM introduces a spectrum-independent tokenizer that assigns each channel a dedicated band embedding, enabling explicit encoding of spectral identity. To effectively capture texture and contextual patterns from sub-meter to hundred-meter imagery, we design a multi-scale adaptive patch embedding mechanism that dynamically modulates the receptive field. Furthermore, to maintain global semantic consistency across varying resolutions, AOM incorporates a multi-scale semantic alignment mechanism alongside a channel-wise self-supervised masking and reconstruction pretraining strategy that jointly models spectral-spatial relationships. Extensive experiments on over 10 public datasets, including those from Sentinel-2, Landsat, and HLS, demonstrate that AOM consistently achieves state-of-the-art (SOTA) performance under challenging conditions such as band missing, cross sensor, and cross resolution settings.

遥感通用模型多尺度跨传感器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。