arXiv:2509.19087cs.CV2025-09被引 2

让通用多模态模型零样本理解遥感多光谱数据,无需训练即可提升分类性能。

Zero-Shot Multi-Spectral Learning: Reimagining a Generalist Multimodal Gemini 2.5 Model for Remote Sensing Applications

  • 将多光谱数据转换为视觉空间表示,通过指令注入实现零样本适配。
  • 在土地覆盖分类等基准上表现显著优于传统方法,性能接近专用模型。
  • 适合遥感从业者快速利用强大模型能力,处理非标准传感器数据。

多光谱影像在土地利用分类、环境监测和城市规划等遥感应用中至关重要,因其额外波段与地表物质(如冰、水、植被)有强关联,能更准确识别目标。该类数据广泛来自哨兵-2和陆地卫星等任务,具有高价值。当前自动分析主要依赖针对多光谱输入专门训练的模型,成本高昂且难以扩展。尽管多光谱数据极具实用价值,却无法直接用于强大的通用多模态大模型(如基于RGB训练的Gemini2.5),这些模型虽具备丰富推理与上下文理解能力,但无法解析特殊多光谱信号。为此,本文提出一种无需训练的零样本方法:将多光谱数据以零样本模式输入仅接受RGB的通用多模态模型,通过视觉空间对齐并注入领域指令,使模型理解新输入。实验表明,该方法在多个主流遥感土地覆盖与土地利用分类基准上取得显著零样本性能提升,展现出Gemini2.5对新输入的易适配性。结果表明,地理空间专业人员可轻松借助此类强大模型加速工作,利用其丰富的推理能力,结合专用传感器数据实现高效分析。

原文摘要 · Abstract (English)

Multi-spectral imagery plays a crucial role in diverse Remote Sensing applications including land-use classification, environmental monitoring and urban planning. These images are widely adopted because their additional spectral bands correlate strongly with physical materials on the ground, such as ice, water, and vegetation. This allows for more accurate identification, and their public availability from missions, such as Sentinel-2 and Landsat, only adds to their value. Currently, the automatic analysis of such data is predominantly managed through machine learning models specifically trained for multi-spectral input, which are costly to train and support. Furthermore, although providing a lot of utility for Remote Sensing, such additional inputs cannot be used with powerful generalist large multimodal models, which are capable of solving many visual problems, but are not able to understand specialized multi-spectral signals. To address this, we propose a training-free approach which introduces new multi-spectral data in a Zero-Shot-only mode, as inputs to generalist multimodal models, trained on RGB-only inputs. Our approach leverages the multimodal models' understanding of the visual space, and proposes to adapt to inputs to that space, and to inject domain-specific information as instructions into the model. We exemplify this idea with the Gemini2.5 model and observe strong Zero-Shot performance gains of the approach on popular Remote Sensing benchmarks for land cover and land use classification and demonstrate the easy adaptability of Gemini2.5 to new inputs. These results highlight the potential for geospatial professionals, working with non-standard specialized inputs, to easily leverage powerful multimodal models, such as Gemini2.5, to accelerate their work, benefiting from their rich reasoning and contextual capabilities, grounded in the specialized sensor data.

遥感多模态零样本Gemini

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。