arXiv:2604.21032cs.CV2026-04中稿 · IGARSS 2026

让普通模型学会分析多光谱图像,无需重新训练即可显著提效

Unlocking Multi-Spectral Data for Multi-Modal Models with Guided Inputs and Chain-of-Thought Reasoning

论文配图:Unlocking Multi-Spectral Data for Multi-Modal Models with Guided Inputs and Chain-of-Thought Reasoning
图 1 · 摘自论文原文
  • 将多光谱数据转换为模型能理解的视觉空间,注入领域知识和推理指令
  • 在多个遥感基准上实现零样本性能大幅提升,最高提升超20%
  • 适合地理信息、环境监测等领域的专业人士快速用上大模型能力

多光谱图像是遥感应用(如土地利用分类与环境监测)中的宝贵输入信号,但通用大模型通常仅在RGB图像上训练,限制了其对多光谱数据的应用。同时,训练专用多光谱多模态模型成本高且结果高度定制化。为此,我们提出一种无需训练的新方法:在标准仅支持RGB的大型多模态模型(LMM)推理流程中引入多光谱数据,通过将非RGB输入适配至视觉空间,并注入领域特定信息与链式思维推理指令,实现性能显著提升。我们在Gemini 2.5模型上验证该方法,在多个主流遥感基准上均取得强零样本性能增益。结果表明,地理空间专业人士可借此借助强大通用模型处理专用传感器输入,获得基于专业数据的丰富推理能力。

原文摘要 · Abstract (English)

Multi-spectral imagery is a valuable input signal for Remote Sensing applications, such as land-use and land-cover classification and environmental monitoring. However, generalist Large Multi-modal Models (LMMs) are typically trained on RGB images, limiting their applicability to the RGB domain. At the same time, training multi-spectral multi-modal models is expensive and produces uniquely specialized models. To address this, we propose a novel training-free approach that introduces multi-spectral data within the inference pipeline of standard RGB-only LMMs, allowing large gains in performance. Our approach leverages the LMMs' understanding of the visual space by adapting non-RGB inputs to that space and injecting domain-specific information and Chain-of-Thought reasoning as instructions. We demonstrate this with the Gemini 2.5 model and observe strong Zero-Shot performance gains on popular Remote Sensing benchmarks. These results highlight the potential for geospatial professionals to leverage powerful generalist models for specialized sensor inputs, benefiting from rich reasoning capabilities grounded in specialized data.

多光谱大模型遥感推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。