arXiv:2508.05202cs.CV2025-08中稿 · IEEE TGRS被引 6

首个专用于光谱遥感土地覆盖提取的视觉语言模型,融合光谱信息提升精度。

SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images

  • 构建含光谱先验的文本数据集SPIE,让LLM理解地物光谱特征。
  • 在5个公开多光谱数据集上,对植被、建筑、水体等类别均优于现有方法。
  • 可生成预测解释文本,适合需要可解释性的遥感应用开发者。

光谱信息在遥感观测中长期被视为关键线索,但现有视觉语言模型对光谱信息利用不足,尤其在多光谱场景下表现受限。为此,我们构建了名为SPIE的视觉语言指令遵循数据集,基于经典光谱指数计算,将地物的光谱先验编码为大语言模型可识别的文本属性。在此基础上,提出SPEX——一个面向指令驱动土地覆盖提取的多模态大语言模型。通过引入多尺度特征聚合、标记上下文压缩及多光谱视觉预训练等组件与训练策略,实现精准灵活的像素级解析。据我们所知,SPEX是首个专注于光谱遥感影像土地覆盖提取的多模态视觉语言模型。在五个公开多光谱数据集上的实验表明,SPEX在植被、建筑物和水体等典型地物提取任务中持续超越现有最优方法。此外,SPEX能生成预测的文本解释,显著提升可解释性与用户友好性。代码将于https://github.com/MiliLab/SPEX发布。

原文摘要 · Abstract (English)

Spectral information has long been recognized as a critical cue in remote sensing observations. Although numerous vision-language models have been developed for pixel-level interpretation, spectral information remains underutilized, resulting in suboptimal performance, particularly in multispectral scenarios. To address this limitation, we construct a vision-language instruction-following dataset named SPIE, which encodes spectral priors of land-cover objects into textual attributes recognizable by large language models (LLMs), based on classical spectral index computations. Leveraging this dataset, we propose SPEX, a multimodal LLM designed for instruction-driven land cover extraction. To this end, we introduce several carefully designed components and training strategies, including multiscale feature aggregation, token context condensation, and multispectral visual pre-training, to achieve precise and flexible pixel-level interpretation. To the best of our knowledge, SPEX is the first multimodal vision-language model dedicated to land cover extraction in spectral remote sensing imagery. Extensive experiments on five public multispectral datasets demonstrate that SPEX consistently outperforms existing state-of-the-art methods in extracting typical land cover categories such as vegetation, buildings, and water bodies. Moreover, SPEX is capable of generating textual explanations for its predictions, thereby enhancing interpretability and user-friendliness. Code will be released at: https://github.com/MiliLab/SPEX.

土地覆盖视觉语言模型遥感多光谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。