用多光谱+语言对齐,让遥感图像描述更准更细。
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
- 融合多光谱数据与语言模型,用轻量投影对齐视觉与语义特征。
- 在BigEarthNet v2上提升分类与描述准确率,尤其在云雪等复杂场景中表现显著。
- 适合遥感、环境监测领域研究者,提升复杂场景理解能力。
遥感场景理解常因复杂环境(如不同用地类型或沿海区域)及天气干扰(如雪、云、霾)而面临挑战。为此,我们提出Spectral LLaVA框架,将多光谱数据与视觉-语言对齐技术结合,以增强场景表征与描述能力。基于Sentinel-2的BigEarthNet v2数据集,我们建立以RGB为基础的场景描述基线,并进一步展示引入多光谱信息后性能的显著提升。框架通过优化轻量级线性投影层实现对齐,同时冻结SpectralGPT的视觉主干。实验涵盖线性探测下的场景分类以及联合进行分类与描述生成的语言建模任务。结果表明,Spectral LLaVA能生成更详细准确的描述,尤其在仅依赖RGB数据难以处理的场景中;同时通过将SpectralGPT特征提炼为语义有意义的表示,提升了分类性能。
原文摘要 · Abstract (English)
Scene understanding in remote sensing often faces challenges in generating accurate representations for complex environments such as various land use areas or coastal regions, which may also include snow, clouds, or haze. To address this, we present a vision-language framework named Spectral LLaVA, which integrates multispectral data with vision-language alignment techniques to enhance scene representation and description. Using the BigEarthNet v2 dataset from Sentinel-2, we establish a baseline with RGB-based scene descriptions and further demonstrate substantial improvements through the incorporation of multispectral information. Our framework optimizes a lightweight linear projection layer for alignment while keeping the vision backbone of SpectralGPT frozen. Our experiments encompass scene classification using linear probing and language modeling for jointly performing scene classification and description generation. Our results highlight Spectral LLaVA's ability to produce detailed and accurate descriptions, particularly for scenarios where RGB data alone proves inadequate, while also enhancing classification performance by refining SpectralGPT features into semantically meaningful representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。