arXiv:2509.22697cs.CVcs.AI2025-09中稿 · ICCV被引 6

用精选提示词提升高光谱图像与文本的对齐效率

Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment

  • 用对比学习框架,仅0.07%参数微调实现跨模态对齐
  • 在印第安纳波利斯和帕维亚大学数据集上准确率提升超1%以上
  • 适合资源有限但需高效多模态理解的高光谱场景应用

随着数据需求增长,高效学习愈发依赖高价值数据的筛选与提炼,而非盲目扩大模型规模。高光谱图像(HSI)因具有高维3D体素结构(每个空间位置对应数百个连续光谱通道),使得跨模态对齐面临更大挑战。尽管视觉语言模型在自然图像或文本任务中表现优异,但在高光谱领域仍属未充分探索方向。本文提出一种基于CLIP风格对比训练的视觉语言模型优化方法,将视觉骨干网络生成的体素级嵌入映射至冻结的大规模嵌入模型(LEM)潜在空间,通过可训练探测器使视觉特征与文本标记表示对齐。采用精选的难负样本(最接近的错误类别)与半难负样本(随机干扰项)及正样本对进行对比损失训练,并引入描述性提示作为语义锚点以增强对齐效果。实验表明,该方法仅更新0.07%参数,即达到领先性能:在印度波因斯(IP)数据集上,整体准确率(OA)提升0.92,卡帕系数(κ)提升1.60;在帕维亚大学(PU)数据集上,分别提升0.69和0.90。模型参数量仅为DCTN的1/50、SS-TMNet的1/90。

原文摘要 · Abstract (English)

As data requirements continue to grow, efficient learning increasingly depends on the curation and distillation of high-value data rather than brute-force scaling of model sizes. In the case of a hyperspectral image (HSI), the challenge is amplified by the high-dimensional 3D voxel structure, where each spatial location is associated with hundreds of contiguous spectral channels. While vision and language models have been optimized effectively for natural image or text tasks, their cross-modal alignment in the hyperspectral domain remains an open and underexplored problem. In this article, we make an attempt to optimize a Vision-Language Model (VLM) for hyperspectral scene understanding by exploiting a CLIP-style contrastive training framework. Our framework maps voxel-level embeddings from a vision backbone onto the latent space of a frozen large embedding model (LEM), where a trainable probe aligns vision features with the model's textual token representations. The two modalities are aligned via a contrastive loss restricted to a curated set of hard (closest wrong classes) and semi-hard (random distractors) negatives, along with positive pairs. To further enhance alignment, descriptive prompts that encode class semantics are introduced and act as structured anchors for the HSI embeddings. It is seen that the proposed method updates only 0.07 percent of the total parameters, yet yields state-of-the-art performance. For example, on Indian Pines (IP) the model produces better results over unimodal and multimodal baselines by +0.92 Overall Accuracy (OA) and +1.60 Kappa ($κ$), while on Pavia University (PU) data it provides gains of +0.69 OA and +0.90 $κ$. Moreover, this is achieved with the set of parameters, nearly 50$\times$ smaller than DCTN and 90$\times$ smaller than SS-TMNet.

高光谱视觉语言模型跨模态对齐小参数微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。