arXiv:2503.21841cs.CV2025-03CVPR被引 57

无需微调的高光谱遥感基础模型,适配不同通道数图像。

HyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing Imagery

  • 通过可学习的权重字典动态构建嵌入层,支持任意通道数输入。
  • 仅用1个提示在11个数据集上达到专用模型5次采样效果。
  • 适合需要快速部署、少资源的高光谱图像分析场景。

高光谱遥感图像的高级解析对精准地球观测任务至关重要。近期视觉基础模型虽推动了遥感解译发展,但主要聚焦于RGB和多光谱图像。由于高光谱通道数量差异大,现有基础模型需逐图微调,极大增加硬件与时间开销。本文提出无需微调的高光谱基础模型HyperFree,基于视觉提示工程改进。为处理不同通道数,设计覆盖0.4~2.5 μm全波段的可学习权重字典,实现嵌入层动态构建;为提升提示可操作性,将特征距离视为语义相似度,生成多个语义感知掩码。在构建的大规模高分辨率高光谱图像上预训练后,HyperFree(仅1个提示)在5项任务、11个数据集上表现媲美专用模型(5次采样)。代码与数据集见https://rsidea.whu.edu.cn/hyperfree.htm。

原文摘要 · Abstract (English)

Advanced interpretation of hyperspectral remote sensing images benefits many precise Earth observation tasks. Recently, visual foundation models have promoted the remote sensing interpretation but concentrating on RGB and multispectral images. Due to the varied hyperspectral channels,existing foundation models would face image-by-image tuning situation, imposing great pressure on hardware and time resources. In this paper, we propose a tuning-free hyperspectral foundation model called HyperFree, by adapting the existing visual prompt engineering. To process varied channel numbers, we design a learned weight dictionary covering full-spectrum from $0.4 \sim 2.5 \, μ\text{m}$, supporting to build the embedding layer dynamically. To make the prompt design more tractable, HyperFree can generate multiple semantic-aware masks for one prompt by treating feature distance as semantic-similarity. After pre-training HyperFree on constructed large-scale high-resolution hyperspectral images, HyperFree (1 prompt) has shown comparable results with specialized models (5 shots) on 5 tasks and 11 datasets.Code and dataset are accessible at https://rsidea.whu.edu.cn/hyperfree.htm.

高光谱基础模型无微调遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。