首个兼顾近距与远距光谱成像的通用光谱基础模型
A General Purpose Spectral Foundational Model for Both Proximal and Remote Sensing Spectral Imaging
- 基于掩码自编码器架构,融合光谱通道编码与时空掩码机制
- 在百通道高光谱数据上训练,支持远距与近距光谱任务
- 适配遥感与工业检测场景,解决小样本光谱建模难题
多光谱和高光谱相机可获取数百个波段的光谱成像数据,每个波段记录特定波长与带宽下的反射率。受限于时间和资源,大规模光谱数据集难以构建,导致从零开始训练预测模型困难。在RGB领域,预训练基础模型可缓解小样本问题,但现有模型大多基于3通道RGB图像预训练,难以直接用于光谱数据。目前少数光谱基础模型存在两大局限:(1) 仅在遥感数据上训练,不适用于近距光谱成像;(2) 使用少于15通道的多光谱数据集,无法处理百通道高光谱图像。为解决上述问题,我们提出一个大规模基础模型及数据集,采用掩码自编码器架构,结合光谱通道编码、空间-光谱掩码和ImageNet预训练,实现对下游光谱成像任务的高效适配与鲁棒性表现。
原文摘要 · Abstract (English)
Spectral imaging data acquired via multispectral and hyperspectral cameras can have hundreds of channels, where each channel records the reflectance at a specific wavelength and bandwidth. Time and resource constraints limit our ability to collect large spectral datasets, making it difficult to build and train predictive models from scratch. In the RGB domain, we can often alleviate some of the limitations of smaller datasets by using pretrained foundational models as a starting point. However, most existing foundation models are pretrained on large datasets of 3-channel RGB images, severely limiting their effectiveness when used with spectral imaging data. The few spectral foundation models that do exist usually have one of two limitations: (1) they are built and trained only on remote sensing data limiting their application in proximal spectral imaging, (2) they utilize the more widely available multispectral imaging datasets with less than 15 channels restricting their use with hundred-channel hyperspectral images. To alleviate these issues, we propose a large-scale foundational model and dataset built upon the masked autoencoder architecture that takes advantage of spectral channel encoding, spatial-spectral masking and ImageNet pretraining for an adaptable and robust model for downstream spectral imaging tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。