用多条件控制生成高光谱图像,提升真实感与多样性。
HSIGene: A Foundation Model For Hyperspectral Image Generation
- 基于潜在扩散模型,支持多条件控制生成。
- 通过超分辨率数据增强,扩充训练样本并保持光谱保真度。
- 适合需要大量真实高光谱图像的农业与环境监测研究者。
高光谱图像(HSI)在农业和环境监测等领域至关重要,但因采集成本高,数据量有限,影响下游任务性能。现有扩散模型合成的HSI仍受限于数据稀缺,生成结果可靠性与多样性不足。部分研究引入多模态数据提升空间多样性,但难以保证光谱准确性。此外,现有模型通常不可控或仅支持单条件控制,限制生成精度。为此,我们提出HSIGene,一种基于潜在扩散的高光谱图像生成基础模型,支持多条件控制,实现更精准可靠的生成。为增强训练数据的空间多样性并保留光谱保真度,我们提出一种基于空间超分辨率的数据增强方法:先对HSI进行上采样,再通过裁剪获取丰富训练块。为进一步提升增强数据的感知质量,提出两阶段超分辨率框架,先处理RGB波段,再利用提出的矩形引导注意力网络(RGAN)进行引导式HSI超分辨率。实验表明,该模型可生成大量适用于去噪、超分辨率等下游任务的真实高光谱图像。代码与模型已开源。
原文摘要 · Abstract (English)
Hyperspectral image (HSI) plays a vital role in various fields such as agriculture and environmental monitoring. However, due to the expensive acquisition cost, the number of hyperspectral images is limited, degenerating the performance of downstream tasks. Although some recent studies have attempted to employ diffusion models to synthesize HSIs, they still struggle with the scarcity of HSIs, affecting the reliability and diversity of the generated images. Some studies propose to incorporate multi-modal data to enhance spatial diversity, but the spectral fidelity cannot be ensured. In addition, existing HSI synthesis models are typically uncontrollable or only support single-condition control, limiting their ability to generate accurate and reliable HSIs. To alleviate these issues, we propose HSIGene, a novel HSI generation foundation model which is based on latent diffusion and supports multi-condition control, allowing for more precise and reliable HSI generation. To enhance the spatial diversity of the training data while preserving spectral fidelity, we propose a new data augmentation method based on spatial super-resolution, in which HSIs are upscaled first, and thus abundant training patches could be obtained by cropping the high-resolution HSIs. In addition, to improve the perceptual quality of the augmented data, we introduce a novel two-stage HSI super-resolution framework, which first applies RGB bands super-resolution and then utilizes our proposed Rectangular Guided Attention Network (RGAN) for guided HSI super-resolution. Experiments demonstrate that the proposed model is capable of generating a vast quantity of realistic HSIs for downstream tasks such as denoising and super-resolution. The code and models are available at https://github.com/LiPang/HSIGene.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。