arXiv:2502.19700cs.CV2025-02被引 4

用文本控制生成高光谱图像,解决小样本不均衡问题

Language-Informed Hyperspectral Image Synthesis for Imbalanced-Small Sample Classification via Semi-Supervised Conditional Diffusion Model

  • 用变分自编码器降维,再通过条件扩散模型生成带文本提示的图像
  • 半监督设计利用未标注数据,随机多边形裁剪和不确定性估计提升多样性
  • 生成图像在分布和视觉上均逼真,适合遥感小样本分类任务

数据增强有效缓解高光谱图像分类中的小样本不均衡问题。现有方法多在潜在空间扩展特征,较少利用文本驱动生成真实且多样的样本。近年来,文本引导的扩散模型因能基于文本提示生成高质量、多样图像而受到关注。本文提出 Txt2HSI-LDM(VAE),一种新型语言引导高光谱图像合成方法,用于解决高光谱图像分类中的小样本不均衡问题。该方法采用去噪扩散模型,通过迭代去除高斯噪声生成条件于文本描述的高光谱样本。首先,为应对高光谱数据的高维度,设计通用变分自编码器(VAE),将数据映射至低维潜在空间,提供稳定特征并降低扩散模型推理复杂度。其次,设计半监督扩散模型,充分利用未标注数据;通过随机多边形空间裁剪(RPSC)和潜在特征不确定性估计(LF-UE)模拟不同混合程度。第三,以语言条件为输入,由VAE解码器从扩散模型生成的潜在空间中重建高光谱图像。实验中,我们从统计特性与二维主成分分析(2D-PCA)空间的数据分布评估合成样本的有效性。此外,可视化像素级视觉-语言交叉注意力,证明模型能捕捉生成数据的空间布局与几何结构。实验表明,所提方法性能优于经典基线模型、前沿卷积神经网络及半监督方法。

原文摘要 · Abstract (English)

Data augmentation effectively addresses the imbalanced-small sample data (ISSD) problem in hyperspectral image classification (HSIC). While most methodologies extend features in the latent space, few leverage text-driven generation to create realistic and diverse samples. Recently, text-guided diffusion models have gained significant attention due to their ability to generate highly diverse and high-quality images based on text prompts in natural image synthesis. Motivated by this, this paper proposes Txt2HSI-LDM(VAE), a novel language-informed hyperspectral image synthesis method to address the ISSD in HSIC. The proposed approach uses a denoising diffusion model, which iteratively removes Gaussian noise to generate hyperspectral samples conditioned on textual descriptions. First, to address the high-dimensionality of hyperspectral data, a universal variational autoencoder (VAE) is designed to map the data into a low-dimensional latent space, which provides stable features and reduces the inference complexity of diffusion model. Second, a semi-supervised diffusion model is designed to fully take advantage of unlabeled data. Random polygon spatial clipping (RPSC) and uncertainty estimation of latent feature (LF-UE) are used to simulate the varying degrees of mixing. Third, the VAE decodes HSI from latent space generated by the diffusion model with the language conditions as input. In our experiments, we fully evaluate synthetic samples' effectiveness from statistical characteristics and data distribution in 2D-PCA space. Additionally, visual-linguistic cross-attention is visualized on the pixel level to prove that our proposed model can capture the spatial layout and geometry of the generated data. Experiments demonstrate that the performance of the proposed Txt2HSI-LDM(VAE) surpasses the classical backbone models, state-of-the-art CNNs, and semi-supervised methods.

高光谱图像扩散模型文本生成小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。