arXiv:2502.17056cs.CV2025-02被引 1

用扩散模型生成带像素级标注的高光谱图像,解决数据稀缺问题。

SpecDM: Hyperspectral Dataset Synthesis with Pixel-level Semantic Annotations

  • 双流VAE分别学习图像与掩码的潜在表示,联合建模扩散过程
  • 首次实现高维高光谱图像与像素级标注的同步生成
  • 适用于语义分割和变化检测任务,提升下游模型性能

在高光谱遥感领域,语义分割(SS)和变化检测(CD)等密集预测任务依赖监督学习,需大量人工标注数据。但高光谱图像(HSIs)的采集与标注因设备特殊、场景受限,成本高且耗时。为此,本文探索生成式扩散模型在合成带像素级标注的高光谱图像中的潜力。核心思路是使用双流变分自编码器(VAE)分别学习图像与对应掩码的潜在表示,训练中建模其联合分布,最终通过各自解码器生成图像与掩码。据我们所知,这是首个实现高维高光谱图像与标注同步生成的工作。该方法可应用于多种数据集生成任务,本文选取语义分割与变化检测两类典型任务,生成适配数据集。实验表明,合成数据集对下游任务有正向促进作用。

原文摘要 · Abstract (English)

In hyperspectral remote sensing field, some downstream dense prediction tasks, such as semantic segmentation (SS) and change detection (CD), rely on supervised learning to improve model performance and require a large amount of manually annotated data for training. However, due to the needs of specific equipment and special application scenarios, the acquisition and annotation of hyperspectral images (HSIs) are often costly and time-consuming. To this end, our work explores the potential of generative diffusion model in synthesizing HSIs with pixel-level annotations. The main idea is to utilize a two-stream VAE to learn the latent representations of images and corresponding masks respectively, learn their joint distribution during the diffusion model training, and finally obtain the image and mask through their respective decoders. To the best of our knowledge, it is the first work to generate high-dimensional HSIs with annotations. Our proposed approach can be applied in various kinds of dataset generation. We select two of the most widely used dense prediction tasks: semantic segmentation and change detection, and generate datasets suitable for these tasks. Experiments demonstrate that our synthetic datasets have a positive impact on the improvement of these downstream tasks.

高光谱数据生成扩散模型语义分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。