用扩散模型生成多样化遥感目标,提升小样本检测效果
Diverse Instance Generation via Diffusion Models for Enhanced Few-Shot Object Detection in Remote Sensing Images
- 用预训练扩散模型生成遥感目标片段,再嵌入真实图像增强数据
- 在多个数据集上平均提升4.4%检测性能,验证方法有效性
- 适合遥感小样本目标检测研究者,尤其关注数据增强的场景
小样本目标检测(FSOD)旨在仅用少量标注样本检测新类别目标,在濒危物种监测、灾情评估等遥感应用中尤为关键。现有遥感图像FSOD方法虽取得进展,但仍受限于实例多样性不足。为此,本文提出一种新框架:利用大规模自然图像预训练的扩散模型生成多样化的遥感实例,从而提升小样本检测器性能。不同于直接合成完整遥感图像,我们首先通过专用切片到切片模块生成实例级片段,再将其嵌入全尺度图像实现增强。为适配遥感场景,设计无类别图像反演模块,将遥感实例片段映射至语义空间;同时引入对比损失,使生成图像与对应类别在语义上对齐。实验表明,该方法在多个数据集和不同方法上平均提升4.4%。消融实验证明,精心设计的反演模块有效提升性能,语义对比损失进一步优化结果。
原文摘要 · Abstract (English)
Few-shot object detection (FSOD) aims to detect novel instances with only a limited number of labeled training samples, presenting a challenge that is particularly prominent in numerous remote sensing applications such as endangered species monitoring and disaster assessment. Existing FSOD methods for remote sensing images (RSIs) have achieved promising progress but remain constrained by the limited diversity of instances. To address this issue, we propose a novel framework that can leverage a diffusion model pretrained on large-scale natural images to synthesize diverse remote sensing instances, thereby improving the performance of few-shot object detectors. Instead of directly synthesizing complete remote sensing images, we first generate instance-level slices via a specialized slice-to-slice module, and then embed these slices into full-scale imagery for enhanced data augmentation. To further adapt diffusion models for remote sensing scenarios, we develop a class-agnostic image inversion module that can invert remote sensing instance slices into semantic space. Additionally, we introduce contrastive loss to semantically align the synthesized images with their corresponding classes. Experimental results show that our method hasachieved an average performance improvement of 4.4% across multiple datasets and various approaches. Ablation experiments indicate that the elaborately designed inversion module can effectively enhance the performance of FSOD methods, and the semantic contrastive loss can further boost the performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。