arXiv:2604.18201cs.CVcs.LG2026-04中稿 · ICLR

用扩散模型提升遥感图像零样本目标定位精度

DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery

论文配图:DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery
图 1 · 摘自论文原文
  • 融合扩散模型与分割模型生成定位线索
  • 在复杂场景下定位准确率提升超14%
  • 适合遥感图像中无标注目标的快速定位

扩散模型在视觉任务中表现出强大能力,包括文本引导的图像生成与编辑。本文探索其在遥感图像目标定位中的潜力,提出一种混合流程:将基于扩散的定位提示与RemoteSAM、SAM3等先进分割模型结合,以获得更精确的边界框。通过利用生成式扩散模型与基础分割模型的互补优势,该方法实现了对复杂场景中目标的鲁棒且自适应定位。实验表明,该流程显著提升了定位性能,在[email protected]指标上较现有最先进方法提升超过14%。

原文摘要 · Abstract (English)

Diffusion models have emerged as powerful tools for a wide range of vision tasks, including text-guided image generation and editing. In this work, we explore their potential for object grounding in remote sensing imagery. We propose a hybrid pipeline that integrates diffusion-based localization cues with state-of-the-art segmentation models such as RemoteSAM and SAM3 to obtain more accurate bounding boxes. By leveraging the complementary strengths of generative diffusion models and foundational segmentation models, our approach enables robust and adaptive object localization across complex scenes. Experiments demonstrate that our pipeline significantly improves localization performance, achieving over a 14% increase in [email protected] compared to existing state-of-the-art methods.

遥感图像扩散模型目标定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。