arXiv:2506.18946cs.CVcs.AI2025-06被引 9

用扩散模型提升遥感图像语义分割的精准度

DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models

  • 引入上下文感知适配器与渐进式跨模态解码器,增强语言与图像对齐
  • 在三个基准数据集上均达到新最优性能,最高提升达8.2个点
  • 适合遥感、灾害响应等需要精准语义理解的领域应用

referring remote sensing image segmentation (RRSIS) 能通过自然语言描述精确划分遥感图像中的区域,广泛应用于灾害响应、城市规划和环境监测。尽管近期取得进展,现有方法在处理航空影像时仍面临尺度变化、方向多样和语义模糊等挑战。为此,我们提出 DiffRIS,利用预训练文本到图像扩散模型的语义理解能力,提升 RRSIS 中的跨模态对齐。该框架包含两项创新:上下文感知适配器(CP-adapter)通过全局上下文建模和对象感知推理动态优化语言特征;渐进式跨模态推理解码器(PCMRD)通过多尺度特征交互,迭代对齐文本描述与视觉区域。CP-adapter 缓解了通用视觉-语言理解与遥感任务间的域差距,而 PCMRD 实现细粒度语义对齐。在三个基准数据集——RRSIS-D、RefSegRS 与 RISBench——上的全面实验表明,DiffRIS 在所有标准指标上持续优于现有方法,确立了新的最先进水平。显著的性能提升验证了通过自适应框架利用预训练扩散模型在遥感任务中的有效性。

原文摘要 · Abstract (English)

Referring remote sensing image segmentation (RRSIS) enables the precise delineation of regions within remote sensing imagery through natural language descriptions, serving critical applications in disaster response, urban development, and environmental monitoring. Despite recent advances, current approaches face significant challenges in processing aerial imagery due to complex object characteristics including scale variations, diverse orientations, and semantic ambiguities inherent to the overhead perspective. To address these limitations, we propose DiffRIS, a novel framework that harnesses the semantic understanding capabilities of pre-trained text-to-image diffusion models for enhanced cross-modal alignment in RRSIS tasks. Our framework introduces two key innovations: a context perception adapter (CP-adapter) that dynamically refines linguistic features through global context modeling and object-aware reasoning, and a progressive cross-modal reasoning decoder (PCMRD) that iteratively aligns textual descriptions with visual regions for precise segmentation. The CP-adapter bridges the domain gap between general vision-language understanding and remote sensing applications, while PCMRD enables fine-grained semantic alignment through multi-scale feature interaction. Comprehensive experiments on three benchmark datasets-RRSIS-D, RefSegRS, and RISBench-demonstrate that DiffRIS consistently outperforms existing methods across all standard metrics, establishing a new state-of-the-art for RRSIS tasks. The significant performance improvements validate the effectiveness of leveraging pre-trained diffusion models for remote sensing applications through our proposed adaptive framework.

遥感分割扩散模型跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。