arXiv:2512.19990cs.CV2025-12

用双分支框架解决低分辨率标签生成高分辨率地表分类问题

A Dual-Branch Local-Global Framework for Cross-Resolution Land Cover Mapping

  • 分两路处理:扩散模型优化局部细节,Transformer保持全局一致
  • 在切萨皮克湾数据集上达到66.52% mIoU,超越现有弱监督方法
  • 适合遥感图像分析、地理信息建模等需要跨分辨率学习的研究者

跨分辨率地表覆盖制图旨在从粗粒度或低分辨率标签中生成高分辨率语义预测,但严重的分辨率差异使有效学习极具挑战。现有弱监督方法常难以将细粒度空间结构与粗标签对齐,导致噪声监督和精度下降。为此,我们提出DDTM,一种双分支弱监督框架,显式分离局部语义精化与全局上下文推理。具体而言,DDTM引入基于扩散的分支,在粗标签监督下逐步优化细尺度局部语义;同时采用基于Transformer的分支,强化大范围空间内的长程上下文一致性。此外,设计伪标签置信度评估模块,缓解跨分辨率不一致带来的噪声,有选择地利用可靠监督信号。大量实验表明,DDTM在切萨皮克湾基准上达到新最佳性能,实现66.52% mIoU,显著优于先前弱监督方法。代码已开源。

原文摘要 · Abstract (English)

Cross-resolution land cover mapping aims to produce high-resolution semantic predictions from coarse or low-resolution supervision, yet the severe resolution mismatch makes effective learning highly challenging. Existing weakly supervised approaches often struggle to align fine-grained spatial structures with coarse labels, leading to noisy supervision and degraded mapping accuracy. To tackle this problem, we propose DDTM, a dual-branch weakly supervised framework that explicitly decouples local semantic refinement from global contextual reasoning. Specifically, DDTM introduces a diffusion-based branch to progressively refine fine-scale local semantics under coarse supervision, while a transformer-based branch enforces long-range contextual consistency across large spatial extents. In addition, we design a pseudo-label confidence evaluation module to mitigate noise induced by cross-resolution inconsistencies and to selectively exploit reliable supervisory signals. Extensive experiments demonstrate that DDTM establishes a new state-of-the-art on the Chesapeake Bay benchmark, achieving 66.52\% mIoU and substantially outperforming prior weakly supervised methods. The code is available at https://github.com/gpgpgp123/DDTM.

地表覆盖弱监督扩散模型遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。