用双模态超分辨率Transformer提升梯田矢量化精度,解决跨尺度数据融合难题。
ΩSFormer: Dual-Modal Ω-like Super-Resolution Transformer Network for Cross-scale and High-accuracy Terraced Field Vectorization Extraction
- 融合遥感与地形数据的双模态特征,通过Ω结构实现跨尺度超分辨率重建。
- 在9个中国研究区上实现22441平方公里覆盖,mIOU提升最高0.297。
- 适合需要高精度梯田矢量提取的土壤水保监测与地理信息系统应用。
梯田是重要的水土保持工程实践,从遥感影像中提取梯田是监测评估水土保持的基础。本研究首次提出一种新型双模态Ω类超分辨率Transformer网络(ΩSFormer),用于智能梯田矢量化提取(TFVE),具有以下优势:(1) 通过在编码器每一步融合原始高分辨率特征与下采样特征,并利用多头注意力机制,降低传统多尺度下采样编码器导致的边缘分割误差;(2) 提出Ω类网络结构,充分整合光谱与地形数据中的丰富高层特征,生成跨尺度超分辨率特征,显著提升提取精度;(3) 验证了跨模态、跨尺度(遥感影像与数字高程模型空间分辨率不一致)超分辨率特征提取的最优融合方案;(4) 采用粗到精及空间拓扑语义关系优化(STSRO)策略,缓解分割边缘像素的不确定性;(5) 利用等高线振荡神经网络持续优化参数,迭代生成梯田矢量结果。此外,首次构建了基于深度学习的梯田矢量化数据集DMRVD,涵盖中国四省九个研究区,总面积达22441平方公里。为评估ΩSFormer性能,对比经典与最先进网络,其mIOU分别较单模遥感、单模DEM及双模结果提升0.165、0.297和0.128。
原文摘要 · Abstract (English)
Terraced field is a significant engineering practice for soil and water conservation (SWC). Terraced field extraction from remotely sensed imagery is the foundation for monitoring and evaluating SWC. This study is the first to propose a novel dual-modal Ω-like super-resolution Transformer network for intelligent TFVE, offering the following advantages: (1) reducing edge segmentation error from conventional multi-scale downsampling encoder, through fusing original high-resolution features with downsampling features at each step of encoder and leveraging a multi-head attention mechanism; (2) improving the accuracy of TFVE by proposing a Ω-like network structure, which fully integrates rich high-level features from both spectral and terrain data to form cross-scale super-resolution features; (3) validating an optimal fusion scheme for cross-modal and cross-scale (i.e., inconsistent spatial resolution between remotely sensed imagery and DEM) super-resolution feature extraction; (4) mitigating uncertainty between segmentation edge pixels by a coarse-to-fine and spatial topological semantic relationship optimization (STSRO) segmentation strategy; (5) leveraging contour vibration neural network to continuously optimize parameters and iteratively vectorize terraced fields from semantic segmentation results. Moreover, a DMRVD for deep-learning-based TFVE was created for the first time, which covers nine study areas in four provinces of China, with a total coverage area of 22441 square kilometers. To assess the performance of ΩSFormer, classic and SOTA networks were compared. The mIOU of ΩSFormer has improved by 0.165, 0.297 and 0.128 respectively, when compared with best accuracy single-modal remotely sensed imagery, single-modal DEM and dual-modal result.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。