用多尺度影像和特征增强提升建筑分割精度
Feature-Augmented Deep Networks for Multiscale Building Segmentation in High-Resolution UAV and Satellite Imagery
- 从RGB图中提取PCA、VDVI等特征辅助深度网络学习
- 在WorldView-3上达到96.5%准确率,F1-score 0.86,IoU 0.80
- 适合遥感图像建筑识别任务,尤其对复杂阴影场景有效
高分辨率RGB影像中的建筑分割因与非建筑物光谱相似、阴影干扰及建筑形态不规则而困难。本研究提出一种综合深度学习框架,用于处理空间分辨率0.4m至2.7m的航拍与卫星影像。构建了多传感器多样化数据集,通过主成分分析(PCA)、可见差异植被指数(VDVI)、形态学建筑指数(MBI)及Sobel边缘滤波器从RGB通道生成辅助特征,增强Res-U-Net对复杂空间模式的学习能力。同时引入层冻结、周期性学习率和SuperConvergence训练策略,降低训练时间与资源消耗。在保留的WorldView-3影像上评估,模型整体准确率达96.5%,F1-score为0.86,交并比(IoU)达0.80,优于现有基于RGB的基准方法。结果表明,多分辨率影像、特征增强与优化训练策略结合,可实现遥感应用中鲁棒的建筑分割。
原文摘要 · Abstract (English)
Accurate building segmentation from high-resolution RGB imagery remains challenging due to spectral similarity with non-building features, shadows, and irregular building geometries. In this study, we present a comprehensive deep learning framework for multiscale building segmentation using RGB aerial and satellite imagery with spatial resolutions ranging from 0.4m to 2.7m. We curate a diverse, multi-sensor dataset and introduce feature-augmented inputs by deriving secondary representations including Principal Component Analysis (PCA), Visible Difference Vegetation Index (VDVI), Morphological Building Index (MBI), and Sobel edge filters from RGB channels. These features guide a Res-U-Net architecture in learning complex spatial patterns more effectively. We also propose training policies incorporating layer freezing, cyclical learning rates, and SuperConvergence to reduce training time and resource usage. Evaluated on a held-out WorldView-3 image, our model achieves an overall accuracy of 96.5%, an F1-score of 0.86, and an Intersection over Union (IoU) of 0.80, outperforming existing RGB-based benchmarks. This study demonstrates the effectiveness of combining multi-resolution imagery, feature augmentation, and optimized training strategies for robust building segmentation in remote sensing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。