用地理人工智能自动检测建筑轮廓错误,提升地图数据质量。
GeoAI-based post-segmentation quality validation of building footprints via spatial feature engineering

- 基于空间特征工程构建多维度预测器,识别建筑轮廓缺陷。
- 在未见测试集上准确率达95.31%,误标率降低83.09%。
- 适合需要高精度地图数据的智慧城市与遥感分析场景。
基于深度学习的高分辨率影像建筑轮廓提取常生成拓扑不一致的矢量数据,无法直接用于地理信息系统(GIS)数据库。为此,本文提出一种多领域地理人工智能质量控制框架,实现误差自动检测与数据库系统性净化。在孟加拉国五个无人机测绘站点,采用U-Net(ResNet-34)和SAM-LoRA(ViT-B)生成候选轮廓,经栅格掩码向量化、几何规整及空间独占约束合并,消除重复表示。引入24个包含几何、空间上下文及栅格光谱纹理特性的预测因子。机器学习分类器在开发集(站点B-D)训练,并在空间独立测试集(站点E)严格验证,未参与超参数调优与类别平衡。实验表明,几何与空间上下文预测因子结合决策树(DT)在识别对象级边界畸变方面表现最优:在未见测试集上准确率达95.31%,F1分数为91.06%,马修斯相关系数(MCC)达0.880。在数据库层面,该框架成功识别87.34%的错误轮廓,同时保留98.31%的有效结构,残余错误比例从27.32%降至4.62%,最终数据库纯度提升至95.38%,相对错误减少83.09%。结果表明,后分割阶段的物体级机器学习方法可为生产级GIS工作流提供高度可迁移、鲁棒的自动化质量保障机制。
原文摘要 · Abstract (English)
Deep learning-based building footprint extraction from high-resolution imagery often produces topologically inconsistent vectors unfit for direct GIS database ingestion. To address this, we present a multidomain GeoAI quality control framework that automates error detection to systematically purify vector footprint databases. Candidate footprints were generated across five UAV survey sites in Bangladesh using U-Net (ResNet-34) and SAM-LoRA (ViT-B). The extracted raster masks were vectorized, geometrically regularized, and consolidated under a spatial-exclusivity constraint to eliminate duplicate representations. We used twenty-four predictors capturing geometric, spatial-contextual, and raster-derived spectral and texture properties. Machine Learning (ML) classifiers were trained on a development partition (Sites B-D) and rigorously validated on a spatially independent test set (Site E) excluded from hyperparameter tuning and class balancing. The experimental results demonstrate that geometric and spatial-contextual predictors using Decision Tree (DT) provide the most effective discriminatory evidence for identifying object-level boundary deformations. DT achieved an accuracy of 95.31%, an F1-score of 91.06%, and a Matthews correlation coefficient (MCC) of 0.880 on the unseen testing site. At the database level, this framework successfully identified 87.34% of erroneous footprints while maintaining 98.31% of acceptable structures, reducing the residual error proportion from 27.32% to 4.62% and improving final database purity to 95.38%. This translates into a relative error reduction of 83.09%. The findings indicate that post-segmentation object-level ML provides a highly transferable, robust mechanism for automated quality assurance in production-ready geographic information system (GIS) workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。