YOLOv11联合分割建筑并分类高度,提升城市建模精度与效率。
Mask-to-Height: A YOLOv11-Based Architecture for Joint Building Instance Segmentation and Height Classification from Satellite Imagery
- 基于YOLOv11设计新架构,融合多尺度特征提升定位精度。
- 在DFC2023数据集上达60.4% mAP@50,38.3% mAP@50-95,五类高度分类准确。
- 擅长处理遮挡、复杂形状和稀有高楼,适合实时大范围城市测绘。
精确的建筑实例分割与高度分类对城市规划、三维城市建模和基础设施监测至关重要。本文详细分析了最新YOLO系列模型YOLOv11在卫星影像中联合进行建筑提取与离散高度分类的应用。该模型通过更高效的架构,在不同尺度间更好地融合特征,提升目标定位精度,并增强在复杂城市场景中的表现。基于包含12个城市超过12.5万栋建筑标注的DFC2023 Track 2数据集,使用精确率、召回率、F1分数和平均精度(mAP)等指标评估性能。结果显示,YOLOv11在实例分割上达到60.4% mAP@50 与 38.3% mAP@50–95,同时在五个预定义高度层级上保持稳健分类准确率。模型在处理遮挡、复杂建筑形状及类别不平衡方面表现优异,尤其对罕见的高层建筑。对比分析表明,其在检测精度与推理速度上均优于早期多任务框架,适用于实时、大规模城市制图。本研究凸显了YOLOv11在简化高度分类建模方面的潜力,为遥感与地理空间智能的未来发展提供可行路径。
原文摘要 · Abstract (English)
Accurate building instance segmentation and height classification are critical for urban planning, 3D city modeling, and infrastructure monitoring. This paper presents a detailed analysis of YOLOv11, the recent advancement in the YOLO series of deep learning models, focusing on its application to joint building extraction and discrete height classification from satellite imagery. YOLOv11 builds on the strengths of earlier YOLO models by introducing a more efficient architecture that better combines features at different scales, improves object localization accuracy, and enhances performance in complex urban scenes. Using the DFC2023 Track 2 dataset -- which includes over 125,000 annotated buildings across 12 cities -- we evaluate YOLOv11's performance using metrics such as precision, recall, F1 score, and mean average precision (mAP). Our findings demonstrate that YOLOv11 achieves strong instance segmentation performance with 60.4\% mAP@50 and 38.3\% mAP@50--95 while maintaining robust classification accuracy across five predefined height tiers. The model excels in handling occlusions, complex building shapes, and class imbalance, particularly for rare high-rise structures. Comparative analysis confirms that YOLOv11 outperforms earlier multitask frameworks in both detection accuracy and inference speed, making it well-suited for real-time, large-scale urban mapping. This research highlights YOLOv11's potential to advance semantic urban reconstruction through streamlined categorical height modeling, offering actionable insights for future developments in remote sensing and geospatial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。