用极少人工标注实现卫星图中的学校精准识别
Label-Efficient School Detection from Aerial Imagery via Weakly Supervised Pretraining and Fine-Tuning

- 通过稀疏点位与语义分割自动生成标签,减少人工标注依赖
- 仅需50张人工标注图即可在低数据场景下实现强检测性能
- 适合教育规划、数字基建等需要大规模地图绘制的场景
精准的学校检测对教育支持举措(如基础设施规划和偏远地区网络覆盖扩展)至关重要。然而,全球许多地区面临官方记录陈旧、不完整或缺失的问题。人工制图虽有价值,但耗时且难以大规模推广。为此,我们提出一种弱监督框架,通过航空影像实现学校检测,极大降低对人工标注的需求,支持全球地图绘制。方法针对低数据场景设计:利用稀疏位置点和语义分割自动生成基础设施掩码,并从中提取边界框;基于这些自动标注图像进行第一阶段训练,学习学校外观表征;随后使用少量人工标注图像对模型进行微调。该两阶段训练流程在极低标注量下实现高效强检测。实验表明,在仅50张人工标注图像的条件下,模型仍表现出色,显著降低标注成本。本框架为全球教育与连通性倡议提供高效可扩展的空基学校测绘方案。所有模型、训练代码及自动生成数据将公开发布,以促进后续研究与实际应用。
原文摘要 · Abstract (English)
Accurate school detection is essential for supporting education initiatives, including infrastructure planning and expanding internet connectivity to underserved areas. However, many regions around the world face challenges due to outdated, incomplete, or unavailable official records. Manual mapping efforts, while valuable, are labor-intensive and lack scalability across large geographic areas. To address this, we propose a weakly supervised framework for school detection from aerial imagery that minimizes the need for human annotations while supporting global mapping efforts. Our method is specifically designed for low-data regimes, where manual annotations are extremely scarce. We introduce an automatic labeling pipeline that leverages sparse location points and semantic segmentation to generate infrastructure masks from which we generate bounding boxes. Using these automatically labeled images, we train our detectors on a first training stage to learn a representation of what schools look like, then using a small set of manually labeled images, we fine-tune the previously trained models on this clean dataset. This two stage training pipeline enables large-scale and strong detection in low-data setting of school infrastructure with minimal supervision. Our results demonstrate strong object detection performance, particularly in the low-data regime, where the models achieve promising results using only 50 manually labeled images, significantly reducing the need for costly annotations. This framework supports education and connectivity initiatives worldwide by providing an efficient and extensible approach to mapping schools from space. All models, training code and auto-labeled data will be publicly released to foster future research and real-world impact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。