arXiv:2604.03505cs.CV2026-04

融合卫星与街景图,用少标注实现高效城市树木检测

Multimodal Urban Tree Detection from Satellite and Street-Level Imagery via Annotation-Efficient Deep Learning Strategies

论文配图:Multimodal Urban Tree Detection from Satellite and Street-Level Imagery via Annotation-Efficient Deep Learning Strategies
图 1 · 摘自论文原文
  • 用卫星图定位树候选区,再调取对应街景图精检,减少无效采样
  • 混合学习策略达F1 0.90,比基线提升12%,显著降低误检漏检
  • 适合需要低成本、高精度城市绿化监测的规划与环保机构

除了直接的生态效益,城市树木在环境可持续性和灾害缓解中发挥基础作用。精确绘制城市树木分布对环境监测、灾后评估和政策制定至关重要。然而,从传统人工调查转向可扩展的自动化系统仍受限于高昂的标注成本及跨城市场景泛化能力差。本研究提出一种多模态框架,结合高分辨率卫星影像与地面级Google Street View,实现有限标注条件下的可扩展、精细化城市树木检测。该框架首先利用卫星影像定位树体候选区域,再针对性获取对应的街景图像进行细节检测,大幅减少无效街景采样。为突破标注瓶颈,采用领域自适应将已有标注数据的知识迁移至新区域。进一步通过三种学习策略评估:半监督学习、主动学习及两者结合的混合策略,均基于基于Transformer的检测模型。结果显示,混合策略表现最佳,F1-score达0.90,较基线提升12%;半监督学习因伪标签中的确认偏差导致性能逐步下降;而主动学习通过有选择地标注不确定或错误预测,持续提升效果。误差分析表明,主动学习与混合策略均有效降低假阳性与假阴性。研究强调多模态融合与引导式标注在构建可扩展、低标注依赖的城市树木制图系统中的关键作用,有助于推动可持续城市规划。

原文摘要 · Abstract (English)

Beyond the immediate biophysical benefits, urban trees play a foundational role in environmental sustainability and disaster mitigation. Precise mapping of urban trees is essential for environmental monitoring, post-disaster assessment, and strengthening policy. However, the transition from traditional, labor-intensive field surveys to scalable automated systems remains limited by high annotation costs and poor generalization across diverse urban scenarios. This study introduces a multimodal framework that integrates high-resolution satellite imagery with ground-level Google Street View to enable scalable and detailed urban tree detection under limited-annotation conditions. The framework first leverages satellite imagery to localize tree candidates and then retrieves targeted ground-level views for detailed detection, significantly reducing inefficient street-level sampling. To address the annotation bottleneck, domain adaptation is used to transfer knowledge from an existing annotated dataset to a new region of interest. To further minimize human effort, we evaluated three learning strategies: semi-supervised learning, active learning, and a hybrid approach combining both, using a transformer-based detection model. The hybrid strategy achieved the best performance with an F1-score of 0.90, representing a 12% improvement over the baseline model. In contrast, semi-supervised learning exhibited progressive performance degradation due to confirmation bias in pseudo-labeling, while active learning steadily improved results through targeted human intervention to label uncertain or incorrect predictions. Error analysis further showed that active and hybrid strategies reduced both false positives and false negatives. Our findings highlight the importance of a multimodal approach and guided annotation for scalable, annotation-efficient urban tree mapping to strengthen sustainable city planning.

城市绿化多模态少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。