融合卫星与街景图像,实现多模态建筑检测与分类
Multi-Modal Building Inspection via Perceiver IO Fusion of Satellite and Street-Level Imagery

- 用Perceiver IO融合共享骨干的遥感与街景图像块
- 在10国3.2万栋建筑上实现多标签屋顶分类,街景任务提升11.3% AP
- 支持可变数量街景图输入,适合真实世界复杂场景
我们提出一种多模态分类框架,通过Perceiver IO架构融合卫星与街景图像,基于共享DINOv2主干网络生成的空间图像块。该设计无需填充或固定尺寸池化,可自然处理每栋建筑的可变数量街景视图,并联合预测多标签屋顶元素与材料类别。我们构建了包含32,135栋建筑(61,672个区域)的大规模数据集,覆盖十个国家,每段配对最多八张街景图,并评估了四种遮罩策略以定位目标建筑。提出一种RGB-M遮罩策略,将建筑轮廓掩码作为第四通道输入,提供软空间先验,在两种模态上均优于硬裁剪。Perceiver IO融合模型优于所有其他融合策略,对街景可见属性带来显著提升(如石板+11.3 AP,天窗+1.3 AP),但卫星单模基线在宏观平均mAP上仍略胜一筹,尤其适用于从上方可见的类别。结果验证了该架构在异构输入和多任务场景下的可扩展性与灵活性。
原文摘要 · Abstract (English)
We present a multi-modal classification framework that fuses satellite and street-level imagery through a Perceiver IO architecture operating on spatial patch tokens from a shared DINOv2 backbone. The design naturally handles a variable number of street-level views per building without padding or fixed-size pooling, and jointly predicts multi-label roof element and roof material classes. We construct a large-scale dataset of 32,135 buildings (61,672 segments) spanning ten countries, pairing satellite images with up to eight street-level views per segment and evaluating four masking strategies for isolating the target building. We propose an RGB-M masking strategy that appends the building footprint mask as a fourth input channel, providing a soft spatial prior that outperforms hard cropping across both modalities. The Perceiver IO fusion model improves over all other fusion strategies and yields substantial per-class gains for attributes visible from street level (e.g., +11.3 AP for slate, +1.3 AP for dormers), though the satellite-only baseline retains a slight advantage in macro-averaged mAP for classes that are predominantly visible from above. These results establish a scalable, flexible architecture for multi-modal building inspection that can accommodate heterogeneous inputs and multiple output tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。