用视觉+周边功能信息,提升老旧小区房屋健康评估精度。
How Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing Inspection

- 融合多视角图像与周边兴趣点数据,构建建筑级健康评估框架。
- 引入周边500~1500米功能特征,显著提升评估准确率至76.79%。
- 适合城市更新、智慧城管等需要精细化评估的场景使用。
住宅级城市体检对识别住房问题和推动精准城市更新至关重要。现有自动化检测研究多依赖单张图像,较少探讨周边城市功能环境是否可为建筑级评估提供补充信息。本研究提出一种视觉-兴趣点(POI)融合框架,结合多视角视觉检测与POI提取的邻域环境特征,实现住宅建筑健康评估。实证数据涵盖中国青岛92个老社区、3,237栋住宅建筑及25,608张实地采集的检查图像,涉及七类住房问题。首先,评估多个目标检测模型以提取图像中问题的位置、类别与置信度,再通过多视角聚合生成可解释的建筑级表征。其次,在500米、1,000米、1,500米缓冲区提取POI特征,结合皮尔逊与斯皮尔曼相关性分析及错误发现率校正,筛选候选上下文特征。最后,在社区隔离的空间交叉验证下,采用代价敏感的随机森林分类器融合视觉与POI特征。结果表明,多视角聚合带来主要性能提升,建筑级宏平均F1从直接检测的60.84%提升至74.95%;引入POI上下文后进一步提升至76.79%,但增益有限且依赖具体问题类别。因此,POI信息作为辅助上下文先验,而非视觉证据的替代或建筑状况的因果决定因素。
原文摘要 · Abstract (English)
Housing-level urban physical examination is essential for identifying residential building problems and supporting targeted urban renewal. Existing automated inspection studies primarily rely on individual images and rarely examine whether surrounding urban functional context can provide supplementary information for building-level assessment. This study proposes a vision-POI fusion framework that combines multi-view visual inspection with POI-derived neighborhood context for residential building health assessment. The empirical dataset covers 92 old residential communities, 3,237 residential buildings, and 25,608 field-acquired inspection images in Qingdao, China, encompassing seven categories of housing-related issues. First, multiple object detection models are evaluated to extract issue locations, categories, and confidence scores from individual images. The image-level outputs are subsequently aggregated across multiple views to construct interpretable building-level representations. Second, POI features are extracted within 500m, 1,000m, and 1,500m neighborhood buffers to characterize surrounding functional environments. Pearson and Spearman correlation analyses, combined with false discovery rate correction, are used to identify candidate contextual features. Finally, visual and POI features are integrated using a cost-sensitive Random Forest classifier under community-isolated spatial cross-validation. The results show that multi-view aggregation provides the main performance improvement, increasing the building-level Macro-F1 from 60.84% under Direct Detection to 74.95%. Incorporating POI context further increases Macro-F1 to 76.79%, although the additional gain is modest and category-dependent. POI information therefore functions as a supplementary contextual prior rather than a substitute for direct visual evidence or a causal determinant of building condition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。