用街景图自动补全建筑属性,支持城市规划与智能分析
OpenFACADES: An Open Framework for Architectural Caption and Attribute Data Enrichment via Street View Imagery
- 融合街景与地图数据,定位最佳观测视角
- 将全景图像转为真实视角,提升建筑外观识别精度
- 利用开源多模态大模型实现属性预测与开放词汇描述
建筑属性如高度、用途和材质在空间数据基础设施中至关重要,支撑各类城市应用。尽管重要,许多城区仍缺乏全面的建筑属性数据。近期进展使通过遥感与街景影像提取客观属性成为可能,但整合多元开放数据、获取完整建筑影像并规模化推断综合属性仍具挑战。本文首次提出 OpenFACADES 框架,利用多模态众包数据,通过多模态大语言模型丰富建筑档案,涵盖客观属性与语义描述。首先,通过等视域分析将 Mapillary 街景元数据与 OpenStreetMap 几何信息对齐,识别适合观测目标建筑的图像;其次,自动化检测全景图像中的建筑立面,并设计重投影方法将其转换为近似真实观察的全景视图;第三,引入创新方法,利用开源视觉-语言模型(VLMs)进行多属性预测与开放词汇字幕生成,基于来自七座城市的 31,180 张标注图像数据集。评估显示,微调后的 VLM 在多属性推理上优于单属性计算机视觉模型和零样本 ChatGPT-4o。进一步实验验证其在文化差异显著区域及不同图像条件下的优越泛化性与鲁棒性。
原文摘要 · Abstract (English)
Building properties, such as height, usage, and material, play a crucial role in spatial data infrastructures, supporting various urban applications. Despite their importance, comprehensive building attribute data remain scarce in many urban areas. Recent advances have enabled the extraction of objective building attributes using remote sensing and street-level imagery. However, establishing a pipeline that integrates diverse open datasets, acquires holistic building imagery, and infers comprehensive building attributes at scale remains a significant challenge. Among the first, this study bridges the gaps by introducing OpenFACADES, an open framework that leverages multimodal crowdsourced data to enrich building profiles with both objective attributes and semantic descriptors through multimodal large language models. First, we integrate street-level image metadata from Mapillary with OpenStreetMap geometries via isovist analysis, identifying images that provide suitable vantage points for observing target buildings. Second, we automate the detection of building facades in panoramic imagery and tailor a reprojection approach to convert objects into holistic perspective views that approximate real-world observation. Third, we introduce an innovative approach that harnesses and investigates the capabilities of open-source large vision-language models (VLMs) for multi-attribute prediction and open-vocabulary captioning in building-level analytics, leveraging a globally sourced dataset of 31,180 labeled images from seven cities. Evaluation shows that fine-tuned VLM excel in multi-attribute inference, outperforming single-attribute computer vision models and zero-shot ChatGPT-4o. Further experiments confirm its superior generalization and robustness across culturally distinct region and varying image conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。