AI看街景判建筑类型,准确率约70%,能帮城市分析提速。
AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery
- 用思维链提示让AI更稳定地判断建筑类型
- AI准确率达70%,接近人类专家水平
- 适合想用AI做城市视觉分析的研究者
本研究探究视觉语言模型(VLMs)从谷歌街景(GSV)图像中推断建筑类型(建造方式、当前用途、层数)的潜力。将VLM的预测结果与人类专家(土木工程师和建筑师)的判断进行对比,作为人工标注的基准数据。评估了GPT-4o、Claude 3.5 Sonnet和Gemini 2.0 Flash等先进VLMs,通过不同缩放策略和提示技术发现,思维链(Chain-of-Thought)提示可带来更稳定的模型表现。通过分析AI解释中关键词的概率,揭示其推理模式,识别出驱动AI与专家一致或分歧的关键主题。结果显示,AI主要依赖视觉线索,而人类专家则更重视上下文信息和领域知识。总体上,VLM可在大规模场景下近似专家能力,平均准确率约为70%。研究展示了VLM在需模式识别与物体识别的城市任务中的自动化潜力,可作为城市分析的互补工具,利用其视觉模式理解优势。本研究推动了对AI视觉预测效率与可扩展性的探索,并为城市分析自动化提供推理机制洞察。
原文摘要 · Abstract (English)
This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Storeys from Google Street View (GSV) images. Predictions generated by VLMs are compared with inference by human experts (civil engineers and architects) as a source of manually labelled ground-truth data. We evaluate several state-of-the-art VLMs, including GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash. By applying different scaling strategies and prompting techniques, we found that Chain-of-Thought prompts provide an overall more stable model performance. We also investigate the reasoning behind VLMs' building-typology predictions by examining the probabilities of keywords appearing in AI explanations. This enabled us to analyse patterns in these reasonings and identify key themes driving both agreements and disagreements between VLM and expert labels. We find that AI tends to focus on visual indicators, whereas human experts place greater emphasis on broader contextual cues and domain knowledge, in addition to visual cues. Overall, VLM can approximate experts' capability in building-typology classification at scale, with an average accuracy of approximately 70%. The study demonstrates the VLM's potential for AI automation in tasks that require pattern recognition and object identification in an urban context. AI have the potential to serve as complementary and collaborative tools for urban analysis, leveraging their strengths in understanding visual patterns. This study contributes to the exploration of the efficiency and scalability of AI visual prediction and provides insights into the reasoning processes that could support automation processes in urban analysis and prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。