arXiv:2504.21040cs.CV2025-04被引 7

用专家知识提升大模型对城市步行环境的评估能力

Can a Large Language Model Assess Urban Design Quality? Evaluating Walkability Metrics Across Expertise Levels

  • 将专家设计知识融入提示词,引导大模型更准确评估步行环境
  • 加入专家知识后,模型评分更集中且一致性更高
  • 适合城市规划、智能评估领域研究者参考

城市街道环境对公共空间中人类活动至关重要。随着街景图像(SVIs)与多模态大语言模型(MLLMs)的兴起,研究者和实践者正借助大数据探索、测量和评估城市环境的语义与视觉特征。由于使用MLLM构建自动化评估流程门槛低,亟需探究其潜在风险与机遇。本文首次探讨将正式化、结构化的专家城市设计知识嵌入MLLM(ChatGPT-4)提示词中,能否提升其基于街景图像评估建成环境步行质量的能力。我们从现有文献中收集步行性指标,并利用相关本体进行分类。选取行人安全与吸引力两个子主题,构建相应提示词。通过不同清晰度与具体性的提示词测试模型在评估步行性子主题上的表现。结果表明,尽管大模型可基于通用知识提供评估与解释,支持多模态图文评估自动化,但通常给出偏乐观分数,且在解读指标时常出错,导致误评。而引入专家知识后,模型评估表现更一致、更集中。

原文摘要 · Abstract (English)

Urban street environments are vital to supporting human activity in public spaces. The emergence of big data, such as street view images (SVIs) combined with multimodal large language models (MLLMs), is transforming how researchers and practitioners investigate, measure, and evaluate semantic and visual elements of urban environments. Considering the low threshold for creating automated evaluative workflows using MLLMs, it is crucial to explore both the risks and opportunities associated with these probabilistic models. In particular, the extent to which the integration of expert knowledge can influence the performance of MLLMs in evaluating the quality of urban design has not been fully explored. This study sets out an initial exploration of how integrating more formal and structured representations of expert urban design knowledge into the input prompts of an MLLM (ChatGPT-4) can enhance the model's capability and reliability in evaluating the walkability of built environments using SVIs. We collect walkability metrics from the existing literature and categorize them using relevant ontologies. We then select a subset of these metrics, focusing on the subthemes of pedestrian safety and attractiveness, and develop prompts for the MLLM accordingly. We analyze the MLLM's ability to evaluate SVI walkability subthemes through prompts with varying levels of clarity and specificity regarding evaluation criteria. Our experiments demonstrate that MLLMs are capable of providing assessments and interpretations based on general knowledge and can support the automation of multimodal image-text evaluations. However, they generally provide more optimistic scores and can make mistakes when interpreting the provided metrics, resulting in incorrect evaluations. By integrating expert knowledge, the MLLM's evaluative performance exhibits higher consistency and concentration.

城市设计大模型评估步行性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。