用多模态大模型自动发现城市设计与道路安全的隐藏关联。
From Street Views to Urban Science: Discovering Road Safety Factors with Multimodal Large Language Models
- 通过大模型分析街景图像生成安全相关问题并提取可解释特征。
- 在曼哈顿路段验证中优于预训练深度模型,且系数显著性可解释。
- 适合城市规划、交通政策制定者用于数据驱动的科学决策。
城市与交通研究长期致力于揭示关键变量与道路安全等社会结果之间的统计关系,以指导城市与交通系统的规划与更新。然而传统方法面临三大挑战:依赖人工专家提出假设(耗时且易受确认偏倚影响);深度学习模型可解释性差;未充分利用蕴含重要城市语境的非结构化数据。为此,本文提出基于多模态大语言模型(MLLM)的可解释假说推断方法,实现城市语境与道路安全关系的自动化假设生成、评估与迭代优化。该方法利用MLLM对街景图像(SVIs)提出安全相关问题,从响应中提取可解释嵌入,并用于回归模型。UrbanX支持基于统计显著性的迭代测试与修正,从而发现此前被忽视的城市设计与安全间的关联。在曼哈顿街段的实验表明,本方法性能超越预训练深度学习模型,同时保持完全可解释性。该框架可扩展至多元社会经济与环境结果,为城市科学研究提供通用范式,提升模型可信度,建立可规模化、基于统计证据的可解释知识发现路径。
原文摘要 · Abstract (English)
Urban and transportation research has long sought to uncover statistically meaningful relationships between key variables and societal outcomes such as road safety, to generate actionable insights that guide the planning, development, and renewal of urban and transportation systems. However, traditional workflows face several key challenges: (1) reliance on human experts to propose hypotheses, which is time-consuming and prone to confirmation bias; (2) limited interpretability, particularly in deep learning approaches; and (3) underutilization of unstructured data that can encode critical urban context. Given these limitations, we propose a Multimodal Large Language Model (MLLM)-based approach for interpretable hypothesis inference, enabling the automated generation, evaluation, and refinement of hypotheses concerning urban context and road safety outcomes. Our method leverages MLLMs to craft safety-relevant questions for street view images (SVIs), extract interpretable embeddings from their responses, and apply them in regression-based statistical models. UrbanX supports iterative hypothesis testing and refinement, guided by statistical evidence such as coefficient significance, thereby enabling rigorous scientific discovery of previously overlooked correlations between urban design and safety. Experimental evaluations on Manhattan street segments demonstrate that our approach outperforms pretrained deep learning models while offering full interpretability. Beyond road safety, UrbanX can serve as a general-purpose framework for urban scientific discovery, extracting structured insights from unstructured urban data across diverse socioeconomic and environmental outcomes. This approach enhances model trustworthiness for policy applications and establishes a scalable, statistically grounded pathway for interpretable knowledge discovery in urban and transportation studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。