MLLMs可批量评估城市安全感知,但会带入性别年龄种族偏见。
Multimodal Large Language Models Predict Urban Safety Perception but Encode Non-Neutral Demographic Priors
- 用多模态大模型分析街景图判断安全感知,支持零样本推理。
- 模型对安全类别的预测准确率65%-69%,但普遍低估危险区域。
- 不同性别年龄种族身份提示导致判断显著差异,非中立立场。
理解人们对城市环境的安全感知对包容性规划至关重要,但传统调查成本高且难以扩展。本文研究多模态大语言模型(MLLMs)是否能从街景图像中评估感知安全,同时考虑感知的主观性。基于Place Pulse 2.0数据集,在56个城市中评估四种开源与专有MLLMs,使用中性提示及由性别、年龄、种族或族裔定义的社会人口学角色提示。分析了模型生成的解释关键词。所有模型均表现出相当的零样本能力,城市级宏观F1分数为65%–69%,并保留有意义的城市间差异。然而,模型系统性倾向“安全”类别,低估不安全情况,并压缩城市间差异。其解释收敛于共享视觉词汇:维护状况、绿化、秩序和居住特征支持“安全”判断;退化、孤立、照明差和行人活动少则支持“不安全”。在固定图像下,角色提示引发显著且结构化的判断变化:女性角色比男性角色更常判定为不安全;年龄影响因模型而异,中年角色总体最接近中性;黑人/非裔美国人和美洲原住民角色偏差最大,而最接近的种族匹配因模型而异。结果表明,尽管MLLMs能提供可扩展的安全感知信号,但并非来自人口学中立立场。
原文摘要 · Abstract (English)
Understanding how people perceive urban environments is essential for inclusive planning, yet conventional surveys are costly and difficult to scale. We investigate whether Multimodal Large Language Models (MLLMs) can assess perceived urban safety from street-view imagery while accounting for the observer-dependent nature of perception. Using Place Pulse 2.0, we evaluate four open and proprietary MLLMs across 56 cities under a Neutral prompt and socio-demographic personas defined by gender, age, and race or ethnicity. We also analyse the keywords generated to justify each classification. All four models display comparable zero-shot capability, with city-macro F1 scores of 65--69%, and preserve meaningful cross-city variation. However, they systematically favour the Safe class, underpredict unsafety, and compress differences between cities. Their explanations converge on a shared visual lexicon: maintenance, greenery, order, and residential character support Safe judgements, whereas deterioration, isolation, poor lighting, and limited pedestrian activity support Unsafe judgements. Persona prompting produces substantial and structured shifts while holding the image fixed. Female personas yield more Unsafe classifications than Male personas across all models; age effects are model-dependent, although Middle-aged personas generally remain closest to Neutral. Black/African American and Native American personas frequently show the largest departures, while the closest race or ethnicity match varies by model. These findings show that MLLMs can provide scalable signals of perceived urban safety, but not from a demographically neutral standpoint.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。