评估城市感知的视觉语言模型需关注人类判断分歧与弃权,让可靠性可量化。
Benchmarks for Vision-Language Models in Urban Perception Should Be Reliability-Aware and Negotiated

- 将人类判断分歧和弃权视为测量结果,而非噪声
- 模型与人工共识一致度随维度可靠性变化,评估需报告可靠性指标
- 适合参与城市治理评估的科研人员与政策制定者参考
视觉语言模型(VLMs)越来越多用于生成街景图像的结构化描述,服务于街道审计、地图绘制和公众咨询等任务。这些应用融合可观测属性与评价类目,而人类评估者常呈现意见分歧和明确弃权。本文主张,评估城市感知的VLM应将分歧与弃权视为测量结果,报告人之间一致性(inter-annotator reliability) alongside 模型对齐程度,并视标签空间与评分策略为可协商的产物,尤其当输出用于城市治理时。研究基于100个蒙特利尔街景,由7个社区组织的12名参与者在30个维度上标注,结合7个VLM的确定性零样本评估。结果显示,模型与人类共识的一致性随维度的人类可靠性变化;在“整体印象”这一评价维度,模型与标注者存在分布不匹配,包括不同的“不适用”率。最后提出对基准创建者、模型开发者及机构的行动建议,使评估中的不确定性与假设透明化。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly used to generate structured descriptions of street-level imagery for tasks such as streetscape auditing, mapping, and public consultation. These uses combine observable attributes with appraisal categories, and the human targets are often distributions of judgments with disagreement and explicit non-response. This paper argues that benchmarking VLMs for urban perception should treat disagreement and abstention as measurement outcomes, report inter-annotator reliability alongside model alignment, and treat the label space and scoring policy as negotiable artifacts when outputs are intended to inform urban governance. We ground the argument in a benchmark of 100 Montreal street scenes annotated along 30 dimensions by 12 participants from seven community organizations, and in a deterministic zero-shot evaluation of seven VLMs. Across dimensions, model agreement with human consensus co-varies with dimension-level human reliability, and for the appraisal dimension Overall Impression models and annotators exhibit distributional mismatch including different rates of Not applicable. We close with actions for benchmark creators, model developers, and institutions to make uncertainty and benchmark assumptions visible in evaluation reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。