测试视觉语言模型对城市场景的理解是否像人一样。
Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark
- 用蒙特利尔街景构建100张真实与合成图像的感知基准。
- 模型在客观属性上表现优于主观评价,最高多标签得分0.48。
- 适合参与式城市设计与模型可解释性研究者使用。
理解人类如何感知城市环境可为城市设计与规划提供支持。我们引入一个小型基准,用于评估视觉语言模型(VLMs)在城市感知任务中的表现,涵盖100张蒙特利尔街景图像,真实与逼真合成图像各半。来自七个社区团体的12名参与者提供了30个维度的230份标注表,包含物理属性与主观感受。法语回答经标准化转为英文。我们在零样本设置下,使用结构化提示与确定性解析器评估了七种VLMs。单选题采用准确率,多标签题使用杰卡德重叠率;人类一致性通过克里彭多夫阿尔法与成对杰卡德衡量。结果表明,模型在可见的客观属性上对齐程度更高,而主观判断差距较大。表现最佳系统(claude-sonnet)在多标签任务中达到宏观准确率0.31、平均杰卡德0.48。人类一致性越高,模型得分越好。合成图像略降低性能。我们公开该基准、提示与工具集,支持可复现、具备不确定性意识的参与式城市分析评估。
原文摘要 · Abstract (English)
Understanding how people read city scenes can inform design and planning. We introduce a small benchmark for testing vision-language models (VLMs) on urban perception using 100 Montreal street images, evenly split between photographs and photorealistic synthetic scenes. Twelve participants from seven community groups supplied 230 annotation forms across 30 dimensions mixing physical attributes and subjective impressions. French responses were normalized to English. We evaluated seven VLMs in a zero-shot setup with a structured prompt and deterministic parser. We use accuracy for single-choice items and Jaccard overlap for multi-label items; human agreement uses Krippendorff's alpha and pairwise Jaccard. Results suggest stronger model alignment on visible, objective properties than subjective appraisals. The top system (claude-sonnet) reaches macro 0.31 and mean Jaccard 0.48 on multi-label items. Higher human agreement coincides with better model scores. Synthetic images slightly lower scores. We release the benchmark, prompts, and harness for reproducible, uncertainty-aware evaluation in participatory urban analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。