arXiv:2608.06934cs.LGcs.CV2026-08

用多模态模型捕捉不同人对步行友好度的主观差异

Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning

论文配图:Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning
图 1 · 摘自论文原文
  • 融合行人视角图像与评价者个人特征建模感知差异
  • 行人视角图片评分比车行视角高,影像源影响判断
  • 考虑评价者身份可提升预测准确率65%,适合城市规划

步行友好度的视觉感知因人而异,反映个体特征、经历与偏好的差异。现有研究常将多样判断简化为平均分,隐含感知一致假设,并依赖车行街景图像,未能体现行人真实视觉体验。本文构建了包含29,870条评分、来自1,196名受访者的数据集,关联澳大利亚城乡及郊区的行人视角图像与评价者属性,提出首个用户条件化的多模态深度学习框架,融合视觉特征与个体表征。视角对比实验表明,行人视角图像获得显著更高的步行友好度评分,说明图像来源是感知调查中的关键设计选择。用户条件化模型相比仅基于图像的基线,排名一致性提升65%(二次加权肯德尔系数0.47 vs. 0.29),证明评价者身份本身具有超越图像内容的预测价值。研究支持从统一的、观察者无关的评分转向体现多元用户的评估模型,推动更包容的步行环境评价。

原文摘要 · Abstract (English)

Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and preferences. Existing studies, however, often reduce these diverse judgements to aggregated scores, implicitly assuming uniform perception, and commonly rely on vehicle-mounted street-view imagery that does not reflect the pedestrian's visual experience. This paper introduces a dataset of 29,870 walkability ratings from 1,196 respondents, linking sidewalk-view imagery across urban, suburban, and regional Australian environments with individual rater attributes, and proposes the first user-conditioned multimodal deep learning framework for walkability perception, fusing visual features with respondent-level representations. A viewpoint-comparison study shows that sidewalk-view images receive significantly higher walkability ratings than matched street-view images, indicating that imagery source is a substantive design decision in perception surveys. The user-conditioned model improves rank agreement with observed ratings by 65% over an image-only baseline (quadratic weighted kappa 0.47 vs. 0.29), demonstrating that who is evaluating an environment carries predictive indication beyond image content alone. These findings support moving from aggregated, observer-independent walkability scores toward models that represent diverse users, enabling more inclusive assessment of pedestrian environments.

步行友好度多模态学习用户建模城市规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。