arXiv:2606.07642cs.CVcs.CY2026-06

用专家指导的视觉语言模型评估街景中的轮椅通行障碍,效果接近真实行为数据。

Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View

论文配图:Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View
图 1 · 摘自论文原文
  • 结合街景图像与专家规则,构建可扩展的无障碍评估框架。
  • VLM评分与轮椅停留时间负相关且分布相似,验证其有效性。
  • 适合城市规划、无障碍设计人员参考,尤其关注可见障碍物识别。

评估建筑环境互动(如轮椅无障碍性)困难,因真实移动受分布式、情境依赖及临时障碍影响,难以规模化捕捉。本文探讨视觉语言模型(VLMs)能否从谷歌街景(GSV)图像中识别无障碍障碍。提出一种专家引导的检索增强框架,融合GSV图像、符合《美国残疾人法案》(ADA)的指导原则及专家制定的评估标准,以评估多个无障碍维度。在佛罗里达大学校园收集了包含407个唯一街景位置的数据集,并通过GPS获取的轮椅停留行为作为移动摩擦信号。结果显示,VLM评分与停留时间呈负相关且分布一致,表明其与真实移动摩擦代理存在部分但稳定的匹配。视觉线索分析显示,缘石坡道、人行横道等明显设施与较高VLM无障碍评分相关,但对细微路面状况、临时障碍和视角依赖性障碍的识别仍有限。总体表明,专家引导的VLM可用于与传感器数据对齐的可扩展无障碍评估。

原文摘要 · Abstract (English)

Assessing built-environment interaction, such as wheelchair accessibility, is difficult because real-world mobility is shaped by distributed, context-dependent, and temporary barriers that are hard to capture at scale. To support scalable assessment, this paper examines whether vision-language models (VLMs) can identify accessibility barriers from Google Street View (GSV) imagery. We propose an expert-guided retrieval-augmented framework that combines GSV images, ADA-informed guidance, and expert-derived rubrics to evaluate accessibility dimensions. We collect a campus-scale dataset at the University of Florida, linking 407 unique GSV locations with GPS-derived wheelchair dwell behavior as a mobility-friction signal. Results show that VLM ratings are both negatively correlated and distributionally similar with dwell time, indicating partial but consistent alignment with a behavioral proxy for mobility friction. Visual cue analysis shows that certain environmental objects, such as curb ramps and crosswalks, are associated with higher VLM accessibility scores, while alignment remains limited for subtle surface conditions, transient obstructions, and viewpoint-dependent barriers. Overall, our findings show the potential of expert-guided VLMs for scalable accessibility assessment aligning with sensor-derived indicators of real-world wheelchair navigation.

无障碍评估视觉语言模型街景分析城市规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。