融合视觉与行为数据,提升城市步行友好度预测精度
WalkCLIP: Multimodal Learning for Urban Walkability Prediction
- 用多模态数据融合卫星、街景与人口动态信息
- 在明尼阿波利斯-圣保罗4660个地点验证,表现优于单一模态模型
- 适合城市规划、交通设计与公共健康研究者参考
城市步行友好度是公共健康、可持续发展和生活质量的核心。传统评估依赖调查与实地审计,成本高且难以扩展。近期研究使用卫星影像、街景图像或人口指标估算步行友好度,但单源方法仅反映环境某一维度:卫星数据从高空描述建成环境,却忽略行人视角;街景图像捕捉地面状况,缺乏空间上下文;人口动态揭示人类活动模式,却不反映视觉形态。本文提出WalkCLIP,一种多模态框架,整合这些互补视角以预测步行友好度。该模型利用GPT-4o生成的图像描述学习具备步行感知的视觉-语言表征,通过空间聚合模块引入邻里上下文,再融合来自人口动态基础模型的表征。在明尼阿波利斯-圣保罗4,660个地点的评估显示,WalkCLIP在预测准确性和空间一致性上均优于单模态与多模态基线。结果表明,视觉与行为信号的融合可实现对步行环境的可靠预测。
原文摘要 · Abstract (English)
Urban walkability is a cornerstone of public health, sustainability, and quality of life. Traditional walkability assessments rely on surveys and field audits, which are costly and difficult to scale. Recent studies have used satellite imagery, street view imagery, or population indicators to estimate walkability, but these single-source approaches capture only one dimension of the walking environment. Satellite data describe the built environment from above, but overlook the pedestrian perspective. Street view imagery captures conditions at the ground level, but lacks broader spatial context. Population dynamics reveal patterns of human activity but not the visual form of the environment. We introduce WalkCLIP, a multimodal framework that integrates these complementary viewpoints to predict urban walkability. WalkCLIP learns walkability-aware vision-language representations from GPT-4o generated image captions, refines these representations with a spatial aggregation module that incorporates neighborhood context, and fuses the resulting features with representations from a population dynamics foundation model. Evaluated at 4,660 locations throughout Minneapolis-Saint Paul, WalkCLIP outperforms unimodal and multimodal baselines in both predictive accuracy and spatial alignment. These results show that the integration of visual and behavioral signals yields reliable predictions of the walking environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。