用眼动数据提升城市感知建模,让机器更懂人怎么感受城市。
Modeling Subjective Urban Perception with Human Gaze

- 融合眼动追踪与街景图像,构建多模态城市感知数据集。
- 眼动信息单独就能预测城市感知,结合场景语义后效果更好。
- 适合做城市规划、智能导航的学者和工程师参考。
城市感知描述了人们主观评价城市环境的方式,影响着城市的体验与理解。现有计算方法主要基于街景图像直接建模城市感知,但忽略了人类判断形成过程中的感知机制。本文提出 Place Pulse-Gaze 数据集,将街景图像与同步的眼动记录及个体感知标签相结合。基于此数据集,我们构建了眼动引导的城市感知框架,系统研究三种互补设置:仅使用眼动建模、眼动与显式语义场景表示融合、眼动与隐式丰富视觉表示融合。实验表明,仅靠眼动已具备有效的预测信号,且与场景表示融合后,在语义与丰富视觉表示下均能进一步提升预测性能。研究强调了将人类感知过程纳入城市场景理解的重要性,并为眼动引导的多模态城市计算开辟新方向。
原文摘要 · Abstract (English)
Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computational approaches primarily model urban perception directly from street view images, but largely ignore the human perceptual process through which such judgments are formed. In this paper, we introduce Place Pulse-Gaze, an urban perception dataset that augments street view images with synchronized eye-tracking recordings and individual perception labels. Based on this dataset, we propose a Gaze-Guided Urban Perception Framework to study how gaze behavior contributes to the modeling of subjective urban perception. The framework systematically investigates three complementary settings: gaze-only modeling, gaze fusion with explicit semantic scene representations, and gaze fusion with implicit richer visual representations. Experiments show that gaze alone already carries useful predictive signals for subjective urban perception, and that integrating gaze with scene representations further improves prediction under both semantic and richer visual representations. Overall, our findings highlight the importance of incorporating human perceptual processes into urban scene understanding and open a direction for gaze-guided multimodal urban computing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。