用眼动数据训练模型,让机器看路更像人。
Learning to See Like Humans: Gaze-Aligned Cycling Safety Prediction

- 用眼动数据引导视觉注意力,让模型关注人类真正注意的区域
- 在安全感知任务中表现媲美顶尖方法,且注意力图更贴近真人注视模式
- 适合做城市感知分析、交通设计优化的研究者与决策者
骑行对公共健康和环境有显著益处,但城市中人们对道路安全的感知常限制其推广。当街景看起来不安全时,人们更不愿骑行,因此感知成为主要障碍。已有研究通过对比街景图像来学习主观安全判断,但未显式建模人类视觉注意力——而这是人类感知安全的核心机制。本文提出眼球追踪引导的感知骑行安全框架(EG-PCS),将眼动数据融入基于视觉变换器的成对学习流程。通过用眼动信号监督模型注意力机制,促使学习到的注意力图与人类注视模式对齐。实验表明,该方法在排序性能上达到现有最优水平,同时注意力图更准确反映人类视觉行为。结果证明,引入眼动信息能提升感知分析的预测精度与可解释性。
原文摘要 · Abstract (English)
Cycling delivers significant public-health and environmental benefits, yet its uptake in cities is often limited by perceived safety. When street environments appear unsafe, individuals are less likely to cycle, making perception a key barrier to adoption. Recent work has shown that pairwise comparisons of street-view images provide a scalable way to learn subjective safety judgments. However, existing approaches do not explicitly model human visual attention, which plays a central role in how humans perceive safety. We propose an Eye-Tracking-Guided Perceived Cycling Safety framework (EG-PCS) that integrates gaze data into a pairwise learning pipeline based on vision transformers. By supervising the model's attention mechanism with eye-tracking signals, we encourage alignment between learned attention maps and human fixation patterns. Experiments show that gaze-guided models achieve similar ranking performance compared to state-of-the-art approaches while producing attention maps that more accurately reflect human visual attention behavior. Our results demonstrate that incorporating eye-tracking information enhances both predictive accuracy and interpretability in perception-based urban analytics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。