arXiv:2606.15202cs.CV2026-06

大模型能像人一样关注高风险场景中的关键区域,无需训练数据。

Comparing Human Gaze and Vision-Language Model Attention in Safety-Relevant Environments

论文配图:Comparing Human Gaze and Vision-Language Model Attention in Safety-Relevant Environments
图 1 · 摘自论文原文
  • 用人类眼动数据与大模型生成的注意力图对比,验证模型是否捕捉真实视觉焦点。
  • GPT-4o与人类注意力分布最接近,三项指标表现优异,尤其在分布匹配上最佳。
  • 无需眼动训练数据,模型即可模拟人类对高危场景的关注模式,适合安全评估应用。

人类视觉注意在感知潜在风险环境时起关键作用。本研究探究大型视觉语言模型能否识别出与人类在安全相关环境中注意力集中区域一致的场景部分。通过十名参与者佩戴Pupil Invisible可穿戴眼镜观看33张不同风险等级的场景图像,采集眼动数据并生成群体平均的人类注视热图。同时,使用OpenAI Vision API调用GPT-4o生成空间注意力预测,并转换为显著性图,与人类注视模式进行对比。采用四种互补指标评估空间一致性:皮尔逊相关系数(r = 0.515 ± 0.117)、归一化扫描路径显著性(NSS = 0.988 ± 0.323)、KL散度(KL = 1.766 ± 0.844)和基于Judd公式的受试者工作特征曲线下面积(AUC-Judd = 0.806 ± 0.076)。与Gemini Pro、Gemini Flash和Claude的交叉比较显示,所有模型均超过随机基线(AUC-Judd=0.5),且获得正的NSS值。Gemini Pro在三项指标中表现最强,而GPT-4o在KL散度上最接近人类注意力分布。结果表明,大型视觉语言模型可在未经过眼动数据训练的情况下,准确识别安全相关场景中人类注意力集中的区域,具备作为可扩展工具模拟人类注意模式的潜力。

原文摘要 · Abstract (English)

Human visual attention plays an important role in how people perceive and respond to environments containing potential risks. This study investigates whether large vision-language models can identify the same regions of a scene that attract human attention in safety-relevant environments. Eye-tracking data were collected from ten participants viewing 33 scene images representing environments with varying levels of potential risk using Pupil Invisible wearable glasses. Gaze coordinates were mapped onto stimulus images to generate population-averaged human gaze heatmaps. In parallel, GPT-4o was prompted through the OpenAI Vision Application Programming Interface (API) to generate spatial predictions of visual attention, which were converted into saliency maps for comparison with human gaze patterns. Spatial alignment between human gaze heatmaps and model-generated saliency maps was evaluated using four complementary metrics: Pearson correlation (r = 0.515 +- 0.117), Normalised Scanpath Saliency (NSS = 0.988 +- 0.323), Kullback-Leibler divergence (KL = 1.766 +- 0.844), and Area Under the Receiver Operating Characteristic Curve using the Judd formulation (AUC-Judd = 0.806 +- 0.076). A cross-model comparison with Gemini Pro, Gemini Flash, and Claude showed that all models exceeded the AUC-Judd chance baseline of 0.5 and achieved positive NSS scores. Gemini Pro demonstrated the strongest spatial localisation according to three of the four metrics, whereas GPT-4o produced the closest distributional match to human attention as measured by KL divergence. These findings suggest that large vision-language models can identify regions that broadly correspond to where humans direct visual attention in safety-relevant scenes without requiring eye-tracking training data. The results highlight the potential of vision-language models as a scalable tool for approximating human attentional patterns.

视觉注意大模型安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。