测试1.14亿个真实注视点,发现中心标记比训练模型更准,且现有模型存在群体偏差。
Human versus Computer Vision

- 用真实眼球数据对比计算机视觉注意力模型,发现未训练的中心点预测更优
- 模型准确率存在系统性偏差,对年轻、白人、温和群体更友好
- 提出基于群体自身注视行为评估模型泛化能力的新方法
本文从观众视角测试主流计算机视觉显著性模型,基于3,023名美国成年人在新闻照片上的1.14亿个网络摄像头注视点数据。结果显示,未经训练的中心标记比所有训练过的深度网络表现更优,因为模型增加的内容落在真实观众从不注视的位置。剩余的预测准确率存在系统性偏差,对年轻、白人及中立立场群体更准确,而对年长、非裔及极端意识形态群体预测效果差。本文提出一种新方法:通过一个群体自身的注视模式来判断模型能否学习该群体,并在样本支持的所有人口统计维度上应用该方法。研究揭示了视觉系统如何真正实现对所有人的可见性,并为相关主张提供了可衡量的标准。
原文摘要 · Abstract (English)
Computer vision saliency models predict where people will look, one map per image, and a billion-dollar predicted-attention industry sells those maps in place of measuring real viewers. I test the leading models from the audience side, against 11.4 million webcam gaze points from 3,023 US adults recruited to national quotas, viewing circulating news photographs. I show that an untrained central marker outperforms every trained network, because the content the networks add on top of the center falls where these audiences never look. What accuracy remains is systematically biased, favoring younger, White, and moderate viewers over older, Black, and ideologically extreme ones. I propose a way forward and build on what a group's own gaze reveals about whether a model can learn that group at all, and I apply it across every demographic axis this sample supports. Ultimately, I show how systems that decide what people see can learn to see everyone, and this study supplies the standard by which such a claim should be judged.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。