探究视觉Transformer注意力与人类审美关注的匹配度
Vision Transformer attention alignment with human visual perception in aesthetic object evaluation
- 用眼动实验和ViT注意力图对比人类与AI对工艺品的关注点
- 第12个注意力头在sigma=2.4时与人眼注意力相关性最强
- 部分注意力头可模拟人类对细节的关注,适合设计评估应用
视觉注意力机制在人类感知与审美评价中起关键作用。尽管视觉Transformer(ViTs)在计算机视觉任务中表现优异,但其与人类视觉注意力模式的对齐程度,尤其是在审美场景下仍缺乏研究。本研究探讨了人类在评估手工制品时的视觉注意力与ViT注意力机制之间的关联性。我们招募了30名参与者(9名女性,21名男性,平均年龄24.6岁),观看20件手工艺品(包括编筐包和姜罐)。使用Pupil Labs眼动仪记录注视轨迹并生成热力图以表征人类注意力。同时,利用预训练的DINO ViT模型,从12个注意力头中提取每张图像的注意力图。通过不同高斯参数(sigma=0.1至3.0)下的KL散度比较人类与ViT注意力分布。统计分析显示,在sigma=2.4±0.03时相关性最佳,其中第12号注意力头与人类注意力模式最接近。注意力头间存在显著差异,第7和第9号头与人类注意力偏差最大(p<0.05,Tukey HSD检验)。结果表明,虽然ViTs表现出更全局化的注意力模式,但某些注意力头能较好拟合人类对特定特征(如编筐扣件)的关注,提示其在产品设计与美学评估中的潜在应用价值,同时也揭示了当前人工智能与人类认知在注意力策略上的本质差异。
原文摘要 · Abstract (English)
Visual attention mechanisms play a crucial role in human perception and aesthetic evaluation. Recent advances in Vision Transformers (ViTs) have demonstrated remarkable capabilities in computer vision tasks, yet their alignment with human visual attention patterns remains underexplored, particularly in aesthetic contexts. This study investigates the correlation between human visual attention and ViT attention mechanisms when evaluating handcrafted objects. We conducted an eye-tracking experiment with 30 participants (9 female, 21 male, mean age 24.6 years) who viewed 20 artisanal objects comprising basketry bags and ginger jars. Using a Pupil Labs eye-tracker, we recorded gaze patterns and generated heat maps representing human visual attention. Simultaneously, we analyzed the same objects using a pre-trained ViT model with DINO (Self-DIstillation with NO Labels), extracting attention maps from each of the 12 attention heads. We compared human and ViT attention distributions using Kullback-Leibler divergence across varying Gaussian parameters (sigma=0.1 to 3.0). Statistical analysis revealed optimal correlation at sigma=2.4 +-0.03, with attention head #12 showing the strongest alignment with human visual patterns. Significant differences were found between attention heads, with heads #7 and #9 demonstrating the greatest divergence from human attention (p< 0.05, Tukey HSD test). Results indicate that while ViTs exhibit more global attention patterns compared to human focal attention, certain attention heads can approximate human visual behavior, particularly for specific object features like buckles in basketry items. These findings suggest potential applications of ViT attention mechanisms in product design and aesthetic evaluation, while highlighting fundamental differences in attention strategies between human perception and current AI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。