arXiv:2504.09393cs.CVcs.LG2025-04

ViT模型展现类人视觉偏见,包括方向、颜色敏感与感知分类。

Vision Transformers Exhibit Human-Like Biases: Evidence of Orientation and Color Selectivity, Categorical Perception, and Phase Transitions

  • 用合成数据测试ViT在角度和颜色上的预测误差分布。
  • 水平方向误差最小,蓝调色误差最高,且颜色分组符合人类感知类别。
  • 注意力头在特定层自发形成通用特征提取能力,类似人脑机制。

本研究探究了视觉变换器(ViTs)是否具备与人脑相似的方向与颜色偏见。通过控制噪声水平、角度、长度、宽度和颜色的合成数据集,分析了使用LoRA微调后的ViT行为。结果揭示四个关键发现:第一,ViTs表现出倾斜效应,在所有条件下180度(水平)方向的预测误差最低;第二,角度预测误差随颜色变化,蓝色系误差最高,黄色系最低;聚类分析显示,ViTs对颜色的分组方式与人类感知类别一致。此外,观察到相变现象:在所有条件下均出现两次相变,但当颜色作为额外数据属性引入时,训练损失曲线的相变出现延迟。最后,某些层的注意力头展现出无需下游任务即可提取通用特征的能力。这些发现表明,偏差和特性主要源于原始数据集的预训练及视觉变换器的固有架构约束,而非仅由下游数据统计决定。

原文摘要 · Abstract (English)

This study explored whether Vision Transformers (ViTs) developed orientation and color biases similar to those observed in the human brain. Using synthetic datasets with controlled variations in noise levels, angles, lengths, widths, and colors, we analyzed the behavior of ViTs fine-tuned with LoRA. Our findings revealed four key insights: First, ViTs exhibited an oblique effect showing the lowest angle prediction errors at 180 deg (horizontal) across all conditions. Second, angle prediction errors varied by color. Errors were highest for bluish hues and lowest for yellowish ones. Additionally, clustering analysis of angle prediction errors showed that ViTs grouped colors in a way that aligned with human perceptual categories. In addition to orientation and color biases, we observed phase transition phenomena. While two phase transitions occurred consistently across all conditions, the training loss curves exhibited delayed transitions when color was incorporated as an additional data attribute. Finally, we observed that attention heads in certain layers inherently develop specialized capabilities, functioning as task-agnostic feature extractors regardless of the downstream task. These observations suggest that biases and properties arise primarily from pre-training on the original dataset which shapes the model's foundational representations and the inherent architectural constraints of the vision transformer, rather than being solely determined by downstream data statistics.

视觉变换器类人偏见注意力机制感知分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。