arXiv:2602.18178cs.CV2026-02被引 2

对比ViT与人类在图形感知任务中的表现,发现ViT存在明显感知差距。

Evaluating Graphical Perception Capabilities of Vision Transformers

  • 基于Cleveland和McGill的视觉编码研究,设计控制实验评估模型
  • ViT在基础视觉任务中表现强,但图形感知能力不如人类和CNN
  • 揭示ViT在可视化系统中的应用局限,适合关注人机感知差异的研究者

视觉变换器(Vision Transformers, ViTs)已成为图像任务中卷积神经网络(CNNs)的强大替代方案。尽管此前已对CNN在图形感知任务中的能力进行过评估,但ViTs在该领域的感知能力仍鲜有研究。本文受Cleveland和McGill奠基性研究启发,设计了一系列受控的图形感知任务,对比了ViTs、CNNs与人类参与者的表现。结果表明,尽管ViTs在通用视觉任务中表现出色,但在可视化领域与人类图形感知的一致性有限。本研究揭示了关键的感知差距,为ViTs在可视化系统及图形感知建模中的应用提供了重要参考。

原文摘要 · Abstract (English)

Vision Transformers, ViTs, have emerged as a powerful alternative to convolutional neural networks, CNNs, in a variety of image-based tasks. While CNNs have previously been evaluated for their ability to perform graphical perception tasks, which are essential for interpreting visualizations, the perceptual capabilities of ViTs remain largely unexplored. In this work, we investigate the performance of ViTs in elementary visual judgment tasks inspired by the foundational studies of Cleveland and McGill, which quantified the accuracy of human perception across different visual encodings. Inspired by their study, we benchmark ViTs against CNNs and human participants in a series of controlled graphical perception tasks. Our results reveal that, although ViTs demonstrate strong performance in general vision tasks, their alignment with human-like graphical perception in the visualization domain is limited. This study highlights key perceptual gaps and points to important considerations for the application of ViTs in visualization systems and graphical perceptual modeling.

视觉感知ViT可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。