测试50多个视觉模型,发现它们颜色感知能力与人眼差距大。
Do Vision Encoders Exhibit Human-like Color Thresholds?

- 用可控色度刺激对比模型与人眼的色彩敏感度
- 最佳模型与人眼匹配度不足0.25(mIoU)
- 自监督模型表现更好,语言监督模型两极分化
理解人类色彩感知是长期研究目标,关键指标是人眼可察觉的最小色差阈值。近年来,深度视觉编码器作为大规模视觉任务的标准模型,将图像映射为潜在特征表示。然而,很少有研究探讨其内部表征是否具备类人色彩阈值。本文对超过50个预训练视觉编码器(包括卷积网络与视觉变换器)进行了大规模探索性研究,使用多级色度刺激,通过区域重叠度量(mIoU)比较模型推导出的色彩辨别阈值与人眼辨别椭圆。结果显示,所有模型家族中,模型表征与人类感知阈值的对齐程度普遍较弱,最优mIoU < 0.25。此外,自监督编码器整体优于监督模型,而语言监督模型表现两极分化,既出现最优也出现最差结果。表明当前大规模视觉训练目标无法自然催生类人色觉敏感性。
原文摘要 · Abstract (English)
Understanding and characterizing human color perception is a longstanding research goal. One of the most traditional approaches is looking for the human color discrimination thresholds, the minimum chromatic differences perceptible to human observers. In recent years, deep neural networks have become the standard networks for computer vision tasks. In particular, deep vision encoders, foundation models trained on large-scale visual data, map images into latent feature representations. Despite the widespread use of deep vision encoders, few studies have investigated whether their internal representations exhibit human-like discrimination thresholds. In this work, we present a large-scale exploratory study probing the chromatic sensitivity of more than 50 pretrained vision encoders, including convolutional networks and vision transformers, against human discrimination thresholds. Using controlled chromatic stimuli at multiple chroma levels, we compare model-derived chromatic discrimination thresholds with human discrimination ellipses through a region-overlap metric (mIoU). Our analysis reveals generally weak alignment between model representations and human perceptual thresholds across all model families, with the best mIoU < 0.25. Moreover, we find that self-supervised encoders consistently outperform supervised ones, while language-supervised models show the most polarized behavior, occupying both the top and bottom of the ranking. These findings suggest that human-like chromatic sensitivity does not emerge naturally from current large-scale visual training objectives for any of the analyzed architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。