用人类感知的彩色模型评估视觉模型,发现不同模型对颜色的理解差异显著。
Beyond Color Geometry: Evaluating Human-Like Color Representations in Vision Models

- 基于86个渐进色类别的感知模型,评估颜色分类边界、紧凑性和渐变一致性。
- 掩码自编码器在颜色渐变对齐上表现最佳,远超其他视觉模型。
- 语言监督模型更关注物体颜色,而自编码器全局表征表面颜色。
视觉模型是否像人一样感知颜色?现有评估多依赖几何空间(如CIELAB)或离散标签,仅反映感知距离或类别归属,无法捕捉人类对颜色的渐进组织方式。本文引入基于86个渐进色类别的模糊感知模型,该模型由人类调查数据拟合而成。该框架适用于任意图像编码器,可衡量三类互补属性:类别边界、类别紧凑性,以及超越颜色几何的渐变对齐。在11个视觉变换器编码器中,类别层面结果相近,但渐变对齐差异显著。掩码自编码器(MAE)在超越几何的对齐上表现最强,置信区间与其他模型无重叠。层分析显示,掩码重建能保持这一结构至输出层。在自然图像上,MAE全局表征表面颜色,而语言监督模型则更强调前景物体的颜色。结果表明,人类类似的颜色表征包含多个独立维度,不应简化为单一评分。
原文摘要 · Abstract (English)
Do vision models see colors the way humans do? Existing evaluations of color representations usually compare them with geometric spaces such as CIELAB or with discrete color labels. These references capture perceptual distance or category membership, but not the graded way in which people organize colors. We evaluate color grounding against a fuzzy perceptual model with 86 graded categories fitted to human survey data. The framework can be applied to any image encoder and measures three complementary properties: category boundaries, category compactness, and graded alignment beyond what color geometry alone can explain. Across eleven Vision Transformer encoders, the category-level results are broadly similar, whereas graded alignment differs substantially. Masked Autoencoders achieve the strongest beyond-geometry alignment, with confidence intervals that do not overlap those of the other encoders. A layer-wise analysis further shows that masked reconstruction preserves this structure toward the output. On natural images, MAE represents surface color globally, while language-supervised models encode color more strongly in relation to the foreground object. These results show that human-like color grounding has several distinct aspects that should not be reduced to a single score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。