探究视觉大模型是否模仿了人眼低层视觉特性
Do computer vision foundation models learn the low-level characteristics of the human visual system?
- 用9种测试评估45个模型的低层视觉能力
- DINOv2等模型在对比度掩蔽上与人类最接近
- 多数模型对低对比度不敏感且响应不规律
计算机视觉基础模型(如DINO或OpenCLIP)在大规模图像数据集上通过自监督方式训练。类似地,大量证据表明人类视觉系统(HVS)受自然世界中颜色和图案统计分布的影响,这些特征也存在于基础模型的训练数据中。本文探讨了在自然图像上训练的基础模型是否模拟了人眼的某些低层视觉特性,如对比度检测、对比度掩蔽和对比度恒常性。我们设计了包含九种测试类型的协议,评估45个基础和生成模型的图像编码器。结果表明,部分基础模型(如DINO、DINOv2和OpenCLIP)表现出与人类视觉部分相似的特性,而其他模型则差异显著。基础模型对低对比度的敏感性较低,且在不同频率下的对比度响应不规则。在对比度掩蔽方面,基础模型与人类数据的吻合度最高。研究提示,人类视觉与计算机视觉在理解真实世界图像时可能走相似也不同的路径。总体而言,尽管仍存在差异,但基于视觉任务训练的基础模型已开始向低层人类视觉靠拢,其中DINOv2表现最为接近。
原文摘要 · Abstract (English)
Computer vision foundation models, such as DINO or OpenCLIP, are trained in a self-supervised manner on large image datasets. Analogously, substantial evidence suggests that the human visual system (HVS) is influenced by the statistical distribution of colors and patterns in the natural world, characteristics also present in the training data of foundation models. The question we address in this paper is whether foundation models trained on natural images mimic some of the low-level characteristics of the human visual system, such as contrast detection, contrast masking, and contrast constancy. Specifically, we designed a protocol comprising nine test types to evaluate the image encoders of 45 foundation and generative models. Our results indicate that some foundation models (e.g., DINO, DINOv2, and OpenCLIP), share some of the characteristics of human vision, but other models show little resemblance. Foundation models tend to show smaller sensitivity to low contrast and rather irregular responses to contrast across frequencies. The foundation models show the best agreement with human data in terms of contrast masking. Our findings suggest that human vision and computer vision may take both similar and different paths when learning to interpret images of the real world. Overall, while differences remain, foundation models trained on vision tasks start to align with low-level human vision, with DINOv2 showing the closest resemblance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。