发现主流神经网络的纹理表征与人类感知不匹配,挑战了现有模型的有效性。
Perceptual misalignment of texture representations in convolutional neural networks

- 用特征相关性量化CNN的纹理表征能力
- 发现纹理感知与视觉系统建模质量无关联
- 提示纹理感知依赖上下文整合等额外机制
对视觉纹理的数学建模可追溯至Julesz的猜想:人类纹理感知基于图像特征间的局部相关性。一种有影响力的方法将此思想推广到卷积神经网络(CNN)的非线性特征之间的线性相关性,通过格拉姆矩阵表示。由于CNN常被用作视觉系统的模型,自然要问:这些“纹理表征”是否自发地与纹理的感知内容对齐?特别是,被认为更接近视觉系统的CNN是否也具有更类人的纹理表征?我们量化了多种CNN在特征相关性上捕捉的感知内容,并将其与它们在Brain-Score指标中对哺乳动物视觉系统的感知对齐度进行比较。出乎意料的是,我们发现传统衡量CNN作为视觉系统模型优劣的指标与其与人类纹理感知的对齐度之间没有关联。结论是:纹理感知涉及与当前基于物体识别训练的CNN方法所建模机制不同的过程,可能依赖于上下文信息的整合。
原文摘要 · Abstract (English)
Mathematical modeling of visual textures traces back to Julesz's intuition that texture perception in humans is based on local correlations between image features. An influential approach for texture analysis and generation generalizes this notion to linear correlations between the nonlinear features computed by convolutional neural networks (CNNs), compiled into Gram matrices. Given that CNNs are often used as models for the visual system, it is natural to ask whether such "texture representations" spontaneously align with the textures' perceptual content, and in particular whether those CNNs that are regarded as better models for the visual system also possess more human-like texture representations. Here we quantify the perceptual content captured by feature correlations computed for a diverse pool of CNNs, and we compare it to the models' perceptual alignment with the mammalian visual system as measured by Brain-Score. Surprisingly, we find that there is no connection between conventional measures of CNN quality as a model of the visual system and its alignment with human texture perception. We conclude that texture perception involves mechanisms that are distinct from those that are commonly modeled using approaches based on CNNs trained on object recognition, possibly depending on the integration of contextual information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。