arXiv:2605.19155cs.CV2026-05

用高效编码原理,让模型从少量图像中学会类人视觉特征。

Efficient coding along the visual hierarchy

  • 用局部统计信息无监督压缩自然图像,逐层生成视觉特征。
  • 生成的特征能被人类识别,且与人脑视觉反应高度匹配。
  • 结合微调后,在数据少时更快速准确,适合小样本学习场景。

生物视觉系统在有限经验下即可学习,而深度学习模型依赖数百万张训练图像。是什么学习原则使这成为可能?我们检验了高效编码——即神经表征捕捉自然输入的统计结构——是否可从有限数据构建出类人的视觉层次特征。我们设计了一种无监督学习方法:深层网络每层仅使用局部统计信息,无需标签、任务或反向传播,将输入压缩到自然图像的主要变化模式上。该方法生成的特征从边缘、颜色逐步发展为纹理和形状。这些特征不仅被人类观察者轻易识别,还能预测人脑视觉皮层对图像的fMRI响应。此外,将高效编码与有监督微调结合的混合学习方式,在低数据环境下表现出更好的脑一致性,并实现更快的类别学习。结果表明,高效编码可能塑造整个视觉层级的表征,有助于解释生物视觉的数据高效性。

原文摘要 · Abstract (English)

Biological visual systems learn from limited experience, unlike deep learning models that rely on millions of training images. What learning principles make this possible? We tested whether efficient coding, the idea that neural representations capture the statistical structure of natural inputs, can build a hierarchy of human-aligned visual features from limited data. We developed an unsupervised learning procedure in which each layer of a deep network compresses its inputs onto the dominant modes of variation in natural images, using only local statistics and no labels, tasks, or backpropagation. This unsupervised procedure yields features that progress from edges and colors to textures and shapes. The features of this deep efficient coding model are readily recognized by human observers and are predictive of image-evoked fMRI responses in human visual cortex. Furthermore, a hybrid learning procedure that combines efficient coding with supervised fine-tuning yields better brain alignment in low-data settings and more rapid category learning. These findings suggest that efficient coding may shape representations across the entire visual hierarchy and help explain the data efficiency of biological vision.

视觉认知高效编码无监督学习脑机接口

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。