用生物视觉机制增强卷积网络,提升图像分类效果
Convolution goes higher-order: a biologically inspired mechanism empowers image classification
- 引入可学习的高阶卷积,模拟生物视觉中的非线性交互
- 在多个数据集上优于传统CNN,3~4阶展开表现最佳
- 适合关注生物启发模型与高效视觉计算的研究者
我们提出一种受复杂非线性生物视觉处理启发的新图像分类方法,将经典卷积神经网络(CNN)升级为具备可学习高阶卷积的架构。模型采用类似Volterra展开的卷积算子,捕捉早期与高级视觉加工中观察到的乘积型交互。在合成数据集上评估了对高阶相关性的敏感度,并在标准基准(MNIST、FashionMNIST、CIFAR10、CIFAR100和Imagenette)上验证性能。结果表明,该架构优于传统CNN基线,在3~4阶展开时达到最优性能,且与自然图像像素强度分布高度一致。通过系统扰动分析,验证了特定图像统计量对模型表现的贡献,揭示不同阶次卷积处理视觉信息的不同方面。表征相似性分析显示网络各层具有不同的几何结构,表明存在质异的视觉信息处理模式。本工作连接神经科学与深度学习,为更有效的生物启发式计算机视觉模型提供路径,尤其适用于资源受限场景。
原文摘要 · Abstract (English)
We propose a novel approach to image classification inspired by complex nonlinear biological visual processing, whereby classical convolutional neural networks (CNNs) are equipped with learnable higher-order convolutions. Our model incorporates a Volterra-like expansion of the convolution operator, capturing multiplicative interactions akin to those observed in early and advanced stages of biological visual processing. We evaluated this approach on synthetic datasets by measuring sensitivity to testing higher-order correlations and performance in standard benchmarks (MNIST, FashionMNIST, CIFAR10, CIFAR100 and Imagenette). Our architecture outperforms traditional CNN baselines, and achieves optimal performance with expansions up to 3rd/4th order, aligning remarkably well with the distribution of pixel intensities in natural images. Through systematic perturbation analysis, we validate this alignment by isolating the contributions of specific image statistics to model performance, demonstrating how different orders of convolution process distinct aspects of visual information. Furthermore, Representational Similarity Analysis reveals distinct geometries across network layers, indicating qualitatively different modes of visual information processing. Our work bridges neuroscience and deep learning, offering a path towards more effective, biologically inspired computer vision models. It provides insights into visual information processing and lays the groundwork for neural networks that better capture complex visual patterns, particularly in resource-constrained scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。