受灵长类视觉皮层启发,构建几何驱动的神经网络,提升模型效率与鲁棒性。
The Geometry of Cortical Computation: Manifold Disentanglement and Predictive Dynamics in VCNet
- 模仿视觉皮层分层处理与双流分离机制,学习低维几何结构表示
- 在纹理识别和光场图像任务中分别达到92.1%和74.4%准确率,超越同类模型
- 适合关注生物启发架构、几何建模与鲁棒视觉的科研人员
尽管现代卷积神经网络表现优异,仍存在数据效率低、分布外泛化差、易受对抗扰动等问题,根源在于缺乏反映视觉世界固有几何结构的先验。相比之下,灵长类视觉系统表现出更高的效率与鲁棒性,其演化形成的架构与计算原理或可为更强大的人工视觉提供蓝图。本文提出视觉皮层网络(VCNet),其设计基于灵长类视觉皮层的宏观组织结构。VCNet被构建成一个几何框架,模拟关键生物机制:跨皮层区域的分层处理、用于学习解耦表征的双流信息分离,以及用于表征优化的自上而下预测反馈。我们从几何与动力系统角度诠释这些机制,认为它们引导学习结构化、低维的神经流形。我们在两个专项基准上评估:Spots-10动物图案数据集(测试对自然纹理的敏感性)与光场图像分类任务(需处理高维视觉数据)。结果表明,VCNet在Spots-10上取得92.1%的准确率,在光场数据集上达74.4%,均超过同等规模的现有模型。本工作证明,将高层次神经科学原理以几何视角整合,可实现更高效、更鲁棒的模型,为解决机器学习长期挑战提供新方向。
原文摘要 · Abstract (English)
Despite their success, modern convolutional neural networks (CNNs) exhibit fundamental limitations, including data inefficiency, poor out-of-distribution generalization, and vulnerability to adversarial perturbations. These shortcomings can be traced to a lack of inductive biases that reflect the inherent geometric structure of the visual world. The primate visual system, in contrast, demonstrates superior efficiency and robustness, suggesting that its architectural and computational principles,which evolved to internalize these structures,may offer a blueprint for more capable artificial vision. This paper introduces Visual Cortex Network (VCNet), a novel neural network architecture whose design is informed by the macro-scale organization of the primate visual cortex. VCNet is framed as a geometric framework that emulates key biological mechanisms, including hierarchical processing across distinct cortical areas, dual-stream information segregation for learning disentangled representations, and top-down predictive feedback for representation refinement. We interpret these mechanisms through the lens of geometry and dynamical systems, positing that they guide the learning of structured, low-dimensional neural manifolds. We evaluate VCNet on two specialized benchmarks: the Spots-10 animal pattern dataset, which probes sensitivity to natural textures, and a light field image classification task, which requires processing higher-dimensional visual data. Our results show that VCNet achieves state-of-the-art accuracy of 92.1\% on Spots-10 and 74.4\% on the light field dataset, surpassing contemporary models of comparable size. This work demonstrates that integrating high-level neuroscientific principles, viewed through a geometric lens, can lead to more efficient and robust models, providing a promising direction for addressing long-standing challenges in machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。