arXiv:2409.13622q-bio.QMcs.LG2024-09

提出一种模拟视觉皮层前区处理机制的高效自编码器模型。

pAE: An Efficient Autoencoder Architecture for Modeling the Lateral Geniculate Nucleus by Integrating Feedforward and Feedback Streams in Human Visual System

  • 基于剪枝自编码器架构,融合前馈与反馈通路建模
  • 在时序模式下达到99.26%预测准确率,比人类高28%
  • 适用于自然图像中动物运动识别,适合神经科学与视觉模型研究者

视觉皮层是大脑中负责分层识别物体的关键区域。理解外侧膝状体(LGN)作为视觉皮层前驱区域,在自下而上与自上而下路径中的作用至关重要。当视觉刺激抵达视网膜后,会先传递至LGN进行初步处理,再送至视觉皮层进一步分析。本文提出一种深度卷积模型,以逼近人类视觉信息处理过程。通过设计基于剪枝自编码器(pAE)的浅层卷积模型,实现对LGN功能的近似。该模型整合了来自初级视觉皮层(V1)的前馈与反馈流。建模框架涵盖自然图像数据集的时序与非时序输入模式,数据由固定相机连续拍摄,分为含动物(运动)和不含动物两类。实验对比了本方法与基于Gabor及双正交小波函数的滤波器组方法。结果表明,所提出的深度调优模型不仅与人类基准表现高度相似,且显著优于其他模型。pAE模型在时序模式下达到99.26%的预测性能,较人类水平提升约28%。

原文摘要 · Abstract (English)

The visual cortex is a vital part of the brain, responsible for hierarchically identifying objects. Understanding the role of the lateral geniculate nucleus (LGN) as a prior region of the visual cortex is crucial when processing visual information in both bottom-up and top-down pathways. When visual stimuli reach the retina, they are transmitted to the LGN area for initial processing before being sent to the visual cortex for further processing. In this study, we introduce a deep convolutional model that closely approximates human visual information processing. We aim to approximate the function for the LGN area using a trained shallow convolutional model which is designed based on a pruned autoencoder (pAE) architecture. The pAE model attempts to integrate feed forward and feedback streams from/to the V1 area into the problem. This modeling framework encompasses both temporal and non-temporal data feeding modes of the visual stimuli dataset containing natural images captured by a fixed camera in consecutive frames, featuring two categories: images with animals (in motion), and images without animals. Subsequently, we compare the results of our proposed deep-tuned model with wavelet filter bank methods employing Gabor and biorthogonal wavelet functions. Our experiments reveal that the proposed method based on the deep-tuned model not only achieves results with high similarity in comparison with human benchmarks but also performs significantly better than other models. The pAE model achieves the final 99.26% prediction performance and demonstrates a notable improvement of around 28% over human results in the temporal mode.

视觉建模自编码器神经科学深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。