通过高阶卷积提升视网膜神经响应预测精度,仅用一半数据就达到0.75相关性。
Higher-Order Convolution Improves Neural Predictivity in the Retina
- 在卷积核内嵌入高阶运算,直接建模像素间时空乘积关系。
- 在蝾螈和小鼠视网膜数据上,相关性达0.75,是基线模型的两倍以上。
- 特别适合预测检测迫近物体的视网膜细胞响应,对缩放参数建模更精准。
我们提出一种新方法,在卷积神经网络(CNN)中直接引入高阶操作,以提升神经响应预测能力。该模型扩展了传统3D CNN,将高阶运算嵌入卷积算子,直接建模空间与时间上邻近像素间的乘积交互。相比加深网络结构,本方法在不增加深度的前提下增强了表达能力,缓解了人工网络与生物视觉系统浅层处理之间的架构差异。我们在两个数据集上评估:蝾螈视网膜神经节细胞(RGC)对自然场景的响应,以及新采集的小鼠RGC对受控几何变换的响应。所提高阶卷积神经网络(HoCNN)性能更优,且仅需标准架构一半的训练数据,与真实神经反应的相关系数最高达0.75(对照组可靠性为0.80±0.02)。集成到先进模型后,跨物种与刺激条件均表现一致提升。分析显示,模型自然编码基本几何变换,尤其在表征物体缩放参数方面显著优于基线模型——对应相关系数从0.32提升至0.72,尤其对瞬时OFF-alpha和瞬时ON细胞的响应预测改进明显。
原文摘要 · Abstract (English)
We present a novel approach to neural response prediction that incorporates higher-order operations directly within convolutional neural networks (CNNs). Our model extends traditional 3D CNNs by embedding higher-order operations within the convolutional operator itself, enabling direct modeling of multiplicative interactions between neighboring pixels across space and time. Our model increases the representational power of CNNs without increasing their depth, therefore addressing the architectural disparity between deep artificial networks and the relatively shallow processing hierarchy of biological visual systems. We evaluate our approach on two distinct datasets: salamander retinal ganglion cell (RGC) responses to natural scenes, and a new dataset of mouse RGC responses to controlled geometric transformations. Our higher-order CNN (HoCNN) achieves superior performance while requiring only half the training data compared to standard architectures, demonstrating correlation coefficients up to 0.75 with neural responses (against 0.80$\pm$0.02 retinal reliability). When integrated into state-of-the-art architectures, our approach consistently improves performance across different species and stimulus conditions. Analysis of the learned representations reveals that our network naturally encodes fundamental geometric transformations, particularly scaling parameters that characterize object expansion and contraction. This capability is especially relevant for specific cell types, such as transient OFF-alpha and transient ON cells, which are known to detect looming objects and object motion respectively, and where our model shows marked improvement in response prediction. The correlation coefficients for scaling parameters are more than twice as high in HoCNN (0.72) compared to baseline models (0.32).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。