用简单模型揭示视觉皮层和扩散模型的推理机制
Toward a mechanistic understanding of inference in visual cortex and diffusion models

- 基于稀疏编码构建可解释的循环动力系统
- 训练后连接矩阵匹配大脑视觉皮层水平连接结构
- 揭示扩散模型生成新图像的内在机理,适合神经科学与机器学习交叉研究
我们提出一个等价于最小扩散模型的初级视觉皮层(V1)感知推理模型,其功能可直接由参数理解。该模型基于带有非因子化先验的稀疏编码,采用无约束成对交互矩阵,扩展为通用的循环动力系统。通过去噪得分匹配和隐式微分高效训练。在自然图像上训练后,学习到的交互矩阵再现了V1浅层神经元间与相同取向偏好相连的水平连接结构。模型表现出极佳的去噪能力,在极端视觉模糊下仍能恢复长轮廓,其泛化性能接近标准黑箱扩散架构。由于模型简洁,可通过潜变量间交互矩阵直接分解网络雅可比矩阵,揭示循环动力如何赋予连续自然结构变形高概率。有趣的是,大量潜变量完全脱离视觉输入,形成层级表示,实现图像特征的全局一致性。该模型连接了神经科学与机器学习:为神经回路在感知推理中的功能连接提供可验证假设;揭示扩散模型从有限训练集生成无限新图像的内部机制。
原文摘要 · Abstract (English)
We describe a model of perceptual inference in primary visual cortex (V1) equivalent to a minimal diffusion model whose function can be readily understood from its parameters. The model is based on sparse coding with a non-factorial prior over latent variables in the form of an unconstrained, pairwise interaction matrix, extending standard sparse coding inference to a general recurrent dynamical system. We efficiently train these recurrent dynamics using a denoising score-matching objective and implicit differentiation. After training on natural images, the learned interaction matrix mirrors the structure of horizontal connections in superficial layers of V1 that link neurons of similar orientation tuning. This model exhibits exceptionally good denoising performance, restoring image features such as extended contours amid extreme visual ambiguity, nearly matching the behavior of standard, black-box diffusion architectures in generalization regime. Owing to the model's simplicity, the network's Jacobian can be decomposed directly in terms of the interaction matrix between latent variables, revealing mechanistically how the recurrent dynamics assign high probability over a continuous family of natural structural deformations. Intriguingly, within this circuit, a large fraction of latent variables learn to disconnect from visual input altogether, essentially forming a hierarchical representation that appears to enforce global consistency among image features. Together, the model and results bridge two distinct domains: for neuroscience, it generates concrete, testable hypotheses regarding functional connectivity in recurrent neural circuits during perceptual inference tasks; for machine learning, it elucidates the internal mechanisms learned by diffusion models that allow them to generate infinitely many novel images from a finite training set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。