arXiv:2602.18250cs.LG2026-02

让神经元变成概率分布,实现不确定性显式计算。

Variational Distributional Neuron

  • 神经元以变分自编码器结构构建,携带先验、后验和局部ELBO
  • 激活值为分布而非标量,通过约束收缩连续可能性空间
  • 适合研究不确定性建模与生成模型的可解释性

我们提出一种变分分布神经元的概念验证:将计算单元构造成一个变分自编码器模块,显式包含先验、近似后验和局部ELBO。该单元不再是确定性的标量,而是一个分布——计算不再传播数值,而是受约束地收缩连续的可能性空间。每个神经元参数化后验,传播重参数化采样,并通过局部ELBO的KL项正则化,因此激活值具有分布特性。这种‘收缩’可通过局部约束测试,并由内部度量监控。单元所携带的上下文信息量及其时间持久性,由不同约束分别调控。该工作解决了一个结构性矛盾:在序列生成中,因果性主要存在于符号空间,即使存在隐变量,它们也常为辅助角色,而有效动态仍由高度确定性的解码器承载;同时,概率隐变量模型虽能捕捉变化因子与不确定性,但不确定性通常由全局或参数机制承担,而单元仍传播标量。因此核心问题在于:若不确定性是计算的本质,为何计算单元不显式携带它?我们由此提出两个维度:(i)概率约束的组合需稳定、可解释且可控;(ii)粒度:若推理是分布间在约束下的协商,基础单元是否应保持确定性,还是变为分布?我们分析了‘坍缩’模式与‘活神经元’的条件,并通过每单元的自回归隐变量先验拓展其贡献随时间演化。

原文摘要 · Abstract (English)

We propose a proof of concept for a variational distributional neuron: a compute unit formulated as a VAE brick, explicitly carrying a prior, an amortized posterior and a local ELBO. The unit is no longer a deterministic scalar but a distribution: computing is no longer about propagating values, but about contracting a continuous space of possibilities under constraints. Each neuron parameterizes a posterior, propagates a reparameterized sample and is regularized by the KL term of a local ELBO - hence, the activation is distributional. This "contraction" becomes testable through local constraints and can be monitored via internal measures. The amount of contextual information carried by the unit, as well as the temporal persistence of this information, are locally tuned by distinct constraints. This proposal addresses a structural tension: in sequential generation, causality is predominantly organized in the symbolic space and, even when latents exist, they often remain auxiliary, while the effective dynamics are carried by a largely deterministic decoder. In parallel, probabilistic latent models capture factors of variation and uncertainty, but that uncertainty typically remains borne by global or parametric mechanisms, while units continue to propagate scalars - hence the pivot question: if uncertainty is intrinsic to computation, why does the compute unit not carry it explicitly? We therefore draw two axes: (i) the composition of probabilistic constraints, which must be made stable, interpretable and controllable; and (ii) granularity: if inference is a negotiation of distributions under constraints, should the primitive unit remain deterministic or become distributional? We analyze "collapse" modes and the conditions for a "living neuron", then extend the contribution over time via autoregressive priors over the latent, per unit.

概率计算神经元建模不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。