随机神经网络集体行为中出现有序结构,温度可自适应优化分类性能。
Emergence of Structure in Ensembles of Random Neural Networks
- 用吉布斯分布加权随机分类器集合,通过损失函数定义能量
- 在高斯数据下,最优温度与教师模型和网络数量无关
- 理论+实验验证其普适性,适合理解随机系统中的自组织现象
随机性广泛存在于数据科学与机器学习中。值得注意的是,由随机组件构成的系统常表现出看似确定的宏观集体行为,体现从微观无序到宏观有序的转变。本文提出一个理论模型,研究随机分类器集合中集体行为的涌现。若通过以分类损失为能量的吉布斯测度对集成进行加权,则存在一个有限温度参数使分类性能达到最优(即损失最小)。当样本服从高斯分布、标签由教师感知机生成时,我们严格证明并数值验证:该最优温度既不依赖未知的教师分类器,也不依赖随机分类器的数量,表明该行为具有普适性。在MNIST数据集上的实验进一步证实此现象在高质量、无噪声数据中的重要性。最后,物理类比揭示了所研究现象的自组织本质。
原文摘要 · Abstract (English)
Randomness is ubiquitous in many applications across data science and machine learning. Remarkably, systems composed of random components often display emergent global behaviors that appear deterministic, manifesting a transition from microscopic disorder to macroscopic organization. In this work, we introduce a theoretical model for studying the emergence of collective behaviors in ensembles of random classifiers. We argue that, if the ensemble is weighted through the Gibbs measure defined by adopting the classification loss as an energy, then there exists a finite temperature parameter for the distribution such that the classification is optimal, with respect to the loss (or the energy). Interestingly, for the case in which samples are generated by a Gaussian distribution and labels are constructed by employing a teacher perceptron, we analytically prove and numerically confirm that such optimal temperature does not depend neither on the teacher classifier (which is, by construction of the learning problem, unknown), nor on the number of random classifiers, highlighting the universal nature of the observed behavior. Experiments on the MNIST dataset underline the relevance of this phenomenon in high-quality, noiseless, datasets. Finally, a physical analogy allows us to shed light on the self-organizing nature of the studied phenomenon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。