揭示专家混合模型在专家数趋近无穷时的统计行为,连接经典模型与量子神经网络
Mean-field limit from general mixtures of experts to quantum neural networks
- 通过梯度流训练专家混合模型,研究其参数分布随专家数量增长的渐近规律
- 证明当专家数无限时,参数经验测度逼近满足非线性连续性方程的概率分布
- 首次将该理论框架应用于量子神经网络生成的专家混合模型,为量子机器学习提供新视角
本文研究了在监督学习问题中,通过梯度流训练的专家混合模型(MoE)的渐近行为。核心结果是:当专家数量趋于无穷时,该模型表现出传播混沌现象。我们证明其参数的经验测度接近一个满足非线性连续性方程的概率测度,并给出了仅依赖于专家数量的显式收敛速率。该理论被应用于由量子神经网络生成的专家混合模型,为理解量子架构下的大规模集成学习提供了数学基础。
原文摘要 · Abstract (English)
In this work, we study the asymptotic behavior of Mixture of Experts (MoE) trained via gradient flow on supervised learning problems. Our main result establishes the propagation of chaos for a MoE as the number of experts diverges. We demonstrate that the corresponding empirical measure of their parameters is close to a probability measure that solves a nonlinear continuity equation, and we provide an explicit convergence rate that depends solely on the number of experts. We apply our results to a MoE generated by a quantum neural network.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。