首次系统分析软最大化专家混合模型的贝叶斯性质,为实际建模提供理论依据。
On Bayesian Softmax-Gated Mixture-of-Experts Models
- 采用贝叶斯框架研究软最大化门控的专家混合模型
- 给出密度估计后验收缩率及参数收敛保证
- 提出两种选专家数策略,适合复杂建模场景
专家混合模型通过输入依赖的门控机制组合多个专家模型,灵活建模复杂的概率输入输出关系。这类模型在现代机器学习中日益重要,但其贝叶斯框架下的理论性质仍基本未被探索。本文研究贝叶斯专家混合模型,聚焦普遍使用的软最大化门控机制。我们分析了三个基础统计任务的后验渐近行为:密度估计、参数估计和模型选择。首先,我们在专家数量固定已知与随机可学习两种情形下,建立了密度估计的后验收缩率。其次,针对参数估计,基于定制的Voronoi型损失,给出了收敛性保证,该损失考虑了专家混合模型的复杂可识别性结构。最后,提出并分析了两种互补的专家数量选择策略。这些结果共同构成了对软最大化门控贝叶斯专家混合模型的首批系统性理论分析,为实际模型设计提供了理论指导。
原文摘要 · Abstract (English)
Mixture-of-experts models provide a flexible framework for learning complex probabilistic input-output relationships by combining multiple expert models through an input-dependent gating mechanism. These models have become increasingly prominent in modern machine learning, yet their theoretical properties in the Bayesian framework remain largely unexplored. In this paper, we study Bayesian mixture-of-experts models, focusing on the ubiquitous softmax-based gating mechanism. Specifically, we investigate the asymptotic behavior of the posterior distribution for three fundamental statistical tasks: density estimation, parameter estimation, and model selection. First, we establish posterior contraction rates for density estimation, both in the regimes with a fixed, known number of experts and with a random learnable number of experts. We then analyze parameter estimation and derive convergence guarantees based on tailored Voronoi-type losses, which account for the complex identifiability structure of mixture-of-experts models. Finally, we propose and analyze two complementary strategies for selecting the number of experts. Taken together, these results provide one of the first systematic theoretical analyses of Bayesian mixture-of-experts models with softmax gating, and yield several theory-grounded insights for practical model design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。