arXiv:2505.13052stat.MLcs.LG2025-05被引 4

提出新方法无需训练多个模型即可准确确定专家数量。

Model Selection for Gaussian-gated Gaussian Mixture of Experts Using Dendrograms of Mixing Measures

  • 基于混合度的树状图扩展,实现专家数一致估计
  • 在过拟合情况下仍达参数估计最优收敛率
  • 适合高维或深层网络中的模型选择问题

混合专家(MoE)模型是统计学与机器学习中广泛应用的集成学习方法,具有灵活性和计算高效性,已成为众多先进深度神经网络架构的核心组件,尤其适用于跨领域的异质数据分析。尽管应用广泛,其模型选择理论,特别是最优混合成分或专家数量的确定,仍不充分且面临重大挑战。这主要源于协变量同时出现在高斯门控函数和专家网络中,引入由参数偏微分方程支配的内在交互作用。本文重新审视混合度的树状图概念,提出一种针对高斯门控高斯MoE模型的新扩展方法,可实现真实混合成分数的一致估计,并在过拟合情形下达到参数估计的逐点最优收敛率。该方法无需训练并比较不同成分数的多个模型,显著降低计算负担,尤其适用于高维或深度网络场景。合成数据实验表明,该方法在准确恢复专家数量方面优于AIC、BIC及积分完成似然等常见准则,同时实现参数估计的最优收敛率并精确逼近回归函数。

原文摘要 · Abstract (English)

Mixture of Experts (MoE) models constitute a widely utilized class of ensemble learning approaches in statistics and machine learning, known for their flexibility and computational efficiency. They have become integral components in numerous state-of-the-art deep neural network architectures, particularly for analyzing heterogeneous data across diverse domains. Despite their practical success, the theoretical understanding of model selection, especially concerning the optimal number of mixture components or experts, remains limited and poses significant challenges. These challenges primarily stem from the inclusion of covariates in both the Gaussian gating functions and expert networks, which introduces intrinsic interactions governed by partial differential equations with respect to their parameters. In this paper, we revisit the concept of dendrograms of mixing measures and introduce a novel extension to Gaussian-gated Gaussian MoE models that enables consistent estimation of the true number of mixture components and achieves the pointwise optimal convergence rate for parameter estimation in overfitted scenarios. Notably, this approach circumvents the need to train and compare a range of models with varying numbers of components, thereby alleviating the computational burden, particularly in high-dimensional or deep neural network settings. Experimental results on synthetic data demonstrate the effectiveness of the proposed method in accurately recovering the number of experts. It outperforms common criteria such as the Akaike information criterion, the Bayesian information criterion, and the integrated completed likelihood, while achieving optimal convergence rates for parameter estimation and accurately approximating the regression function.

模型选择混合专家统计学习高斯混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。