提出无需调参的专家数量选择方法,解决软门控混合专家模型的识别难题。
Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency Without Model Sweeps
- 基于沃罗诺伊损失与混合度量树状图,建立统一统计框架
- 在过参数化下实现参数估计最优收敛率,准确恢复专家数量
- 适用于真实数据(如玉米抗旱蛋白组学),避免多次训练
我们构建了软门控高斯混合专家(SGMoE)的统一统计框架,解决了参数估计与模型选择中的三大长期问题:(i) 门控参数的平移不变性,(ii) 门控与专家间的内在耦合关系,(iii) softmax诱导的条件密度中分子分母强耦合。提出与门控划分几何一致的沃罗诺伊型损失函数,建立了最大似然估计(MLE)的有限样本收敛率。在过参数化模型中,揭示了MLE收敛率与刻画近非可识别方向的多项式方程组可解性之间的关联。针对模型选择,将混合度量树状图适配至SGMoE,实现无需模型遍历的一致性选择,在过拟合下达到点态最优参数率且避免多规模训练。合成数据模拟验证理论,准确恢复专家数,实现预测收敛率并逼近回归函数。在模型误设(如ε-污染)下,树状图准则仍能恢复真实成分数,而AIC、BIC与积分完整似然则随样本量增大过度选型。在玉米抗旱性状的蛋白质组数据上,树状图引导的SGMoE选择两个专家,呈现清晰的混合度量层级,早期稳定似然,并生成可解释的基因型-表型映射,优于标准准则且无需多尺度训练。
原文摘要 · Abstract (English)
We develop a unified statistical framework for softmax-gated Gaussian mixture of experts (SGMoE) that addresses three long-standing obstacles in parameter estimation and model selection: (i) non-identifiability of gating parameters up to common translations, (ii) intrinsic gate-expert interactions that induce coupled differential relations in the likelihood, and (iii) the tight numerator-denominator coupling in the softmax-induced conditional density. Our approach introduces Voronoi-type loss functions aligned with the gate-partition geometry and establishes finite-sample convergence rates for the maximum likelihood estimator (MLE). In over-specified models, we reveal a link between the MLE's convergence rate and the solvability of an associated system of polynomial equations characterizing near-nonidentifiable directions. For model selection, we adapt dendrograms of mixing measures to SGMoE, yielding a consistent, sweep-free selector of the number of experts that attains pointwise-optimal parameter rates under overfitting while avoiding multi-size training. Simulations on synthetic data corroborate the theory, accurately recovering the expert count and achieving the predicted rates for parameter estimation while closely approximating the regression function. Under model misspecification (e.g., $ε$-contamination), the dendrogram selection criterion is robust, recovering the true number of mixture components, while the Akaike information criterion, the Bayesian information criterion, and the integrated completed likelihood tend to overselect as sample size grows. On a maize proteomics dataset of drought-responsive traits, our dendrogram-guided SGMoE selects two experts, exposes a clear mixing-measure hierarchy, stabilizes the likelihood early, and yields interpretable genotype-phenotype maps, outperforming standard criteria without multi-size training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。