提出快速模型选择与稳定优化方法,提升软门控专家混合模型的训练可靠性。
Fast Model Selection and Stable Optimization for Softmax-Gated Multinomial-Logistic Mixture of Experts Models
- 用显式二次下界构造批量MM算法,实现坐标闭式更新,保证目标单调上升。
- 理论证明有限样本下条件密度估计与参数恢复率,实验准确率优于基线。
- 无需遍历,通过融合冗余专家原子实现近最优模型选择,适合高维分类任务。
专家混合(MoE)架构通过学习门控机制组合专用预测器,在回归与分类中均表现良好。然而,对于采用软门控多分类逻辑门控的模型,其最大似然训练的稳定性保障与合理模型选择仍缺乏严格理论支持。本文在全数据(批量)设置下解决这两个问题:首先,基于显式二次下界推导出批量最小化-最大化(MM)算法,实现坐标闭式更新,确保目标函数单调递增并全局收敛至驻点,避免了类似EM算法中常见的近似M步;其次,证明了条件密度估计与参数恢复的有限样本速率,并将混合测度的树状图方法拓展至分类场景,实现无遍历的专家数量选择,经合并冗余拟合原子后达到近参数最优率。在生物蛋白-蛋白相互作用预测任务上的实验验证了完整流程的有效性,相比强统计与机器学习基线,模型精度更高且概率校准更优。
原文摘要 · Abstract (English)
Mixture-of-Experts (MoE) architectures combine specialized predictors through a learned gate and are effective across regression and classification, but for classification with softmax multinomial-logistic gating, rigorous guarantees for stable maximum-likelihood training and principled model selection remain limited. We address both issues in the full-data (batch) regime. First, we derive a batch minorization-maximization (MM) algorithm for softmax-gated multinomial-logistic MoE using an explicit quadratic minorizer, yielding coordinate-wise closed-form updates that guarantee monotone ascent of the objective and global convergence to a stationary point (in the standard MM sense), avoiding approximate M-steps common in EM-type implementations. Second, we prove finite-sample rates for conditional density estimation and parameter recovery, and we adapt dendrograms of mixing measures to the classification setting to obtain a sweep-free selector of the number of experts that achieves near-parametric optimal rates after merging redundant fitted atoms. Experiments on biological protein--protein interaction prediction validate the full pipeline, delivering improved accuracy and better-calibrated probabilities than strong statistical and machine-learning baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。