提出一种更灵活的增量随机优化算法,适用于混合专家模型。
Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts
- 用随机增量方法改进MM算法,无需显式隐变量表示
- 在混合专家回归中表现优于SGD、Adam等主流优化器
- 适合处理高维异构数据,如生物信息学中的干旱胁迫研究
现代统计与机器学习中,高流量流式数据处理日益普遍,批量算法因需多次遍历全数据集而难以应用。本文重新审视并分析了一种增量随机极大化极小(MM)算法,该算法是增量随机期望最大化(EM)的推广,且不依赖显式隐变量表示,适用性更广、灵活性更高。理论证明了该算法的收敛性:迭代序列趋于目标函数梯度为零的驻点。在软最大门控混合专家(MoE)回归任务上验证了其有效性,该任务无现成的随机EM算法可用。实验表明,该方法在性能上持续优于随机梯度下降、RMSProp、Adam及二阶截断随机优化等常用优化器。此外,在两个真实世界数据集上均取得稳定预测提升,包括一项整合高维蛋白质组学与生态生理特征的玉米耐旱基因型研究,充分展示了其实际价值。
原文摘要 · Abstract (English)
Processing high-volume, streaming data is increasingly common in modern statistics and machine learning, where batch-mode algorithms are often impractical because they require repeated passes over the full dataset. This has motivated incremental stochastic estimation methods, including the incremental stochastic Expectation-Maximization (EM) algorithm formulated via stochastic approximation. In this work, we revisit and analyze an incremental stochastic variant of the Majorization-Minimization (MM) algorithm, which generalizes incremental stochastic EM as a special case. Our approach relaxes key EM requirements, such as explicit latent-variable representations, enabling broader applicability and greater algorithmic flexibility. We establish theoretical guarantees for the incremental stochastic MM algorithm, proving consistency in the sense that the iterates converge to a stationary point characterized by a vanishing gradient of the objective. We demonstrate these advantages on a softmax-gated mixture of experts (MoE) regression problem, for which no stochastic EM algorithm is available. Empirically, our method consistently outperforms widely used stochastic optimizers, including stochastic gradient descent, root mean square propagation, adaptive moment estimation, and second-order clipped stochastic optimization. These results support the development of new incremental stochastic algorithms, given the central role of softmax-gated MoE architectures in contemporary deep neural networks for heterogeneous data modeling. Beyond synthetic experiments, we also validate practical effectiveness on two real-world datasets, including a bioinformatics study of dent maize genotypes under drought stress that integrates high-dimensional proteomics with ecophysiological traits, where incremental stochastic MM yields stable gains in predictive performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。