arXiv:2601.09043stat.MLcs.LG2026-01

提出一种自适应稀疏专家选择的贝叶斯模型,提升大语言模型中专家激活效率。

Horseshoe Mixtures-of-Experts (HS-MoE)

  • 结合霍尔舍收缩先验与输入相关门控,实现专家使用的数据自适应稀疏化。
  • 设计粒子学习算法,仅追踪充分统计量即可完成序列推断。
  • 适用于大模型中极低稀疏度场景,适合关注高效专家调度的研究者。

霍尔舍混合专家(HS-MoE)模型为混合专家架构中的稀疏专家选择提供了贝叶斯框架。通过将霍尔舍先验的自适应全局-局部收缩特性与输入依赖的门控机制结合,实现了专家使用上的数据自适应稀疏性。主要方法论贡献是一种用于序列推断的粒子学习算法,该算法在时间上向前传播滤波器的同时仅追踪充分统计量。此外,我们探讨了HS-MoE与大型语言模型中现代混合专家层的关系,这些层在极端稀疏约束下部署(例如,每标记仅激活少量专家,从大量专家池中选出)。

原文摘要 · Abstract (English)

Horseshoe mixtures-of-experts (HS-MoE) models provide a Bayesian framework for sparse expert selection in mixture-of-experts architectures. We combine the horseshoe prior's adaptive global-local shrinkage with input-dependent gating, yielding data-adaptive sparsity in expert usage. Our primary methodological contribution is a particle learning algorithm for sequential inference, in which the filter is propagated forward in time while tracking only sufficient statistics. We also discuss how HS-MoE relates to modern mixture-of-experts layers in large language models, which are deployed under extreme sparsity constraints (e.g., activating a small number of experts per token out of a large pool).

贝叶斯模型专家系统稀疏性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。