发现稀疏专家模型需按任务复杂度动态激活专家数,才能更好泛化。
Sparse Mixture-of-Experts for Compositional Generalization: Empirical Evidence and Theoretical Foundations of Optimal Sparsity
- 按任务难度动态调整激活专家数量,而非固定最少激活。
- 实验显示最优激活专家数与任务复杂度成正比。
- 理论证明最优稀疏性介于极简与全激活之间,适合复杂推理任务。
稀疏混合专家(SMoE)架构因其在不显著增加计算成本的前提下扩展神经网络的能力而备受关注,尤其在变压器模型中表现突出。尽管如此,其在组合泛化(即适应已知成分的新组合)方面的作用仍缺乏深入研究。本研究挑战了‘最少专家激活即可完成任务泛化’的假设,探究了任务复杂度与最优稀疏性之间的关系。通过在SRAVEN符号推理任务和SKILL-MIX基准上的实证评估,我们发现:(i) 激活的专家数量随任务难度提升而持续增加以维持性能;(ii) 最优激活专家数与任务复杂度呈比例增长。理论分析推导出最优稀疏性的缩放定律,通过平衡近似误差与估计误差,揭示其与实证结果的一致性。我们正式证明,最优稀疏性位于最小激活(1-2个专家)与全激活之间,具体数量与任务复杂度成正比,并受训练数据规模与模型复杂度影响。这些发现为设计兼具计算效率与强组合泛化能力的SMoE模型提供了实用指导。
原文摘要 · Abstract (English)
Sparse Mixture-of-Experts (SMoE) architectures have gained prominence for their ability to scale neural networks, particularly transformers, without a proportional increase in computational cost. Despite their success, their role in compositional generalization, i.e., adapting to novel combinations of known components, remains under-explored. This study challenges the assumption that minimal expert activation suffices for task generalization and investigates the relationship between task complexity and optimal sparsity in SMoE models. Through empirical evaluations on the SRAVEN symbolic reasoning task and the SKILL-MIX benchmark, we demonstrate that (i) the number of activated experts consistently increases with the perceived task difficulty to maintain performance; and (ii) the optimal number of activated experts scales proportionally with task complexity. Our theoretical analysis derives a scaling law for optimal sparsity by balancing approximation and estimation errors, revealing alignment with empirical observations. We formally show that the optimal sparsity lies between minimal activation (1-2 experts) and full activation, with the exact number scaling proportionally to task complexity and further influenced by the size of the training data and the complexity of the model. These findings offer practical insights for designing SMoE models that achieve computational efficiency while enabling robust compositional generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。