MoE模型在特征噪声下更鲁棒,因稀疏专家激活具噪声过滤作用。
Robustness of Mixtures of Experts to Feature Noise
- 利用稀疏专家激活机制过滤特征噪声,实现模块化计算。
- 相比密集网络,泛化误差更低,收敛更快,抗扰动能力更强。
- 适合高噪声场景下的高效模型设计,尤其语言任务中表现优异。
尽管混合专家(MoE)模型在实践中取得成功,但其性能超越密集网络的原因尚不明确,尤其是超出参数量规模的影响。本文在参数量一致的设定下,研究输入具有潜在模块结构但受特征噪声干扰的情形,将噪声视为内部激活的代理。结果表明,稀疏的专家激活起到了噪声过滤的作用:相较于密集估计器,MoE在特征噪声下表现出更低的泛化误差、更强的抗扰动能力以及更快的收敛速度。合成数据和真实语言任务的实验验证了理论发现,显示稀疏模块化计算带来稳定的鲁棒性与效率提升。
原文摘要 · Abstract (English)
Despite their practical success, it remains unclear why Mixture of Experts (MoE) models can outperform dense networks beyond sheer parameter scaling. We study an iso-parameter regime where inputs exhibit latent modular structure but are corrupted by feature noise, a proxy for noisy internal activations. We show that sparse expert activation acts as a noise filter: compared to a dense estimator, MoEs achieve lower generalization error under feature noise, improved robustness to perturbations, and faster convergence speed. Empirical results on synthetic data and real-world language tasks corroborate the theoretical insights, demonstrating consistent robustness and efficiency gains from sparse modular computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。