arXiv:2510.03151cs.LG2025-10

用量化理论分析专家模型,揭示专家数量与精度的权衡关系。

Mixture of Many Zero-Compute Experts: A High-Rate Quantization Theory Perspective

  • 将输入空间分块,每块用零计算常数专家,基于高率量化理论建模。
  • 一维情况下可精确求解最优分块与测试误差;多维时给出误差上界。
  • 理论结合实验,说明专家数量影响近似与估计误差的平衡。

本文运用经典高率量化理论,为回归任务中的混合专家(MoE)模型提供新视角。我们的MoE通过输入空间分割定义,每个区域对应一个单参数专家,推理时为零计算的常数预测器。在专家数量足够大的假设下,各区域极小,便于研究模型类的近似误差:(i) 对一维输入,我们推导出测试误差及其最优分割和专家设置;(ii) 对多维输入,我们给出测试误差的上界并研究其最小化。此外,在给定输入空间分割的前提下,我们研究专家参数从训练数据学习的统计性质,从而理论上和实证上揭示了MoE学习中近似误差与估计误差之间的权衡如何依赖于专家数量。

原文摘要 · Abstract (English)

This paper uses classical high-rate quantization theory to provide new insights into mixture-of-experts (MoE) models for regression tasks. Our MoE is defined by a segmentation of the input space to regions, each with a single-parameter expert that acts as a constant predictor with zero-compute at inference. Motivated by high-rate quantization theory assumptions, we assume that the number of experts is sufficiently large to make their input-space regions very small. This lets us to study the approximation error of our MoE model class: (i) for one-dimensional inputs, we formulate the test error and its minimizing segmentation and experts; (ii) for multidimensional inputs, we formulate an upper bound for the test error and study its minimization. Moreover, we consider the learning of the expert parameters from a training dataset, given an input-space segmentation, and formulate their statistical learning properties. This leads us to theoretically and empirically show how the tradeoff between approximation and estimation errors in MoE learning depends on the number of experts.

专家模型量化理论近似误差统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。