测试发现前沿专家模型的模块化功能大多不成立,结果依赖测量方式。
How Modular Is a Frontier Mixture-of-Experts? A Pre-registered Causal Test in Which Apparent Expert Modularity Mostly Dissolves

- 通过因果干预和预注册假设,系统测试专家功能模块性。
- 仅阿拉伯语专家家族在独立数据集上保持稳定选择性,其余均不成立。
- 结果受数据集、指标和统计标准影响,需严格控制变量。
稀疏混合专家(MoE)模型将每个标记路由至少数专家,暗示专家可能形成与能力或语言相关的功能模块。我们在前沿开源权重模型Command A+(总参数218B,活跃参数25B;128个专家,每次激活8个,外加1个共享专家)上进行因果测试。构建路由质量图谱,预注册六组家族-轴心假设,在推理时对各家族进行消融,并与大小匹配的随机专家基线对比,检验其是否仅影响自身轴心(最差的非目标影响不超过目标影响的三分之一)。关键是在四种度量标准下,以及在独立语料库上的自助法置信区间验证。结果警示:稳健的功能模块性极为罕见且依赖测量方式。六组预注册家族中,仅阿拉伯语家族在独立语料库和保守统计标准下表现稳定(1/6;更宽松标准可得3/6,但该结果对阈值敏感)。其余家族虽有实际因果效应,却缺乏选择性,其表观模块性随语料、指标和统计阈值而变化。在Qwen3-30B-A3B上的正向对照重现了已发表的分离结构,确认方法能检测到真实模块性。结论在未量化BF16模型上同样成立,排除4比特量化干扰。我们得出结论:除非控制语料、指标和统计标准,否则基于消融的模块性判断不可靠。相关图谱和消融数据已公开。
原文摘要 · Abstract (English)
Sparse Mixture-of-Experts (MoE) models route each token to a few of many experts, inviting the hypothesis that experts form functional modules tied to capabilities or languages. We test this causally on Command A+, a frontier open-weights MoE (218B total / 25B active; 128 experts, 8 active, +1 shared). We build a routing-mass atlas, pre-register six family-to-axis hypotheses before any intervention, and ablate each family at inference time against a size-matched random-expert null, measuring whether it selectively breaks its own axis (worst off-target effect at most one third of on-target). Crucially, we test the same families under four metrics and a held-out, independent-corpus run with bootstrap confidence intervals. Our finding is cautionary: robust functional modularity is rare and measurement-dependent. Of six pre-registered families, only one, the Arabic-language family, is a clean selective module that survives an independent corpus and a conservative statistical bar (1/6; a more permissive pre-registered point rule admits 3/6, but that count is threshold-sensitive). Every other family has a real causal effect yet fails selectivity, and its apparent modularity flips with the measurement: with the corpus, the metric, and the statistical bar. A positive control on Qwen3-30B-A3B recovers its published disjoint structure, confirming the method detects modularity when present. The verdict reproduces on the un-quantized BF16 model, ruling out a 4-bit quantization artifact. We conclude that ablation-based modularity verdicts are not safe unless the corpus, metric, and statistical bar are controlled. We release the atlas and ablation data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。