发现专家模型剪枝中常用指标无法预测真实重要性,挑战了主流可解释性方法的推理逻辑。
From Observation to Intervention: A Causal Audit of Expert Importance in Mixture-of-Experts Models

- 通过逐标记干预实验检验路由统计量对专家重要性的预测能力
- 60组测试中所有指标效应量均低于0.23,无一可靠预测专家可删
- 揭示早期层冗余导致剪枝方法成功但非因识别出真正不重要专家
可解释性方法常以群体层面的观测统计量推断特定计算干预的效果,在皮尔的层级框架中,这相当于将第1层关联证据当作第2层干预结论使用,其有效性极少被验证。本文聚焦混合专家(MoE)剪枝中的路由统计量:利用率、激活范数与路由权重分布是否能预测哪些专家可移除而无功能损失。在三个高冗余MoE架构(OLMoE-1B-7B-0924, Qwen1.5-MoE-A2.7B, DeepSeek-V2-Lite)上进行逐标记干预审计,发现所有60个指标-层组合的效应量均低于Cohen's d=0.23,且无任何指标在修正后的双重检验标准下稳定为正。仅当使用相同样本量的逐标记路由权重控制实验时,在OLMoE最终层恢复信号(d=+0.231,95% CI [+0.09, +0.37],p=0.0013)。现有剪枝方法的成功并非源于识别可删专家,而是因为早期层冗余使多数选择标准等价。研究提供了一个明确反例,证明从群体观测到个体干预的推论不成立,并展示了干预审计如何校准可解释性声明的证据标准。
原文摘要 · Abstract (English)
Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of targeted interventions on specific computations; in Pearl's terms, they treat rung-1 associational evidence as if it supported rung-2 interventional conclusions, a move whose validity is rarely tested. We examine one concrete instance: the use of routing statistics in Mixture-of-Experts (MoE) pruning, where utilization rates, activation norms, and routing weight distributions are treated as predictors of which experts can be removed without functional cost. A token-level interventional audit across three high-redundancy MoE architectures (OLMoE-1B-7B-0924, Qwen1.5-MoE-A2.7B, DeepSeek-V2-Lite) finds no observational metric predicts causal expert importance in any model: across all 60 metric-layer combinations effect sizes stay below Cohen's $d = 0.23$, and no metric is reliably positive under our corrected, dual-test criterion. A per-token routing weight control, run with identical $n$, rules out insufficient power, recovering a signal whose CI excludes zero at OLMoE's final MoE layer ($d = +0.231$, 95\% CI $[+0.09, +0.37]$, $p = 0.0013$). Existing pruning methods succeed in this regime not by identifying dispensable experts but because early-layer redundancy renders most selection criteria interchangeable. Our results provide an explicit counterexample to the common inferential step from population-level observational summaries to token-level interventional claims about expert importance, and illustrate how interventional audits can calibrate the evidential standards for interpretability claims.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。