arXiv:2604.14434cs.AI2026-04被引 2

让专家分工可解释可控,用几何路由实现因果干预。

Geometric Routing Enables Causal Expert Control in Mixture of Experts

  • 用余弦相似度路由,在低维空间显化专家语义专属性。
  • 15%专家为单义专家,覆盖10类语义,干预后概率提升超3倍。
  • 几何透明路由支持推理时零开销控制,适合可解释模型研究。

稀疏混合专家(MoE)模型在保持每标记计算量不变的前提下扩展参数规模,但各专家的专属性仍不明确。我们证明:尽管路由拓扑对模型质量无影响(五种结构配置收敛至统计等效语言建模性能),但专家身份具有因果意义。通过构建余弦相似度路由的低维度量空间,可直接观察到专家的语义专属性。第一,将专家输出向量经未嵌入矩阵投影生成语义词典,15%的专家为单义专家,涵盖时间、地理、基数、话语、情感、金融、军事、科学等10个类别;第二,路由呈现频率到句法的梯度:早期层按词频分离,深层按句法类别分离(均通过Zipf混淆控制,所有p < 0.001);第三,因果干预验证标签有效性:引导至时间专家中心使时间概率提升+321%(44个提示的中位数);抑制地理专家使地理概率下降-23%;重写专家输出使目标类别概率减半,且效应在层间叠加;第四,该干预效果不仅限于余弦路由,线性路由也可实现类似控制,但仅余弦路由提供几何透明性——专家专属性可直接从质心矩阵读取。因此,专家级专属性是首类可解释性原语:架构上单义、因果验证、推理时零开销可控。

原文摘要 · Abstract (English)

Sparse Mixture-of-Experts (MoE) models scale parameters while fixing active computation per token, but the specialization of individual experts remains opaque. In a companion paper we showed that routing topology is quality-neutral: five structurally different configurations converge to statistically equivalent language modeling quality. Here we show that expert identity is nonetheless causally meaningful: individual rank-1 experts are monosemantic by construction, and cosine-similarity routing in a low-dimensional metric space makes their specialization directly inspectable. We present four lines of evidence. First, projecting expert output vectors through the unembedding matrix yields a Semantic Dictionary: 15% of experts are monosemantic specialists spanning 10 categories (temporal, geographic, cardinal, discourse, emotional, financial, military, scientific). Second, routing exhibits a frequency-to-syntax gradient: early layers separate tokens by word frequency, deeper layers by syntactic class (Zipf-confound controls, all $p < 0.001$). Third, causal interventions confirm these labels: steering toward a temporal expert's centroid increases P(temporal) by +321% (median across 44 prompts); suppressing a geographic expert drops P(geographic) by -23%; rewriting an expert's output vector halves target-category probability, and effects compose additively across layers. Fourth, the interventions are not unique to cosine routing: linear routers support comparable steering, but only cosine routing provides geometric transparency -- expert specialization is readable directly from the centroid matrix. MoE expert-level specialization is a first-class interpretability primitive: architecturally monosemantic, causally validated, and controllable at inference with zero overhead.

MoE可解释性因果干预几何路由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。