发现MoE模型路由层敏感但难控,无法有效实现公平性干预。
Routing Sensitivity Without Controllability: A Diagnostic Study of Fairness in MoE Language Models
- 提出FARE诊断框架,评估不同MoE架构中路由级偏见干预的可行性。
- 多数模型中路由偏好调整要么不可行,要么损失性能或效果不显著。
- 偏见与核心知识深度耦合,导致路由干预无法传递到生成结果。
混合专家(MoE)语言模型在路由层级普遍对人口统计内容敏感,但利用此敏感性进行公平性控制在结构上受限。本文提出公平感知路由均衡(FARE)诊断框架,用于探究不同MoE架构下路由级刻板印象干预的边界。FARE揭示:路由偏好调整在Mixtral、Qwen1.5、Qwen3中不可实现,在DeepSeekMoE中统计不稳健,而在OLMoE中则伴随显著性能代价(CrowS-Pairs下降4.4%p,TQA下降6.3%p)。关键的是,即使对数似然偏好调整具有鲁棒性,也无法传递至生成结果:在非空模型上扩展评估均未在任何生成指标上取得有效提升。群体级专家屏蔽分析表明:偏见与核心知识深度耦合于专家组内。研究说明路由敏感性虽为必要前提,却不足以实现刻板印象控制,并明确了未来可控MoE系统的设计条件。
原文摘要 · Abstract (English)
Mixture-of-Experts (MoE) language models are universally sensitive to demographic content at the routing level, yet exploiting this sensitivity for fairness control is structurally limited. We introduce Fairness-Aware Routing Equilibrium (FARE), a diagnostic framework designed to probe the limits of routing-level stereotype intervention across diverse MoE architectures. FARE reveals that routing-level preference shifts are either unachievable (Mixtral, Qwen1.5, Qwen3), statistically non-robust (DeepSeekMoE), or accompanied by substantial utility cost (OLMoE, -4.4%p CrowS-Pairs at -6.3%p TQA). Critically, even where log-likelihood preference shifts are robust, they do not transfer to decoded generation: expanded evaluations on both non-null models yield null results across all generation metrics. Group-level expert masking reveals why: bias and core knowledge are deeply entangled within expert groups. These findings indicate that routing sensitivity is necessary but insufficient for stereotype control, and identify specific architectural conditions that can inform the design of more controllable future MoE systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。