arXiv:2608.11212cs.AIcs.CL2026-08

量化导致专家路由错误,但现有方法无法判断错误是否有害。

Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts

论文配图:Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts
图 1 · 摘自论文原文
  • 通过四次实验框架量化路由错误造成的损失比例。
  • 4位量化下约31%的性能损失由路由翻转引起。
  • 检测路由翻转易,判断其好坏难,限制了修复策略。

Top-k混合专家(MoE)路由具有不连续性,部署导向的数值扰动——由保护的BF16门控读取的4比特键值缓存量化——会将标记推过决策边界,引发专家激活翻转。本文未提出新缓解方法,而是提供因果分析工具、实证发现与检测极限结果。四次运行的装置量化了量化损伤中路由介导部分(RMF),token级归因分解其机制,预注册探针在三个架构间传递发现。在OLMoE-1B-7B 4比特KV(试点)下,约三分之一的损伤为路由介导:RMF ~ 0.31(发现值0.31 [0.20, 0.41];过程复现均值0.313 ± 0.020;预注册重执行0.231)。可部署的路由器边际能检测翻转发生(AUC 0.772),但无法区分有害或有益翻转(随机水平):在测试的局部、推理可观测路由器统计中,无预测翻转损失符号的指标显著优于随机——构成对选择性修复的实证障碍。有符号翻转税与符号不可分性具跨模型普适性;干净参考修复的收益受架构调节;受控同检查点标志交换重构门控归一化惯例,使其成为损伤幅度调节器而非路由恢复机制。真实int4 KV核产生的损伤比例与虚构量化剂量曲线一致,但功效不足(95% CI [-0.111, 0.394] 包含零)——排除严重分歧,非独立复现。假设、阈值与评估均在测量前预注册,漏失情况已报告;预注册保留读取复现了分区及样本外近似抵消的税,严格不可能排除略显不足。

原文摘要 · Abstract (English)

Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quantization read by a protected BF16 gate -- pushes tokens across decision boundaries and flips which experts fire. This paper proposes no new mitigation; it supplies a causal apparatus, empirical findings, and a detection-limit result. A four-run apparatus prices the route-mediated fraction (RMF) of quantization damage, a token-level attribution decomposes it by mechanism, and pre-registered probes carry the findings across three architectures. On OLMoE-1B-7B at 4-bit KV (pilot), about a third of the damage is routing-mediated: RMF ~ 0.31 (discovery 0.31 [0.20, 0.41]; process-replicated mean 0.313 +/- 0.020; pre-registered re-execution 0.231). The deployable router margin detects that a flip occurred (AUC 0.772) but cannot tell a harmful flip from a helpful one (at chance): among the tested local, inference-observable router statistics we find no predictor of a flip's loss sign above chance -- an empirical benefit-detection barrier bounding selective repair restricted to this feature family. The signed-flip tax and sign-inseparability carry cross-model; the clean-reference remedy's payout is architecture-modulated; a controlled same-checkpoint flag-swap re-scopes the gate's normalization convention to a damage-magnitude moderator, not a route-recoverability mechanism. A real int4 KV kernel yields a fraction compatible with the fake-quant dose curve but underpowered (95% CI [-0.111, 0.394] includes zero) -- ruling out gross disagreement, not an independent replication. Hypotheses, thresholds, and evaluations were pre-registered before measurement, with misses reported; a pre-registered held-out read replicates the partition and the near-cancelling tax out of sample, while the strict impossibility exclusion narrowly misses.

量化路由专家模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。