arXiv:2608.21840cs.LG2026-08

专家越多越差劲:混合专家状态空间模型在动态系统预测中表现反而下降。

More Experts, Worse Dynamics: Inverse Scaling and Spectral Bias in Mixture-of-Experts State-Space Models

论文配图:More Experts, Worse Dynamics: Inverse Scaling and Spectral Bias in Mixture-of-Experts State-Space Models
图 1 · 摘自论文原文
  • 在三种动态模式下测试专家混合模型,发现增加专家数量反而降低性能。
  • 专家越多,路由失效,模型无法区分不同动态模式,整体误差上升。
  • 现有评估指标可能误导,需结合相空间分析才能真实判断模型表现。

混合专家(MoE)架构通常被视为通过分解复杂系统为更简单的局部动态来提升表达能力。这一直觉被扩展到谱状态空间模型中,即混合稳定算子可适应异质或模式切换的时间序列。我们在此研究一个受控合成设置,以分离动力学而非表征挑战。任务为对包含三类模式的序列进行下一步预测:由Mackey-Glass系统生成的混沌动态、稳定振荡模式和噪声主导的自回归模式。通过大量消融实验(包括容量缩放、最优路由、冻结专家变体及输出级MoE基线对比),发现算子级混合模型始终未能超越单专家基线。专家数量增加导致逆向缩放现象,路由崩溃或无法引发有意义的专业化,即使具备完美模式监督也无法避免全局性能下降。此外,混沌轨迹上均方误差的看似改善具有误导性。相空间分析表明,较低误差常源于时间平滑,破坏了底层吸引子的几何结构,而非真实建模动态。这些结果揭示了在当前参数化与训练协议下算子插值的潜在局限,并强调在评估模式切换动力系统时需采用几何感知的评价方式。

原文摘要 · Abstract (English)

Mixture-of-Experts (MoE) architectures are commonly motivated as a way to increase expressivity by decomposing complex systems into simpler local dynamics. This intuition has recently been extended to spectral state-space models, where mixing stable operators is assumed to enable adaptation to heterogeneous or regime-switching time series. We critically evaluate this assumption in a controlled synthetic setting designed to isolate dynamical rather than representational challenges. We study a next-step prediction task on sequences composed of three regimes: chaotic dynamics generated by the Mackey-Glass system, a stable oscillatory regime, and a noise-dominated autoregressive regime. Across extensive ablations including capacity scaling, oracle routing, frozen-expert variants, and comparisons to output-level MoE baselines, operator-level mixture models consistently fail to outperform a single-expert baseline. Increasing the number of experts leads to inverse scaling, routing collapses or fails to induce meaningful specialization, and even perfect regime supervision does not prevent degradation in global performance. Furthermore, we show that apparent improvements in mean squared error on chaotic trajectories can be misleading. Phase-space analysis reveals that lower error often arises from temporal smoothing that destroys the geometry of the underlying attractor rather than from faithful modeling of the dynamics. These results identify a likely limitation of operator interpolation under the studied parameterization and training protocol, and underscore the need for geometry-aware evaluation when assessing regime-switching dynamical systems.

动态建模混合专家状态空间模型逆向缩放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。