arXiv:2507.03221cs.LGcs.AI2025-07

用神经抑制提升专家模型动态路由效率

Neural Inhibition Improves Dynamic Routing and Mixture of Experts

  • 引入神经抑制机制,让模型根据数据自动选择专用专家路径
  • 实验显示该方法显著提升通用任务性能,优于基线模型
  • 适合研究动态路由、Mixture of Experts及Transformer的学者

为实现高效且多样化的深度学习模型,需根据神经元群体信号动态选择架构。我们提出,通过在神经元群体中引入神经抑制,可抑制各类数据统计中共享的信号,从而让路由模型为每个数据样本选择更专精的专家路径。只有通过抑制,路由机制才能有效筛选神经通路。这一方法在混合专家模型、动态路由及Transformer语言模型中尚未被充分研究与验证。我们提供了实验证据,证明神经抑制算法能显著提升通用任务表现,激励更多研究投入该方向。

原文摘要 · Abstract (English)

To be effective, efficient, and diverse, deep learning models need to dynamically choose its architecture based on signals from a population of neurons. We hypothesize dynamic routing models can be improved with neural inhibition in those neural populations. This means signals commonly shared among the various modes of data statistics can be inhibited so that the routing model can choose a specialized expert path for each data sample. Only through inhibition is the routing mechanism able to effectively select neural pathways. We believe this is an under-studied and under-verified implementation methodology for Mixture-of-Experts, dynamic routing, and transformer language models. We provide experimental evidence that the neural inhibition algorithm significantly boosts the performance of general tasks and motivates more effort to be invested in this research direction.

动态路由专家模型神经抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。