动态专家系统通过成本惩罚实现分子记忆,恢复旧领域速度提升10倍。
Cost-Penalized Fitness in FMA-Orchestrated Mixture of Experts: Experimental Evidence for Molecular Memory in Domain Adaptation
- 用成本惩罚和新生专家缓冲期管理专家池,避免替换。
- 回溯旧领域时恢复速度达9-11倍,无需新增或更换专家。
- 适合需要长期知识保留的工业级大模型应用。
我们报告了nanoFMT(一种由自由市场算法(FMA)协调的动态混合专家(MoE)Transformer)在七次受控实验中的结果。该研究针对高级大模型发展中的核心问题:当系统在不断变化的数据分布下满负荷运行时,如何管理专家池?实验表明,结合成本惩罚适应度与线性缓冲期的新生专家机制,可使系统通过多样化积累领域知识,而非简单替换。核心成果为一次往返领域迁移实验,返回先前学习领域时恢复速度提升9-11倍,且无需任何专家新生或替换。这种“分子记忆”效应——休眠专家在领域回归时重新激活——在现有MoE管理方法中尚无先例。初步成本分析显示,在中等情景下,对类似OpenAI规模的提供商,每年可节省3910万美元并减少27.1吉瓦时能耗。
原文摘要 · Abstract (English)
We present experimental results from seven controlled runs of nanoFMT, a Free-Market Algorithm (FMA) orchestrated transformer with dynamic Mixture-of-Experts (MoE) management. The experiments address a fundamental question for advanced LLM development: how should an MoE system manage its expert pool when operating at full capacity under changing data distributions? We demonstrate that cost-penalized fitness metrics, combined with a linear grace period for newborn experts, produce a system that accumulates domain expertise through diversification rather than replacement. The central result is a round-trip domain shift experiment showing 9-11x faster recovery when returning to a previously learned domain, with zero expert births or replacements required. This "molecular memory" effect -- where dormant experts survive and reactivate when their domain returns -- has no analogue in current MoE management approaches. A preliminary cost analysis estimates annual savings of $39.1M and 27.1 GWh energy reduction for an OpenAI-scale provider under a moderate scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。