提出新模型MixER,提升多层级动态系统重建的效率与泛化能力。
Towards Foundational Models for Dynamical System Reconstruction: Hierarchical Meta-Learning via Mixture of Experts
- 采用基于K-means和最小二乘的自定义门控机制,优化专家路由速度与一致性。
- 在最多含10个参数的常微分方程系统中实现高效训练与可扩展性。
- 特别适合稀疏、层次结构明显的科学数据,如神经科学时间序列分析。
随着基础模型重塑科学发现,动态系统重建(DSR)仍面临跨系统层级学习能力的瓶颈。尽管元学习在单一系统上表现良好,但在数据稀疏且关联松散、需学习多个层级的情况下效果下降。混合专家(MoE)为解决此问题提供了自然范式,但传统MoE因依赖梯度下降更新门控机制,导致更新缓慢且路由冲突。为此,本文提出MixER:一种采用稀疏top-1 MoE结构的新型重构器,其门控更新算法基于K-means与最小二乘法。大量实验验证了MixER的有效性,在最多包含10个参数的常微分方程系统中实现高效训练与可扩展性。然而,在高数据量场景下,当每个专家仅处理高度相关的数据子集时,其性能仍不及现有最优元学习方法。进一步分析表明,混合专家生成的上下文表示质量与数据中是否存在层次结构密切相关。
原文摘要 · Abstract (English)
As foundational models reshape scientific discovery, a bottleneck persists in dynamical system reconstruction (DSR): the ability to learn across system hierarchies. Many meta-learning approaches have been applied successfully to single systems, but falter when confronted with sparse, loosely related datasets requiring multiple hierarchies to be learned. Mixture of Experts (MoE) offers a natural paradigm to address these challenges. Despite their potential, we demonstrate that naive MoEs are inadequate for the nuanced demands of hierarchical DSR, largely due to their gradient descent-based gating update mechanism which leads to slow updates and conflicted routing during training. To overcome this limitation, we introduce MixER: Mixture of Expert Reconstructors, a novel sparse top-1 MoE layer employing a custom gating update algorithm based on $K$-means and least squares. Extensive experiments validate MixER's capabilities, demonstrating efficient training and scalability to systems of up to ten parametric ordinary differential equations. However, our layer underperforms state-of-the-art meta-learners in high-data regimes, particularly when each expert is constrained to process only a fraction of a dataset composed of highly related data points. Further analysis with synthetic and neuroscientific time series suggests that the quality of the contextual representations generated by MixER is closely linked to the presence of hierarchical structure in the data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。