arXiv:2507.05685cs.LGcs.AI2025-07被引 3

解决联邦MoE训练中客户端与专家动态匹配难题,提升通信效率。

Efficient Training of Large-Scale AI Models Through Federated Mixture-of-Experts: A System-Level Approach

  • 设计智能动态匹配机制,根据客户端能力实时分配专家
  • 实现全局专家负载监控,减少通信轮次达成收敛
  • 适合边缘计算场景下大规模隐私保护模型训练

联邦学习(FL)与专家混合模型(MoE)的结合为在去中心化数据上训练更强大的大规模人工智能模型(LAMs)提供了可能,同时保障隐私。然而,高效联邦训练这类复杂结构的LAMs面临显著系统级挑战,尤其在于协调异构客户端资源与众多专用专家之间的复杂关系。本文指出一个关键但未被充分研究的问题:缺乏兼顾客户端能力差异与整体系统负载均衡的动态客户端-专家对齐量化策略。为此,提出一种系统级概念设计,包含动态适应度评分、全局专家负载监测和客户端能力画像。通过解决这些系统性问题,可实现更可扩展、高效且鲁棒的训练机制,以更少通信轮次完成收敛,推动大规模联邦MoE结构化LAMs在边缘计算中的广泛应用,具备超高通信效率。

原文摘要 · Abstract (English)

The integration of Federated Learning (FL) and Mixture-of-Experts (MoE) presents a compelling pathway for training more powerful, large-scale artificial intelligence models (LAMs) on decentralized data while preserving privacy. However, efficient federated training of these complex MoE-structured LAMs is hindered by significant system-level challenges, particularly in managing the interplay between heterogeneous client resources and the sophisticated coordination required for numerous specialized experts. This article highlights a critical, yet underexplored concept: the absence of robust quantitative strategies for dynamic client-expert alignment that holistically considers varying client capacities and the imperative for system-wise load balancing. Specifically, we propose a conceptual system design for intelligent client-expert alignment that incorporates dynamic fitness scoring, global expert load monitoring, and client capacity profiling. By tackling these systemic issues, we can unlock more scalable, efficient, and robust training mechanisms {with fewer communication rounds for convergence}, paving the way for the widespread deployment of large-scale federated MoE-structured LAMs in edge computing with ultra-high communication efficiency.

联邦学习MoE边缘计算通信效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。