arXiv:2506.02787cs.DCcs.AI2025-06被引 2

解决大模型训练中异构节点与动态网络的并行难题

Rethinking Dynamic Networks and Heterogeneous Computing with Automatic Parallelization

  • 建模异构节点与动态网络,实现细粒度负载分配
  • 在稳定与复杂场景下均达到先进水平性能
  • 通过策略剪枝加速搜索,适合云环境等动态场景

混合并行技术对高效训练大语言模型至关重要。然而,现有自动并行规划框架常忽视节点异构性与动态网络拓扑变化的同步考虑,限制了实际应用效果。本文通过建模动态网络环境下异构节点,采用基于仿真的策略确定最优并行配置,实现面向异构节点与复杂网络场景的细粒度工作负载分配,在常规稳定网络条件下性能媲美当前最先进方法。此外,引入策略剪枝技术,快速剔除不可行配置,显著缩小搜索空间,通过模拟器内并行执行大幅加速搜索过程。初步评估表明,该方法在异构节点上显著提升训练性能,并在云计算等复杂动态场景中展现更强适应性。

原文摘要 · Abstract (English)

Hybrid parallelism techniques are essential for efficiently training large language models (LLMs). Nevertheless, current automatic parallel planning frameworks often overlook the simultaneous consideration of node heterogeneity and dynamic network topology changes, limiting their effectiveness in practical applications. In this paper, we address these limitations by modeling heterogeneous nodes within dynamically changing network environments and leveraging simulation-based strategies to determine optimal parallel configurations. Our approach enables fine-grained workload allocation tailored for heterogeneous nodes and complex network scenarios, achieving performance competitive with state-of-the-art methods under regular and stable network conditions. Additionally, we introduce a strategy pruning technique to rapidly discard infeasible parallel configurations, substantially reducing the search space and accelerating the search process through parallel execution within the simulator. Preliminary evaluations confirm that our method notably enhances training performance on heterogeneous nodes and demonstrates improved adaptability in complex, dynamic scenarios such as cloud computing environments.

大模型训练异构计算动态网络并行优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。