arXiv:2608.06046cs.NIcs.DC2026-08

让网络和训练协同优化,提速42%完成模型训练。

ML-for-ML

论文配图:ML-for-ML
图 1 · 摘自论文原文
  • 联合调整网络与训练参数,统一优化目标
  • 在共享云环境中提升至目标损失速度达42%更快
  • 适合关注训练效率的机器学习系统研发者

AI训练负载快速增长,其耗时、能耗与基础设施成本日益突出。在共享云集群中,训练与微调任务与其它工作负载竞争网络资源,而网络机制与机器学习训练策略通常被独立优化:网络控制字节传输,机器学习系统决定通信时机与量级。我们指出这种分离导致端到端性能未被充分挖掘。本文提出ML-for-ML,一种跨层视角,将网络侧与机器学习侧的调控参数在统一的时间-目标损失目标下共同选择。初步原型表明,通过联合优化机器学习与网络参数,可在达到目标损失上比传统方法快最多42%。

原文摘要 · Abstract (English)

AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running workloads for network resources, while network mechanisms and ML training choices are typically optimized separately: networking controls how bytes move, whereas ML systems control when and how much communication occurs. We argue that this separation leaves end-to-end performance on the table. We present ML-for-ML, a cross-layer perspective in which network-side and ML-side knobs are selected jointly under a shared time-to-target-loss objective. Our preliminary prototype shows that by co-optimizing the ML and network parameters, we reach the target loss up to 42% faster.

机器学习优化网络协同训练加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。