arXiv:2507.20424cs.LGcs.DC2025-07中稿 · UAI 2026 - 9 pages…被引 1

通过推拉机制提升分布式训练的泛化能力,通信效率不降反而更好。

Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning

  • 用逆均值谷深度量平坦极小值,轻量正则化引导协同寻找宽极小。
  • 在多个数据集上相比其他方法提升泛化性能,通信开销保持低位。
  • 适合追求高泛化且资源受限的分布式深度学习场景。

我们研究中心化分布式数据并行训练深度神经网络(DNNs),旨在改进局部梯度方法在通信效率与模型性能之间的权衡。为此,我们重新审视平坦极小值假说,该假说认为具有良好泛化能力的模型通常位于损失曲面更平坦的区域。我们提出一种简单而有效的平坦度度量——逆均值谷深(Inverse Mean Valley),并证实其与DNN泛化差距具有强相关性。我们将该度量的高效松弛形式引入分布式训练目标作为轻量正则项,鼓励各工作节点协同搜索宽极小值。该正则项产生向外推力,对抗同步步骤中的拉力,形成分布式推-拉力(DPPF)算法。实验表明,DPPF优于其他通信高效的训练方法,在保持通信效率的同时,其泛化性能超越局部梯度法和同步梯度平均。此外,损失曲面可视化验证了DPPF定位平坦极小值的能力。理论上,我们证明了DPPF能引导工作节点覆盖平坦山谷,最终谷宽由推力与拉力的平衡决定,且其推拉动态具备自稳定性。我们还建立了与谷宽相关的泛化保证,并证明了非凸设置下的收敛性。

原文摘要 · Abstract (English)

We study centralized distributed data parallel training of deep neural networks (DNNs), aiming to improve the trade-off between communication efficiency and model performance of the local gradient methods. To this end, we revisit the flat-minima hypothesis, which suggests that models with better generalization tend to lie in flatter regions of the loss landscape. We introduce a simple, yet effective, sharpness measure, Inverse Mean Valley, and demonstrate its strong correlation with the generalization gap of DNNs. We incorporate an efficient relaxation of this measure into the distributed training objective as a lightweight regularizer that encourages workers to collaboratively seek wide minima. The regularizer exerts a pushing force that counteracts the consensus step pulling the workers together, giving rise to the Distributed Pull-Push Force (DPPF) algorithm. Empirically, we show that DPPF outperforms other communication-efficient approaches and achieves better generalization performance than local gradient methods and synchronous gradient averaging, while maintaining communication efficiency. In addition, our loss landscape visualizations confirm the ability of DPPF to locate flatter minima. On the theoretical side, we show that DPPF guides workers to span flat valleys, with the final valley width governed by the interplay between push and pull strengths, and that its pull-push dynamics is self-stabilizing. We further provide generalization guarantees linked to the valley width and prove convergence in the non-convex setting.

分布式训练泛化能力极小值搜索通信效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。