arXiv:2507.05508cs.LG2025-07ICML

用多层蒙特卡洛法让有偏压缩无偏化,兼顾通信效率与理论保障

Beyond Communication Overhead: A Multilevel Monte Carlo Approach for Mitigating Compression Bias in Distributed Learning

  • 设计多层蒙特卡洛压缩框架,用有偏压缩生成无偏梯度估计
  • 在分布式深度学习任务中验证,性能优于传统有偏压缩方法
  • 适配主流压缩器如Top-k和位压缩,适合大规模分布式训练场景

分布式学习近年来发展迅速,通信开销常成为主要瓶颈。梯度压缩技术可降低通信成本,但存在有偏压缩器的实证效率与无偏压缩器的理论保证之间的权衡。本文提出一种新的多层蒙特卡洛(MLMC)压缩方案,利用有偏压缩器构造统计无偏估计,有效弥合有偏与无偏方法的差距,融合二者优势。为展示方法的通用性,我们将该方案应用于主流压缩器(如Top-k和位压缩),得到性能增强的新变体。此外,我们推导出自适应版本以进一步提升效果。实验在分布式深度学习任务上验证了该方法的有效性。

原文摘要 · Abstract (English)

Distributed learning methods have gained substantial momentum in recent years, with communication overhead often emerging as a critical bottleneck. Gradient compression techniques alleviate communication costs but involve an inherent trade-off between the empirical efficiency of biased compressors and the theoretical guarantees of unbiased compressors. In this work, we introduce a novel Multilevel Monte Carlo (MLMC) compression scheme that leverages biased compressors to construct statistically unbiased estimates. This approach effectively bridges the gap between biased and unbiased methods, combining the strengths of both. To showcase the versatility of our method, we apply it to popular compressors, like Top-$k$ and bit-wise compressors, resulting in enhanced variants. Furthermore, we derive an adaptive version of our approach to further improve its performance. We validate our method empirically on distributed deep learning tasks.

分布式学习梯度压缩蒙特卡洛无偏估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。