arXiv:2601.18950stat.MLcs.IT2026-01被引 1

提出四种无需先验相关性的分布式均值压缩方法,显著降低通信开销。

Collaborative Compressors in Distributed Mean Estimation with Limited Communication Budget

  • 通过协作压缩利用向量间相似性,无需预先知道相关性
  • 理论证明误差随相似度下降呈平滑退化,覆盖ℓ₂、ℓ∞和余弦误差
  • 适用于通信受限的分布式优化场景,实现高效低开销计算

分布式高维均值估计是分布式优化中常见的聚合操作。在通信受限场景下,需对向量进行压缩后再共享。传统独立编码解码忽略向量间的相似性。现有相关性感知压缩方案虽能提升效率,但依赖已知相关性,且对相似度下降时的性能退化分析仅限于ℓ₂误差。本文提出四种无需先验相关性的协作压缩方案,均简单易实现且计算高效,可大幅减少通信量。理论分析揭示了ℓ₂、ℓ∞及余弦估计误差随向量相似度变化的规律,为实际应用提供指导。

原文摘要 · Abstract (English)

Distributed high dimensional mean estimation is a common aggregation routine used often in distributed optimization methods. Most of these applications call for a communication-constrained setting where vectors, whose mean is to be estimated, have to be compressed before sharing. One could independently encode and decode these to achieve compression, but that overlooks the fact that these vectors are often close to each other. To exploit these similarities, recently Suresh et al., 2022, Jhunjhunwala et al., 2021, Jiang et al, 2023, proposed multiple correlation-aware compression schemes. However, in most cases, the correlations have to be known for these schemes to work. Moreover, a theoretical analysis of graceful degradation of these correlation-aware compression schemes with increasing dissimilarity is limited to only the $\ell_2$-error in the literature. In this paper, we propose four different collaborative compression schemes that agnostically exploit the similarities among vectors in a distributed setting. Our schemes are all simple to implement and computationally efficient, while resulting in big savings in communication. The analysis of our proposed schemes show how the $\ell_2$, $\ell_\infty$ and cosine estimation error varies with the degree of similarity among vectors.

分布式优化压缩算法通信效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。