数据异质性越强,分布式学习泛化性能反而越好。
Heterogeneity Matters even More in Distributed Learning: Study from Generalization Perspective
- 从泛化误差角度分析客户端数据分布差异的影响。
- 发现客户端数据越不同,模型泛化误差越小,且可量化。
- 适用于研究联邦学习中异质数据对模型性能的影响。
本文从泛化误差视角研究分布式学习系统中客户端间数据异质性的影响,即单轮联邦学习的性能表现。具体地,$K$ 个客户端各自拥有 $n$ 个独立生成的训练样本,来自可能不同的数据分布,其本地模型由中心服务器聚合。我们分析了客户端数据分布差异对聚合模型泛化误差的影响。首先,基于分布建立了泛化误差的期望上界和尾部上界,部分扩展了传统的条件互信息(CMI)界,使其适用于任意客户端数 $K \geq 1$ 的分布式场景。接着,结合信息论中的率失真理论,推导出可能更紧的有损版本边界。进一步,将该有损界应用于分布式支持向量机(DSVM)分类问题,得到显式的泛化误差界,明确依赖于数据异质程度。结果表明:随着客户端间数据异质性增加,泛化误差界减小,说明在数据越不一致时,DSVM 泛化能力越强。这一反直觉发现不仅超越了特定模型,还在多个实验中得到验证。
原文摘要 · Abstract (English)
In this paper, we investigate the effect of data heterogeneity across clients on the performance of distributed learning systems, i.e., one-round Federated Learning, as measured by the associated generalization error. Specifically, $K$ clients have each $n$ training samples generated independently according to a possibly different data distribution, and their individually chosen models are aggregated by a central server. We study the effect of the discrepancy between the clients' data distributions on the generalization error of the aggregated model. First, we establish in-expectation and tail upper bounds on the generalization error in terms of the distributions. In part, the bounds extend the popular Conditional Mutual Information (CMI) bound, which was developed for the centralized learning setting, i.e., $K=1$, to the distributed learning setting with an arbitrary number of clients $K \geq 1$. Then, we connect with information-theoretic rate-distortion theory to derive possibly tighter \textit{lossy} versions of these bounds. Next, we apply our lossy bounds to study the effect of data heterogeneity across clients on the generalization error for the distributed classification problem in which each client uses Support Vector Machines (DSVM). In this case, we establish explicit generalization error bounds that depend explicitly on the data heterogeneity degree. It is shown that the bound gets smaller as the degree of data heterogeneity across clients increases, thereby suggesting that DSVM generalizes better when the dissimilarity between the clients' training samples is bigger. This finding, which goes beyond DSVM, is validated experimentally through several experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。