arXiv:2410.20659stat.MLcs.AI2024-10被引 2

揭示联邦学习中数据内在维度对模型收敛速度的关键影响

A Statistical Analysis of Deep Federated Learning for Intrinsically Low-dimensional Data

  • 基于双阶段采样模型分析深度联邦回归的泛化性能
  • 误差率与数据内在维数相关,而非名义高维特征
  • 适用于研究异构联邦学习泛化能力的科研人员

尽管联邦学习的优化问题已受到广泛关注,但其泛化误差,尤其是在异构联邦学习中的表现,仍缺乏深入研究,主要局限于参数化范式。本文在两阶段采样模型下探究深度联邦回归的泛化性质。结果表明,由熵维数表征的内在维度在选择合适网络规模时,对深度学习者的收敛速率起决定性作用。具体而言,当响应变量与解释变量间的真实关系为β-霍尔德函数,且从m个参与客户端获取n个独立同分布样本时,参与客户端的误差率最多为$ ilde{O}((mn)^{-2β/(2β+ ar{d}_{2β}(λ))})$,非参与客户端则为$ ilde{O}(Δullet m^{-2β/(2β+ ar{d}_{2β}(λ))} + (mn)^{-2β/(2β+ ar{d}_{2β}(λ))})$。其中$ar{d}_{2β}(λ)$表示解释变量边缘分布λ的2β-熵维数,Δ刻画采样方案两阶段间的依赖关系。因此,研究不仅显式纳入客户端异质性,还表明深度联邦学习的误差收敛速率取决于数据的内在维度,而非名义上的高维特征。

原文摘要 · Abstract (English)

Despite significant research on the optimization aspects of federated learning, the exploration of generalization error, especially in the realm of heterogeneous federated learning, remains an area that has been insufficiently investigated, primarily limited to developments in the parametric regime. This paper delves into the generalization properties of deep federated regression within a two-stage sampling model. Our findings reveal that the intrinsic dimension, characterized by the entropic dimension, plays a pivotal role in determining the convergence rates for deep learners when appropriately chosen network sizes are employed. Specifically, when the true relationship between the response and explanatory variables is described by a $β$-Hölder function and one has access to $n$ independent and identically distributed (i.i.d.) samples from $m$ participating clients, for participating clients, the error rate scales at most as $\Tilde{O}((mn)^{-2β/(2β+ \bar{d}_{2β}(λ))})$, whereas for non-participating clients, it scales as $\Tilde{O}(Δ\cdot m^{-2β/(2β+ \bar{d}_{2β}(λ))} + (mn)^{-2β/(2β+ \bar{d}_{2β}(λ))})$. Here $\bar{d}_{2β}(λ)$ denotes the corresponding $2β$-entropic dimension of $λ$, the marginal distribution of the explanatory variables. The dependence between the two stages of the sampling scheme is characterized by $Δ$. Consequently, our findings not only explicitly incorporate the ``heterogeneity" of the clients, but also highlight that the convergence rates of errors of deep federated learners are not contingent on the nominal high dimensionality of the data but rather on its intrinsic dimension.

联邦学习泛化误差内在维度异构数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。