arXiv:2603.11307cs.LG2026-03

用本地数据统计条件化全局模型,零通信开销解决数据异构难题。

Client-Conditional Federated Learning via Local Training Data Statistics

  • 基于客户端本地数据的PCA统计量条件化全局模型
  • 在多种异构场景下性能超越最优基准1-6%,稀疏数据下最稳健
  • 无需聚类、不存每客户端模型,适合资源受限场景

数据异构下的联邦学习仍具挑战:现有方法或忽略客户端差异(FedAvg),或需高成本聚类发现(IFCA),或需维护每个客户端模型(Ditto)。这些方法在数据稀疏或多维异构时表现下降。本文提出通过客户端本地计算的主成分分析(PCA)统计量来条件化单个全局模型,无需额外通信开销。在涵盖四种异构类型(标签偏移、协变量偏移、概念偏移及综合异构)、四个数据集(MNIST、Fashion-MNIST、CIFAR-10、CIFAR-100)和七种联邦学习基线方法的97种配置上评估,本方法在所有设置中达到与已知真实聚类分配的“黄金标准”(Oracle)相当的性能,在综合异构场景中更优于其1–6%(因连续统计量比离散聚类标识更具信息量),且是唯一对数据稀疏性鲁棒的方法。

原文摘要 · Abstract (English)

Federated learning (FL) under data heterogeneity remains challenging: existing methods either ignore client differences (FedAvg), require costly cluster discovery (IFCA), or maintain per-client models (Ditto). All degrade when data is sparse or heterogeneity is multi-dimensional. We propose conditioning a single global model on locally-computed PCA statistics of each client's training data, requiring zero additional communication. Evaluating across 97~configurations spanning four heterogeneity types (label shift, covariate shift, concept shift, and combined heterogeneity), four datasets (MNIST, Fashion-MNIST, CIFAR-10, CIFAR-100), and seven FL baseline methods, we find that our method matches the Oracle baseline -- which knows true cluster assignments -- across all settings, surpasses it by 1--6% on combined heterogeneity where continuous statistics are richer than discrete cluster identifiers, and is uniquely sparsity-robust among all tested methods.

联邦学习数据异构无通信开销主成分分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。