arXiv:2503.06916cs.LG2025-03ICCV被引 7

自监督蒸馏让联邦学习逼近中心化性能

You Are Your Own Best Teacher: Achieving Centralized-level Performance in Federated Learning under Heterogeneous and Long-tailed Data

  • 用自增强样本间知识蒸馏提升本地表征能力
  • 在长尾分布下性能接近中心化训练,超参方法5.4%以上
  • 无需额外数据或模型,适合真实异构联邦场景

数据异质性(局部非独立同分布与全局长尾分布)是联邦学习中的主要挑战,导致性能显著落后于中心化学习。已有研究指出表征差和分类器偏差是主因,并提出受神经坍缩启发的合成单纯形ETF方法以逼近最优表征。然而我们发现,此类方法仍无法有效达到神经坍缩状态,与中心化训练存在巨大差距。本文从自蒸馏视角重新思考该问题,提出FedYoYo(你就是自己最好的老师),引入增强型自蒸馏(ASD)机制,通过弱增强与强增强样本间的知识迁移改进本地表征学习,无需额外数据或模型。进一步提出分布感知对数调整(DLA)以平衡自蒸馏过程并校正偏差表征。实验表明,FedYoYo几乎消除性能差距,在混合异质性下实现中心化级性能,提升本地表征质量,减少模型漂移,加速收敛,特征原型更接近神经坍缩最优。大量实验验证其达到当前最优,甚至在全局长尾设置下超越中心化对数调整方法5.4%。

原文摘要 · Abstract (English)

Data heterogeneity, stemming from local non-IID data and global long-tailed distributions, is a major challenge in federated learning (FL), leading to significant performance gaps compared to centralized learning. Previous research found that poor representations and biased classifiers are the main problems and proposed neural-collapse-inspired synthetic simplex ETF to help representations be closer to neural collapse optima. However, we find that the neural-collapse-inspired methods are not strong enough to reach neural collapse and still have huge gaps to centralized training. In this paper, we rethink this issue from a self-bootstrap perspective and propose FedYoYo (You Are Your Own Best Teacher), introducing Augmented Self-bootstrap Distillation (ASD) to improve representation learning by distilling knowledge between weakly and strongly augmented local samples, without needing extra datasets or models. We further introduce Distribution-aware Logit Adjustment (DLA) to balance the self-bootstrap process and correct biased feature representations. FedYoYo nearly eliminates the performance gap, achieving centralized-level performance even under mixed heterogeneity. It enhances local representation learning, reducing model drift and improving convergence, with feature prototypes closer to neural collapse optimality. Extensive experiments show FedYoYo achieves state-of-the-art results, even surpassing centralized logit adjustment methods by 5.4\% under global long-tailed settings.

联邦学习自蒸馏长尾分布表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。