arXiv:2506.00440cs.LG2025-06被引 1

用群体稳定性指数选更同质客户端,提升非独立同分布下的联邦学习性能。

PSI-PFL: Population Stability Index for Client Selection in non-IID Personalized Federated Learning

  • 基于群体稳定性指数(PSI)筛选数据分布相近的客户端
  • 在多种数据类型上实现最高10%的准确率提升,缓解标签偏移影响
  • 适合数据隐私强且数据异质性高的实际应用场景

联邦学习(FL)可在保护数据隐私的前提下实现分布式机器学习模型训练,但客户端间非独立同分布(non-IID)数据会引发模型更新偏差与性能下降。为此,本文提出针对个性化联邦学习(PFL)的客户端选择框架PSI-PFL,利用群体稳定性指数(PSI)量化并缓解数据异质性(即非独立同分布程度)。该方法根据PSI值筛选更具同质性的客户端,有效降低标签偏移带来的负面影响。在表格、图像、文本等多种数据模态上的实验表明,PSI-PFL在非IID场景下显著提升全局模型准确率,相较现有最优基线最高提升达10%,同时保障本地模型表现更均衡。该方法不仅增强了联邦学习性能,还为数据隐私与异质性共存的实际应用提供了实用价值。

原文摘要 · Abstract (English)

Federated Learning (FL) enables decentralized machine learning (ML) model training while preserving data privacy by keeping data localized across clients. However, non-independent and identically distributed (non-IID) data across clients poses a significant challenge, leading to skewed model updates and performance degradation. Addressing this, we propose PSI-PFL, a novel client selection framework for Personalized Federated Learning (PFL) that leverages the Population Stability Index (PSI) to quantify and mitigate data heterogeneity (so-called non-IIDness). Our approach selects more homogeneous clients based on PSI, reducing the impact of label skew, one of the most detrimental factors in FL performance. Experimental results over multiple data modalities (tabular, image, text) demonstrate that PSI-PFL significantly improves global model accuracy, outperforming state-of-the-art baselines by up to 10\% under non-IID scenarios while ensuring fairer local performance. PSI-PFL enhances FL performance and offers practical benefits in applications where data privacy and heterogeneity are critical.

联邦学习数据异质性客户端选择隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。