arXiv:2511.11778cs.LG2025-11被引 1

在标签极少时,利用客户端未标记数据提升联邦学习效果。

CATCHFed: Efficient Unlabeled Data Utilization for Semi-Supervised Federated Learning in Limited Labels Environments

  • 根据类别难易度自适应调整阈值,优化伪标签生成。
  • 通过混合阈值和一致性正则化,显著提升小样本下性能。
  • 适合标签稀缺的隐私保护场景,如医疗、金融联邦建模。

联邦学习是一种利用分布式客户端资源并保护数据隐私的有前景范式。现有大多数联邦学习方法假设客户端拥有标签数据,但在真实场景中,客户端标签常不可用。半监督联邦学习允许仅服务器持有标签数据,以应对这一问题。然而,当标签数据减少时,性能会显著下降。为此,我们提出CATCHFed,引入考虑类别难度的客户端感知自适应阈值,采用混合阈值提升伪标签质量,并利用未伪标签数据进行一致性正则化。在多种数据集和配置下的大量实验表明,CATCHFed能有效利用客户端未标记数据,在极端低标签设置下仍取得优异性能。

原文摘要 · Abstract (English)

Federated learning is a promising paradigm that utilizes distributed client resources while preserving data privacy. Most existing FL approaches assume clients possess labeled data, however, in real-world scenarios, client-side labels are often unavailable. Semi-supervised Federated learning, where only the server holds labeled data, addresses this issue. However, it experiences significant performance degradation as the number of labeled data decreases. To tackle this problem, we propose \textit{CATCHFed}, which introduces client-aware adaptive thresholds considering class difficulty, hybrid thresholds to enhance pseudo-label quality, and utilizes unpseudo-labeled data for consistency regularization. Extensive experiments across various datasets and configurations demonstrate that CATCHFed effectively leverages unlabeled client data, achieving superior performance even in extremely limited-label settings.

联邦学习半监督数据效率隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。