arXiv:2505.09733cs.LGcs.AI2025-05被引 2

解决联邦学习中的噪声与缺失数据问题,提升模型鲁棒性。

Robust Federated Learning with Confidence-Weighted Filtering and GAN-Based Completion under Noisy and Incomplete Data

  • 用置信度加权过滤去除噪声标签,提升数据质量。
  • 通过条件GAN生成合成数据,补全缺失类别,缓解不平衡。
  • 适合资源受限设备,兼顾隐私保护与实际部署需求。

联邦学习(FL)在保持数据隐私的前提下,为分布式客户端数据协同训练提供了有效方案。然而,标签噪声、类别缺失和分布不均等数据质量问题严重制约其效果。本文提出一种系统性方法,通过自适应噪声清洗、基于条件GAN的合成数据生成及鲁棒联邦训练,全面提升数据完整性。在MNIST与Fashion-MNIST基准数据集上的实验表明,在不同噪声水平和类别不平衡条件下,模型性能显著提升,尤其在宏平均F1分数上表现突出。该框架在计算开销与性能增益间取得良好平衡,适用于资源受限的边缘设备,同时严格保障数据隐私。结果表明,该方法能有效应对常见数据质量问题,具备鲁棒性、可扩展性与隐私合规性,适用于多样化的现实联邦学习场景。

原文摘要 · Abstract (English)

Federated learning (FL) presents an effective solution for collaborative model training while maintaining data privacy across decentralized client datasets. However, data quality issues such as noisy labels, missing classes, and imbalanced distributions significantly challenge its effectiveness. This study proposes a federated learning methodology that systematically addresses data quality issues, including noise, class imbalance, and missing labels. The proposed approach systematically enhances data integrity through adaptive noise cleaning, collaborative conditional GAN-based synthetic data generation, and robust federated model training. Experimental evaluations conducted on benchmark datasets (MNIST and Fashion-MNIST) demonstrate significant improvements in federated model performance, particularly macro-F1 Score, under varying noise and class imbalance conditions. Additionally, the proposed framework carefully balances computational feasibility and substantial performance gains, ensuring practicality for resource constrained edge devices while rigorously maintaining data privacy. Our results indicate that this method effectively mitigates common data quality challenges, providing a robust, scalable, and privacy compliant solution suitable for diverse real-world federated learning scenarios.

联邦学习数据清洗GAN生成隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。