arXiv:2503.12220cs.LGcs.CR2025-03被引 3

针对零售数据异构性,提出自适应隐私分组联邦学习框架

PA-CFL: Privacy-Adaptive Clustered Federated Learning for Transformer-Based Sales Forecasting on Heterogeneous Retail Data

  • 按隐私水平与数据特征分组客户端,构建独立联邦学习子系统
  • 相比本地学习,R²提升5.4%,RMSE降低69%,MAE减少45%
  • 适合需要兼顾隐私保护与高精度销售预测的跨区域零售场景

联邦学习(FL)使零售商在保护隐私的前提下共享模型参数以进行需求预测。然而,受消费者行为差异等因素影响,不同地区数据存在显著异构性,制约了联邦学习效果。为此,我们提出面向异构零售数据的需求预测隐私自适应分组联邦学习(PA-CFL)。该方法结合差分隐私与特征重要性分布,将零售商划分为多个“数据气泡”,每个气泡内形成独立的联邦学习系统,有效隔离数据异构性。每个气泡内采用Transformer模型为各客户端预测本地销售。实验表明,PA-CFL显著优于FedAvg,且在所有参与客户端上均超越本地学习:相比本地学习,其R²提升5.4%,RMSE降低69%,MAE减少45%。该方法通过动态调整噪声水平及每组参与客户端范围,实现自适应优化;通过主动筛选高风险客户端,缓解系统安全威胁。结果证明,PA-CFL能有效提升时间序列预测中异构数据下的联邦学习性能,在零售应用中实现预测精度与隐私保护的平衡。此外,其检测并消除恶意数据的能力进一步增强了系统的鲁棒性与可靠性。

原文摘要 · Abstract (English)

Federated learning (FL) enables retailers to share model parameters for demand forecasting while maintaining privacy. However, heterogeneous data across diverse regions, driven by factors such as varying consumer behavior, poses challenges to the effectiveness of federated learning. To tackle this challenge, we propose Privacy-Adaptive Clustered Federated Learning (PA-CFL) tailored for demand forecasting on heterogeneous retail data. By leveraging differential privacy and feature importance distribution, PA-CFL groups retailers into distinct ``bubbles'', each forming its own federated learning system to effectively isolate data heterogeneity. Within each bubble, Transformer models are designed to predict local sales for each client. Our experiments demonstrate that PA-CFL significantly surpasses FedAvg and outperforms local learning in demand forecasting performance across all participating clients. Compared to local learning, PA-CFL achieves a 5.4% improvement in R^2, a 69% reduction in RMSE, and a 45% decrease in MAE. Our approach enables effective FL through adaptive adjustments to diverse noise levels and the range of clients participating in each bubble. By grouping participants and proactively filtering out high-risk clients, PA-CFL mitigates potential threats to the FL system. The findings demonstrate PA-CFL's ability to enhance federated learning in time series prediction tasks with heterogeneous data, achieving a balance between forecasting accuracy and privacy preservation in retail applications. Additionally, PA-CFL's capability to detect and neutralize poisoned data from clients enhances the system's robustness and reliability.

联邦学习销售预测隐私保护Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。