arXiv:2412.16079cs.LGcs.CV2024-12

用博弈论方法解决分布式学习中的数据不均衡问题。

Fair Distributed Machine Learning with Imbalanced Data as a Stackelberg Evolutionary Game

  • 将分布式学习建模为斯塔克尔伯格演化博弈,动态分配节点权重。
  • 自适应算法使弱势节点AUC提升2.713%,大样本节点仅降0.441%。
  • 适合医疗等数据分布不均场景,提升公平性与模型鲁棒性。

去中心化学习可在不集中数据的前提下训练深度学习模型,提升数据隐私、运行效率并推动数据所有权政策。然而,数据不均衡在该框架下带来挑战:小数据参与者性能显著落后于大数据参与者。这一问题在医疗领域尤为突出,由患者群体差异、技术不平等及数据采集实践不同所致。本文将分布式学习视为斯塔克尔伯格演化博弈,提出两种节点贡献权重设置算法:确定性斯塔克尔伯格加权模型(DSWM)与自适应斯塔克尔伯格加权模型(ASWM)。通过三个医学数据集验证动态加权对欠代表节点的影响。结果表明,ASWM显著提升欠代表节点表现,其AUC提升2.713%;而大数据节点平均性能仅下降0.441%。

原文摘要 · Abstract (English)

Decentralised learning enables the training of deep learning algorithms without centralising data sets, resulting in benefits such as improved data privacy, operational efficiency and the fostering of data ownership policies. However, significant data imbalances pose a challenge in this framework. Participants with smaller datasets in distributed learning environments often achieve poorer results than participants with larger datasets. Data imbalances are particularly pronounced in medical fields and are caused by different patient populations, technological inequalities and divergent data collection practices. In this paper, we consider distributed learning as an Stackelberg evolutionary game. We present two algorithms for setting the weights of each node's contribution to the global model in each training round: the Deterministic Stackelberg Weighting Model (DSWM) and the Adaptive Stackelberg Weighting Model (ASWM). We use three medical datasets to highlight the impact of dynamic weighting on underrepresented nodes in distributed learning. Our results show that the ASWM significantly favours underrepresented nodes by improving their performance by 2.713% in AUC. Meanwhile, nodes with larger datasets experience only a modest average performance decrease of 0.441%.

分布式学习数据不平衡博弈论医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。