arXiv:2505.09295cs.CYcs.AI2025-05

解决医疗联邦学习中的数据不均与偏见问题,提升模型公平性。

Toward Fair Federated Learning under Demographic Disparities and Data Imbalance

  • 结合公平性正则化与分组条件过采样,应对多维度敏感属性偏差。
  • 实验证明在临床数据上显著降低公平性指标方差,性能不下降。
  • 适用于医疗等高风险领域,兼顾隐私保护与公平建模。

在医疗等高风险领域应用人工智能时,确保公平性至关重要。基于不平衡且人口结构偏斜的数据训练的预测模型可能加剧现有不平等。联邦学习(FL)可在保护隐私的前提下实现跨机构协作,但仍易受算法偏见和子群体数据不均的影响,尤其当多个敏感属性交叉时。我们提出FedIDA(联邦学习中对不平衡与偏见的感知),一种框架无关的方法,将公平性感知正则化与分组条件过采样相结合。该方法支持多重敏感属性及异构数据分布,不改变底层联邦学习算法的收敛性。我们通过Lipschitz连续性和浓度不等式提供理论分析,证明了公平性提升的边界,并显示FedIDA降低了测试集上公平性度量的方差。在基准数据集和真实临床数据集上的实验表明,FedIDA在保持竞争力预测性能的同时持续改善公平性,验证了其在医疗领域实现公平且隐私保护建模的有效性。源代码已开源于GitHub。

原文摘要 · Abstract (English)

Ensuring fairness is critical when applying artificial intelligence to high-stakes domains such as healthcare, where predictive models trained on imbalanced and demographically skewed data risk exacerbating existing disparities. Federated learning (FL) enables privacy-preserving collaboration across institutions, but remains vulnerable to both algorithmic bias and subgroup imbalance - particularly when multiple sensitive attributes intersect. We propose FedIDA (Fed erated Learning for Imbalance and D isparity A wareness), a framework-agnostic method that combines fairness-aware regularization with group-conditional oversampling. FedIDA supports multiple sensitive attributes and heterogeneous data distributions without altering the convergence behavior of the underlying FL algorithm. We provide theoretical analysis establishing fairness improvement bounds using Lipschitz continuity and concentration inequalities, and show that FedIDA reduces the variance of fairness metrics across test sets. Empirical results on both benchmark and real-world clinical datasets confirm that FedIDA consistently improves fairness while maintaining competitive predictive performance, demonstrating its effectiveness for equitable and privacy-preserving modeling in healthcare. The source code is available on GitHub.

联邦学习公平性医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。