首个测试联邦学习公平性的基准框架,揭示模型在客户端的隐性偏见。
FeDa4Fair: Client-Level Federated Datasets for Fairness Evaluation
- 构建可定制的联邦数据集,模拟不同客户端对敏感属性的异质偏见
- 提供标准化评估套件,支持公平性方法在真实复杂场景下的测试
- 适用于研究联邦学习公平性、算法鲁棒性及偏见检测的学者与工程师
联邦学习(FL)可在保护隐私的前提下实现协同训练,但带来一个关键挑战:‘公平性假象’。全局模型在服务器端评估时看似公平,但在客户端层面仍存在持续歧视。现有公平性增强的FL方法通常仅缓解单一敏感属性(如二元)的偏差,忽略两个现实且矛盾的情形:属性偏差(客户端对不同敏感属性表现出不公平)和取值偏差(客户端对同一属性的不同取值有冲突的偏见)。为推动更稳健、可复现的联邦学习公平性研究,我们提出FeDa4Fair,首个专为在异质客户端偏见条件下压力测试公平性方法而设计的基准框架。贡献包括:(1) 提出FeDa4Fair库,可生成用于评估公平联邦学习方法的定制化数据集;(2) 发布由该库生成的基准套件,以标准化公平性方法的评估流程;(3) 提供开箱即用的公平性评估函数,直接用于这些数据集。
原文摘要 · Abstract (English)
Federated Learning (FL) enables collaborative training while preserving privacy, yet it introduces a critical challenge: the "illusion of fairness''. A global model, usually evaluated on the server, appears fair on average while keeping persistent discrimination at the client level. Current fairness-enhancing FL solutions often fall short, as they typically mitigate biases for a single, usually binary, sensitive attribute, while ignoring two realistic and conflicting scenarios: attribute-bias (where clients are unfair toward different sensitive attributes) and value-bias (where clients exhibit conflicting biases toward different values of the same attribute). To support more robust and reproducible fairness research in FL, we introduce FeDa4Fair, the first benchmarking framework designed to stress-test fairness methods under these heterogeneous conditions. Our contributions are three-fold: (1) We introduce FeDa4Fair, a library designed to create datasets tailored to evaluating fair FL methods under heterogeneous client bias; (2) we release a benchmark suite generated by the FeDa4Fair library to standardize the evaluation of fair FL methods; (3) we provide ready-to-use functions for evaluating fairness outcomes for these datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。