解决联邦学习中用户入选与参与偏差问题,提升模型对目标人群的代表性。
Who Trains Matters: Federated Learning under Enrollment and Participation Selection Biases

- 提出两阶段选择模型,区分用户入选和参与偏差
- 设计逆概率加权算法FedIPW,纠正训练目标与目标人群的偏差
- 适用于数据分布不均、用户参与不均衡的真实联邦学习场景
联邦学习(FL)从分布式客户端的更新中训练共享模型,通常隐含假设参与客户端能代表目标人群。但在实际中,这一假设可能在两个阶段失效,导致选择偏差。首先,设备限制、软件要求或用户同意等资格规则决定哪些客户端可被纳入训练,引发‘入选偏差’;其次,在已纳入客户端中,电池状态、网络状况和本地时间等因素决定每轮通信中哪些客户端参与,引发‘参与偏差’。尽管现有研究多关注轮次级参与偏差,却较少关注影响更持久的群体级入选偏差,后者会导致训练目标与目标人群目标间的持续失配。本文形式化了两阶段选择下的联邦学习,提出 extsc{FedIPW}——一种逆概率加权聚合方案,在标准忽略性和正性假设下可恢复目标人群的平均更新。由于非入选客户端的个体协变量常不可得,还引入有限信息的汇总校准扩展,利用已知的目标人群统计量对已入选样本重新加权,部分纠正入选偏差。进一步提供算法无关的优化分析,表明不完全的偏差修正会导致非消失的偏差下界。合成联邦逻辑回归实验验证了预测的目标偏差,并显示入选修正能降低两阶段选择下的目标人群误差。
原文摘要 · Abstract (English)
Federated learning (FL) trains a shared model from updates contributed by distributed clients, often implicitly assuming that contributing clients are representative of the target population. In practice, this representativeness assumption can fail at two distinct stages, inducing selection bias. First, eligibility rules such as device constraints, software requirements, or user consent determine which clients are ever enrolled and reachable for training, inducing \emph{enrollment bias}. Second, among enrolled clients, user and system factors such as battery state, network status, and local time determine which clients participate in each communication round, inducing \emph{participation bias}. Although existing work has largely addressed round-level participation bias, it has paid far less attention to population-level enrollment bias, which can induce a persistent mismatch between the training objective and the target-population objective. We formalize FL under a two-stage selection model and derive \textsc{FedIPW}, an inverse-probability-weighted aggregation scheme that recovers the target-population mean update under standard ignorability and positivity assumptions. Because client-level covariates are often unavailable for non-enrolled clients, we also introduce a limited-information aggregate-calibration extension that uses known target-population summaries to reweight the enrolled sample, partially correcting enrollment bias. We further provide an algorithm-agnostic optimization analysis under residual weighting error and show that incomplete selection correction can induce a non-vanishing bias floor. Finally, experiments on synthetic federated logistic regression validate the predicted objective mismatch and show that enrollment correction reduces target-population error under two-stage selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。