arXiv:2506.00379stat.MLcs.LG2025-06

解决联邦学习中标签分布不一致下的特征筛选难题

Label-shift robust federated feature screening for high-dimensional classification

  • 基于条件分布设计抗标签偏移的特征筛选方法
  • 在不同客户端数据分布下保持高筛选准确率
  • 适合隐私敏感场景下的高维数据分类任务

分布式与联邦学习是处理大规模高维分类数据的重要工具。为降低计算成本并克服维度灾难,特征筛选在数据预处理中至关重要,可剔除无关特征。然而,客户端间的数据异质性,尤其是标签分布偏移,给特征筛选带来严峻挑战。本文提出统一框架,整合现有筛选方法,并引入新型鲁棒联邦特征筛选方法(LR-FFS)及其联邦估计流程。该框架支持方法的统一分析,系统刻画其在标签偏移下的行为。基于此,LR-FFS利用条件分布函数与期望值,在不增加计算开销的前提下抵御标签偏移、模型误设及异常值影响。联邦实现保障计算效率与隐私安全,同时保持与集中式处理相当的筛选效果。此外,本文还提供一种联邦特征筛选的错误发现率(FDR)控制方法。实验与理论分析表明,LR-FFS在各类客户端环境(包括不同类别分布、样本量及缺失类别数据)中均表现优异。

原文摘要 · Abstract (English)

Distributed and federated learning are important tools for high-dimensional classification of large datasets. To reduce computational costs and overcome the curse of dimensionality, feature screening plays a pivotal role in eliminating irrelevant features during data preprocessing. However, data heterogeneity, particularly label shifting across different clients, presents significant challenges for feature screening. This paper introduces a general framework that unifies existing screening methods and proposes a novel utility, label-shift robust federated feature screening (LR-FFS), along with its federated estimation procedure. The framework facilitates a uniform analysis of methods and systematically characterizes their behaviors under label shift conditions. Building upon this framework, LR-FFS leverages conditional distribution functions and expectations to address label shift without adding computational burdens and remains robust against model misspecification and outliers. Additionally, the federated procedure ensures computational efficiency and privacy protection while maintaining screening effectiveness comparable to centralized processing. We also provide a false discovery rate (FDR) control method for federated feature screening. Experimental results and theoretical analyses demonstrate LR-FFS's superior performance across diverse client environments, including those with varying class distributions, sample sizes, and missing categorical data.

联邦学习特征筛选标签偏移高维分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。