arXiv:2504.17448cs.LGcs.DB2025-04中稿 · TKDE 2025

针对联邦学习中客户端数据异构问题,提出高效标注选择方法

CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning

  • 通过分析推理不一致性追踪高认知不确定性样本
  • 在多个数据集上显著提升模型准确率与标注效率
  • 适合数据分布差异大的联邦学习场景使用

主动学习(AL)通过选择最具信息量的未标注数据减少人工标注成本,但单独执行仍受限于数据多样性不足和标注预算有限。联邦主动学习(FAL)通过协同数据选择与模型训练,在保护原始数据隐私的同时缓解该问题。然而现有FAL方法忽视客户端间数据分布异构性及其导致的全局与局部模型参数波动,影响模型精度。为此,我们提出针对客户端异构性的数据选择方法CHASe(Client Heterogeneity-Aware Data Selection)。CHASe聚焦于识别在训练过程中决策边界附近剧烈振荡的高认知不确定性样本(EVs)。为兼顾有效性与效率,模型采用三项技术:1)通过分析训练周期间的推理不一致性追踪EVs;2)引入新对齐损失校准低精度模型的决策边界;3)通过子集采样与数据冻结唤醒机制提升数据选择效率。实验表明,CHASe在多种数据集、模型复杂度及异构联邦设置下均优于多个基线方法,兼具高效性与有效性。

原文摘要 · Abstract (English)

Active learning (AL) reduces human annotation costs for machine learning systems by strategically selecting the most informative unlabeled data for annotation, but performing it individually may still be insufficient due to restricted data diversity and annotation budget. Federated Active Learning (FAL) addresses this by facilitating collaborative data selection and model training, while preserving the confidentiality of raw data samples. Yet, existing FAL methods fail to account for the heterogeneity of data distribution across clients and the associated fluctuations in global and local model parameters, adversely affecting model accuracy. To overcome these challenges, we propose CHASe (Client Heterogeneity-Aware Data Selection), specifically designed for FAL. CHASe focuses on identifying those unlabeled samples with high epistemic variations (EVs), which notably oscillate around the decision boundaries during training. To achieve both effectiveness and efficiency, \model{} encompasses techniques for 1) tracking EVs by analyzing inference inconsistencies across training epochs, 2) calibrating decision boundaries of inaccurate models with a new alignment loss, and 3) enhancing data selection efficiency via a data freeze and awaken mechanism with subset sampling. Experiments show that CHASe surpasses various established baselines in terms of effectiveness and efficiency, validated across diverse datasets, model complexities, and heterogeneous federation settings.

联邦学习主动学习数据选择异构性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。