arXiv:2506.03954cs.LGcs.AI2025-06KDD被引 12

首个异构联邦学习库,支持多场景公平比较与评估。

HtFLlib: A Comprehensive Heterogeneous Federated Learning Library and Benchmark

  • 整合12个跨领域数据集与40种异构模型架构,支持多样化实验。
  • 包含10种主流异构联邦学习方法,覆盖准确率、通信开销等多维度评估。
  • 适合研究者快速验证新算法,尤其适用于医疗与传感器场景。

随着人工智能发展,异构模型协作可通过跨机构和设备的知识迁移缓解数据稀缺问题。传统联邦学习仅支持同构模型,限制了异构架构间的协作。为此,异构联邦学习(HtFL)方法被提出,以实现多样模型间的协同并应对数据异构性。然而,当前缺乏对快速增长的HtFL方法进行标准化评估与分析的综合性基准。首先,数据集差异大、模型异构场景多样、方法实现不一,导致难以公平比较。其次,现有方法在医疗、传感器信号等领域的有效性与鲁棒性尚未充分探索。为此,我们推出首个异构联邦学习库(HtFLlib),一个易用且可扩展的框架,集成多个数据集与异构场景,提供稳健的研究与应用基准。具体包括:(1)12个涵盖多种领域、模态和数据异构性的数据集;(2)40种从小型到大型的跨模态模型架构;(3)模块化、可扩展的代码库,实现10种代表性HtFL方法;(4)在准确率、收敛速度、计算成本和通信成本等方面的系统性评估。我们强调先进HtFL方法的优势与潜力,期望HtFLlib推动该领域研究进展并拓展其应用。代码已开源:https://github.com/TsingZ0/HtFLlib。

原文摘要 · Abstract (English)

As AI evolves, collaboration among heterogeneous models helps overcome data scarcity by enabling knowledge transfer across institutions and devices. Traditional Federated Learning (FL) only supports homogeneous models, limiting collaboration among clients with heterogeneous model architectures. To address this, Heterogeneous Federated Learning (HtFL) methods are developed to enable collaboration across diverse heterogeneous models while tackling the data heterogeneity issue at the same time. However, a comprehensive benchmark for standardized evaluation and analysis of the rapidly growing HtFL methods is lacking. Firstly, the highly varied datasets, model heterogeneity scenarios, and different method implementations become hurdles to making easy and fair comparisons among HtFL methods. Secondly, the effectiveness and robustness of HtFL methods are under-explored in various scenarios, such as the medical domain and sensor signal modality. To fill this gap, we introduce the first Heterogeneous Federated Learning Library (HtFLlib), an easy-to-use and extensible framework that integrates multiple datasets and model heterogeneity scenarios, offering a robust benchmark for research and practical applications. Specifically, HtFLlib integrates (1) 12 datasets spanning various domains, modalities, and data heterogeneity scenarios; (2) 40 model architectures, ranging from small to large, across three modalities; (3) a modularized and easy-to-extend HtFL codebase with implementations of 10 representative HtFL methods; and (4) systematic evaluations in terms of accuracy, convergence, computation costs, and communication costs. We emphasize the advantages and potential of state-of-the-art HtFL methods and hope that HtFLlib will catalyze advancing HtFL research and enable its broader applications. The code is released at https://github.com/TsingZ0/HtFLlib.

联邦学习异构模型基准测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。