arXiv:2410.09124cs.LGcs.AI2024-10

让多方协作训练模型时,能验证各方没作弊。

SoK: Verifiable Cross-Silo FL

  • 设计可验证的跨孤岛联邦学习协议,确保各方按规则参与
  • 对比不同方案在效率与安全防御上的表现差异
  • 适合关注隐私保护与可信计算的研究者和从业者

联邦学习(FL)是一种广泛采用的方法,允许在分布在多个设备上的数据上训练机器学习模型。在跨孤岛联邦学习中,参与者数量适中,每个参与方通常代表一个明确的组织,如医疗或金融领域中的医院或数据枢纽。尽管这些实体是公认的,但仍可能存在恶意行为者试图干扰训练过程以获取利益,例如制造偏倚结果或减少计算负担。当训练数据公开时,恶意行为较易被发现;但若需保持数据隐私,则问题更为严峻。为此,近年来出现了对可验证协议的兴趣,使各方能够验证其他方是否遵守训练流程并正确执行计算。本文系统化梳理了可验证跨孤岛联邦学习领域的知识,分析多种协议,构建分类体系,并比较其效率与威胁模型。同时评估零知识证明(ZKP)方案,探讨如何在联邦学习场景中降低其整体开销。最后,识别研究空白并提出未来方向。

原文摘要 · Abstract (English)

Federated Learning (FL) is a widespread approach that allows training machine learning (ML) models with data distributed across multiple devices. In cross-silo FL, which often appears in domains like healthcare or finance, the number of participants is moderate, and each party typically represents a well-known organization. For instance, in medicine data owners are often hospitals or data hubs which are well-established entities. However, malicious parties may still attempt to disturb the training procedure in order to obtain certain benefits, for example, a biased result or a reduction in computational load. While one can easily detect a malicious agent when data used for training is public, the problem becomes much more acute when it is necessary to maintain the privacy of the training dataset. To address this issue, there is recently growing interest in developing verifiable protocols, where one can check that parties do not deviate from the training procedure and perform computations correctly. In this paper, we present a systematization of knowledge on verifiable cross-silo FL. We analyze various protocols, fit them in a taxonomy, and compare their efficiency and threat models. We also analyze Zero-Knowledge Proof (ZKP) schemes and discuss how their overall cost in a FL context can be minimized. Lastly, we identify research gaps and discuss potential directions for future scientific work.

联邦学习可验证性隐私保护零知识证明

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。