提出新框架,量化联邦学习中跨客户端的数据记忆风险。
Exploring Cross-Client Memorization of Training Data in Large Language Models for Federated Learning
- 设计细粒度跨样本记忆检测方法,覆盖所有客户端数据。
- 发现联邦模型更易记住本客户端数据,受解码策略等影响。
- 适合关注隐私泄露与模型安全的研究者和工程师。
联邦学习(FL)可在不共享原始数据的情况下实现协作训练,但仍存在训练数据记忆的风险。现有检测技术仅针对单个样本,低估了跨样本记忆的潜在威胁。尽管中心化学习中已有细粒度跨样本记忆评估方法,但其依赖数据集中访问,无法直接用于联邦学习。本文提出一种新框架,通过细粒度跨样本记忆测量,量化联邦学习中客户端内部及跨客户端的记忆行为。基于此框架,开展两项研究:(1) 检测跨客户端的细微记忆现象;(2) 探究影响记忆的关键因素,包括解码策略、前缀长度和联邦算法。结果表明,联邦模型确实会记忆客户数据,尤其在客户端内部记忆更强,且记忆程度受训练与推理因素影响。
原文摘要 · Abstract (English)
Federated learning (FL) enables collaborative training without raw data sharing, but still risks training data memorization. Existing FL memorization detection techniques focus on one sample at a time, underestimating more subtle risks of cross-sample memorization. In contrast, recent work on centralized learning (CL) has introduced fine-grained methods to assess memorization across all samples in training data, but these assume centralized access to data and cannot be applied directly to FL. We bridge this gap by proposing a framework that quantifies both intra- and inter-client memorization in FL using fine-grained cross-sample memorization measurement across all clients. Based on this framework, we conduct two studies: (1) measuring subtle memorization across clients and (2) examining key factors that influence memorization, including decoding strategies, prefix length, and FL algorithms. Our findings reveal that FL models do memorize client data, particularly intra-client data, more than inter-client data, with memorization influenced by training and inferencing factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。