arXiv:2504.00921cs.LG2025-04被引 1

首个针对表格数据联邦学习的遗忘方法基准研究,解决隐私与效率难题。

Benchmarking Federated Machine Unlearning methods for Tabular Data

  • 在联邦学习中对比特征与实例级遗忘,用随机森林和逻辑回归测试
  • 树模型精确遗忘能力强,梯度方法计算更快但精度稍低
  • 适合关注隐私安全的工业级模型部署者参考

机器遗忘使模型可按需删除特定数据,对隐私敏感的机器学习至关重要,尤其在联邦学习(FL)环境中。本文首次系统性地对跨孤岛联邦学习场景下的表格数据遗忘方法进行基准测试,应对数据隐私与通信效率的核心挑战。研究在特征和实例两个层面开展遗忘,采用随机森林与逻辑回归模型,对比了微调与基于梯度的多种遗忘算法,在多个数据集上评估其保真度、可验证性和计算效率。实验表明,尽管各类方法保真度均较高,但树模型在可验证性方面表现优异,能实现精确遗忘;而基于梯度的方法则在计算效率上更具优势。本研究为联邦学习环境下的遗忘算法设计与选择提供了关键洞见,奠定了隐私保护机器学习进一步研究的基础。

原文摘要 · Abstract (English)

Machine unlearning, which enables a model to forget specific data upon request, is increasingly relevant in the era of privacy-centric machine learning, particularly within federated learning (FL) environments. This paper presents a pioneering study on benchmarking machine unlearning methods within a federated setting for tabular data, addressing the unique challenges posed by cross-silo FL where data privacy and communication efficiency are paramount. We explore unlearning at the feature and instance levels, employing both machine learning, random forest and logistic regression models. Our methodology benchmarks various unlearning algorithms, including fine-tuning and gradient-based approaches, across multiple datasets, with metrics focused on fidelity, certifiability, and computational efficiency. Experiments demonstrate that while fidelity remains high across methods, tree-based models excel in certifiability, ensuring exact unlearning, whereas gradient-based methods show improved computational efficiency. This study provides critical insights into the design and selection of unlearning algorithms tailored to the FL environment, offering a foundation for further research in privacy-preserving machine learning.

联邦学习机器遗忘隐私保护表格数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。