arXiv:2508.08875cs.LGcs.AI2025-08AAAI被引 4

让联邦大模型能删除特定用户数据,满足隐私法规要求。

Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models

  • 将联邦学习与数据删除结合为双重优化目标。
  • 在多个算法上验证,删除效果好且模型性能损失小。
  • 适合关注数据合规与模型可撤销性的研究者。

大语言模型(LLMs)越来越多地采用联邦学习(FL)来利用私有的、任务特定的数据集进行微调,同时保护数据隐私。然而,尽管联邦学习框架能实现无需共享原始数据的协作训练,却缺乏内置机制以满足如欧盟GDPR中“被遗忘权”的监管要求。引入私有数据加剧了数据质量与长期治理的担忧,而现有分布式训练框架无法提供原则性方法来在训练后选择性移除特定客户端的贡献。由于数据孤岛、严格的隐私限制以及模型聚合的相互依赖性,联邦大模型的去学习(unlearning)远比集中式场景复杂。为此,我们提出Oblivionis,一个轻量级的学习与去学习框架,使客户端能够在联邦大模型训练过程中选择性地删除特定私有数据,提升可信度与合规性。通过将联邦学习与去学习统一为双目标优化,我们整合了6种联邦学习与5种去学习算法,进行全面评估与对比分析,建立了一条稳健的联邦大模型去学习流水线。大量实验表明,Oblivionis优于本地训练,在遗忘效果与模型效用之间取得稳健平衡,跨算法比较为未来大模型发展提供了清晰方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) increasingly leverage Federated Learning (FL) to utilize private, task-specific datasets for fine-tuning while preserving data privacy. However, while federated LLM frameworks effectively enable collaborative training without raw data sharing, they critically lack built-in mechanisms for regulatory compliance like GDPR's right to be forgotten. Integrating private data heightens concerns over data quality and long-term governance, yet existing distributed training frameworks offer no principled way to selectively remove specific client contributions post-training. Due to distributed data silos, stringent privacy constraints, and the intricacies of interdependent model aggregation, federated LLM unlearning is significantly more complex than centralized LLM unlearning. To address this gap, we introduce Oblivionis, a lightweight learning and unlearning framework that enables clients to selectively remove specific private data during federated LLM training, enhancing trustworthiness and regulatory compliance. By unifying FL and unlearning as a dual optimization objective, we incorporate 6 FL and 5 unlearning algorithms for comprehensive evaluation and comparative analysis, establishing a robust pipeline for federated LLM unlearning. Extensive experiments demonstrate that Oblivionis outperforms local training, achieving a robust balance between forgetting efficacy and model utility, with cross-algorithm comparisons providing clear directions for future LLM development.

联邦学习大模型数据删除隐私合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。