提出可保留模型性能的数据删除框架,兼顾合规与商业价值。
From Machine Learning to Machine Unlearning: Complying with GDPR's Right to be Forgotten while Maintaining Business Value of Predictive Models
- 用集成学习构建高精度模型,支持高效数据删除
- 基于知识蒸馏的去训练方法,实现快速且完整的数据擦除
- 适合需遵守隐私法规的企业级预测模型维护
近年来的隐私法规(如GDPR)赋予数据主体‘被遗忘的权利’(RTBF),要求企业响应数据擦除请求。然而,企业在从已训练好的预测模型中删除特定训练数据时面临巨大挑战。现有机器去训练方法虽能快速擦除数据,但常导致模型性能下降,引发经济损失并影响合规性。本文提出一个从机器学习到去训练的全流程框架ETID(Ensemble-based iTerative Information Distillation),通过新型集成学习构建准确模型,并设计适配该模型的蒸馏式去训练机制,实现高效、有效的数据擦除。大量实验表明,ETID优于多种前沿方法,在保持模型质量的同时显著提升去训练效率。本工作展示了其在推动合法、可持续的数据与预测服务市场中的潜力。
原文摘要 · Abstract (English)
Recent privacy regulations (e.g., GDPR) grant data subjects the `Right to Be Forgotten' (RTBF) and mandate companies to fulfill data erasure requests from data subjects. However, companies encounter great challenges in complying with the RTBF regulations, particularly when asked to erase specific training data from their well-trained predictive models. While researchers have introduced machine unlearning methods aimed at fast data erasure, these approaches often overlook maintaining model performance (e.g., accuracy), which can lead to financial losses and non-compliance with RTBF obligations. This work develops a holistic machine learning-to-unlearning framework, called Ensemble-based iTerative Information Distillation (ETID), to achieve efficient data erasure while preserving the business value of predictive models. ETID incorporates a new ensemble learning method to build an accurate predictive model that can facilitate handling data erasure requests. ETID also introduces an innovative distillation-based unlearning method tailored to the constructed ensemble model to enable efficient and effective data erasure. Extensive experiments demonstrate that ETID outperforms various state-of-the-art methods and can deliver high-quality unlearned models with efficiency. We also highlight ETID's potential as a crucial tool for fostering a legitimate and thriving market for data and predictive services.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。