提出离线强化学习中数据删除的新方法,兼顾隐私与决策质量。
Unlearning Offline Stochastic Multi-Armed Bandits
- 基于高斯机制与回滚策略,自适应切换算法应对不同数据场景。
- 在固定样本和分布模型下均实现理论保障,证明可有效平衡隐私与性能。
- 适合关注数据隐私与强化学习结合的研究者参考。
机器遗忘旨在从已训练模型中移除特定数据点,为数据删除请求提供一种无需全量重训的隐私保护方案。现有研究主要集中在无监督/有监督学习的遗忘机制,而对序列决策系统中的遗忘问题关注较少。本文首次系统研究基础的序列决策问题:离线随机多臂老虎机(offline stochastic multi-armed bandits, MAB)。我们形式化了离线MAB的隐私约束,并以遗忘后的决策质量衡量效用。针对单源与多源遗忘场景,在固定样本模型与分布模型下展开分析。算法设计基于高斯机制与回滚两种经典基线,提出根据数据分布与隐私约束自适应切换的算法。进一步引入混合机制阐明基线原理。在各类设置下提供性能保证,并建立双模型下的下界。实验验证了预测的权衡关系,证明所提方法的有效性。
原文摘要 · Abstract (English)
Machine unlearning aims to unlearn data points from a learned model, offering a principled way to process data-deletion requests and mitigate privacy risks without full retraining. Prior work has mainly studied unsupervised / supervised machine unlearning, leaving unlearning for sequential decision-making systems far less understood. We initiate the first study of a foundational sequential decision-making problem: offline stochastic multi-armed bandits (MAB). We formalize the privacy constraint for offline MAB and measure utility by the post-unlearning decision quality. We conduct a systematic study of both single- and multi-source unlearning scenarios under two data-generation models, the fixed-sample model and the distribution model. For these settings, our algorithmic design is built on two canonical base algorithms: Gaussian mechanism and rollback, and we propose adaptive algorithms that switch between them according to the data regime and privacy constraint. We further introduce a mixing procedure that elucidates the rationale behind these baselines. We provide performance guarantees across the above settings and establish lower bounds under both dataset models. Experiments validate the predicted tradeoffs and demonstrate the effectiveness of the proposed methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。