提出隐私保护强化学习新算法,兼顾隐私与性能。
Differentially Private Policy Gradient
- 将隐私保护转化为信任区域计算,避免性能损失。
- 实测在多个基准上优于现有在线强化学习隐私算法。
- 适合关注数据隐私的强化学习应用开发者。
随着强化学习在现实世界中的广泛应用,对个人数据的大量消耗引发隐私担忧。本文提出一种差分隐私(DP)策略梯度算法,证明在该场景下引入差分隐私可转化为计算合适的信任区域,从而避免无隐私保护方法的理论性质损失。因此,可通过调节隐私噪声与信任区域大小的权衡,实现高性能的差分隐私策略梯度算法。我们在多个基准任务上验证了该方法的性能,结果表明其在复杂任务上的表现显著优于现有的在线强化学习差分隐私算法。
原文摘要 · Abstract (English)
Motivated by the increasing deployment of reinforcement learning in the real world, involving a large consumption of personal data, we introduce a differentially private (DP) policy gradient algorithm. We show that, in this setting, the introduction of Differential Privacy can be reduced to the computation of appropriate trust regions, thus avoiding the sacrifice of theoretical properties of the DP-less methods. Therefore, we show that it is possible to find the right trade-off between privacy noise and trust-region size to obtain a performant differentially private policy gradient algorithm. We then outline its performance empirically on various benchmarks. Our results and the complexity of the tasks addressed represent a significant improvement over existing DP algorithms in online RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。