随机探索自带隐私保护,无需额外加噪
Randomized Least Squares Value Iteration itself is Joint Differentially Private
- 利用随机探索的噪声自然实现联合差分隐私
- 在表格型马尔可夫决策过程上达到ε(δ)隐私,其中ε(δ)=2AK/H²log(2HSA)+...
- 适合医疗、推荐等敏感场景的隐私强化强化学习
随着强化学习在医疗、推荐系统等敏感领域应用日益广泛,隐私保护技术变得至关重要。本文研究了基于随机探索的强化学习算法在周期性设置下的隐私保护问题,重点关注随机最小二乘值迭代(RLSVI)。我们提出一种新的隐私分析,揭示了用于探索的噪声同时具备隐私保护作用。具体而言,我们证明了在表格型马尔可夫决策过程(MDP)中,RLSVI本身即为(ε(δ), δ)-联合差分私有,其中ε(δ) = 2AK/(H²log(2HSA)) + 2√(2AK log(1/δ)/(H²log(2HSA))),S和A分别为状态数和动作数,H为每轮长度,K为总轮次数。
原文摘要 · Abstract (English)
As reinforcement learning (RL) increasingly applies to sensitive domains, such as health care and recommendation systems, privacy-preserving techniques have become essential to protect users' sensitive information. We investigate privacy-preserving RL under an episodic setting, focusing on algorithms based on randomized exploration, such as Randomized Least Squares Value Iteration (RLSVI). The overall goal is to study how randomized exploration interacts with the injected noise required by privacy mechanisms. In this work, we show a new privacy analysis that characterizes how the noise in RLSVI set for exploration simultaneously provides privacy protection. Specifically, we prove that RLSVI is $(\varepsilon(δ),δ)$-joint differentially private in tabular MDP as is with $\varepsilon(δ) = \frac{2AK}{H^2\log(2HSA)} + 2\sqrt{\frac{2AK\log(1/δ)}{H^2\log(2HSA)}}$, where $S$ and $A$ are the number of states and actions respectively, $H$ is the length of an episode and $K$ is the number of episodes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。