arXiv:2605.07049cs.LGcs.AI2026-05被引 1

首次实现通用函数逼近下的隐私强化学习理论保证

Towards Differentially Private Reinforcement Learning with General Function Approximation

  • 结合分批更新与指数机制,设计新型隐私保护策略
  • 隐私条件下模型无监督学习的累计损失为~O(K^{3/5})
  • 适用于研究隐私与强化学习交叉问题的学者

我们首次为具有通用函数逼近能力的差分隐私在线强化学习提供了理论保障,突破了以往仅限于表格和线性设定的局限。方法结合分批策略与指数机制,并提出新的遗憾分析框架。即使在通用函数逼近下,模型无监督设置中的遗憾仍可达到与线性情况相当的最优水平,约为~O(K^{3/5}),其中K为回合数。作为重要副产品,我们还首次建立了依赖标准覆盖性复杂度的分批更新在线强化学习遗憾界,补充了基于新提出的Eluder-Condition类的结果。此外,揭示了近期关于线性函数逼近的私有强化学习结果中的根本性缺陷,厘清了该领域的研究格局。

原文摘要 · Abstract (English)

We present the first theoretical guarantees for differentially private online reinforcement learning (RL) with general function approximation, extending beyond prior work restricted to tabular and linear settings. Our approach combines a batched policy update scheme with the exponential mechanism, together with a novel regret analysis. We show that, even under general function approximation, the regret in the model-free setting under differential privacy matches the state of the art for the linear case, scaling as $\widetilde{O}(K^{3/5})$, where $K$ denotes the number of episodes. As an important by-product, we also establish the first regret bound for online RL with batch update that depends on the standard complexity measure of coverability, complementing existing results based on a newly introduced Eluder-Condition class. In addition, we uncover fundamental gaps in recent results for private RL with linear function approximation, thereby clarifying its landscape.

强化学习差分隐私函数逼近理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。