研究隐私保护强化学习的样本需求,发现隐私成本常为次要影响。
On the Sample Complexity of Differentially Private Policy Optimization
- 针对策略优化设计专用差分隐私定义,解决在线学习隐私难题。
- 理论证明多数算法在隐私约束下样本复杂度提升有限。
- 为医疗、机器人等敏感场景的隐私安全训练提供依据。
策略优化(PO)是现代强化学习的核心,广泛应用于机器人、医疗和大语言模型训练。随着其在敏感领域的部署增多,隐私问题日益突出。本文首次对差分隐私策略优化的样本复杂度进行理论研究。我们提出了适用于PO的差分隐私定义,解决了在线学习动态带来的隐私单位界定难题。通过统一框架,系统分析了策略梯度(PG)、自然策略梯度(NPG)等主流算法在不同设置下的样本复杂度。理论结果表明,隐私代价通常仅表现为样本复杂度中的低阶项,同时揭示了私有化策略优化中的细微但关键现象,为隐私保护算法设计提供了重要实践启示。
原文摘要 · Abstract (English)
Policy optimization (PO) is a cornerstone of modern reinforcement learning (RL), with diverse applications spanning robotics, healthcare, and large language model training. The increasing deployment of PO in sensitive domains, however, raises significant privacy concerns. In this paper, we initiate a theoretical study of differentially private policy optimization, focusing explicitly on its sample complexity. We first formalize an appropriate definition of differential privacy (DP) tailored to PO, addressing the inherent challenges arising from on-policy learning dynamics and the subtlety involved in defining the unit of privacy. We then systematically analyze the sample complexity of widely-used PO algorithms, including policy gradient (PG), natural policy gradient (NPG) and more, under DP constraints and various settings, via a unified framework. Our theoretical results demonstrate that privacy costs can often manifest as lower-order terms in the sample complexity, while also highlighting subtle yet important observations in private PO settings. These offer valuable practical insights for privacy-preserving PO algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。