在部分观测与单侧奖励下实现强隐私保护的强化学习
Privacy Preserving Reinforcement Learning with One-Sided Feedback
- 设计新算法POOL,支持高维连续空间下的隐私保护学习
- 理论证明样本复杂度逼近非私密强化学习下界
- 适合需要隐私保障的高维强化学习场景
我们研究在多维连续状态和动作空间中,仅接收部分状态观测且仅在部分状态-动作空间获得奖励信息的强化学习问题。该设置带来学习效率与隐私保护双重挑战。为此,我们提出新的隐私保护强化学习算法POOL。对POOL进行系统性理论分析,推导出样本复杂度上界,其与非私密强化学习的已知下界一致。其中,E_rho为隐私参数,H为时间跨度,alpha为最优差距参数。结果表明,在保持高学习效率的同时可实现强隐私保障,为多维环境中具有单侧反馈的实用化隐私感知强化学习迈出关键一步。
原文摘要 · Abstract (English)
We study reinforcement learning (RL) in multi-dimensional continuous state and action spaces with one-sided feedback, where the agent receives partial observations of the state and obtains reward information for only a subset of the state-action space at each time step. This setting introduces substantial challenges in both learning efficiency and privacy preservation. To address these challenges, we propose POOL, a novel privacy-preserving RL algorithm. We conduct a comprehensive theoretical analysis of POOL, deriving a sample complexity bound that matches the known lower bounds for non-private RL. Here, E_rho denotes the privacy parameter, H is the time horizon, and alpha is the optimality-gap parameter. Our findings show that it is possible to enforce strong privacy guarantees while maintaining high learning efficiency, marking a significant step toward practical, privacy-aware RL in multi-dimensional environments with one-sided feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。