提出可保护专家隐私的离线强化学习方法,兼顾安全与性能。
Preserving Expert-Level Privacy in Offline Reinforcement Learning
- 基于共识机制实现专家级差分隐私,兼容现有算法
- 在多个任务中优于基线,保持强学习性能
- 适用于医疗、广告等敏感数据场景
离线强化学习旨在从一个或多个行为策略(专家)收集的历史数据中学习最优策略。然而,专家可能对隐私敏感,因学习到的策略可能泄露其具体决策信息。在个性化检索、广告和医疗等领域,专家选择被视为敏感数据。为可证明地保护此类专家隐私,我们提出一种新型基于共识的专家级差分隐私离线强化学习训练方法,兼容任何现有离线强化学习算法。我们证明了严格的差分隐私保证,同时保持强大的实验性能。与现有差分隐私强化学习工作不同,我们在经典强化学习环境中进行了概念验证实验,涵盖大连续状态空间,在多个任务上显著优于自然基线。
原文摘要 · Abstract (English)
The offline reinforcement learning (RL) problem aims to learn an optimal policy from historical data collected by one or more behavioural policies (experts) by interacting with an environment. However, the individual experts may be privacy-sensitive in that the learnt policy may retain information about their precise choices. In some domains like personalized retrieval, advertising and healthcare, the expert choices are considered sensitive data. To provably protect the privacy of such experts, we propose a novel consensus-based expert-level differentially private offline RL training approach compatible with any existing offline RL algorithm. We prove rigorous differential privacy guarantees, while maintaining strong empirical performance. Unlike existing work in differentially private RL, we supplement the theory with proof-of-concept experiments on classic RL environments featuring large continuous state spaces, demonstrating substantial improvements over a natural baseline across multiple tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。