arXiv:2506.12358cs.LGcs.SY2025-06被引 2

用相对熵正则化实现高效加密策略生成,保护隐私同时保持计算效率。

Relative Entropy Regularized Reinforcement Learning for Efficient Encrypted Policy Synthesis

  • 基于相对熵正则化设计线性无极小值结构,便于集成全同态加密
  • 在加密环境下实现策略合成,量化与重加密误差可被理论控制
  • 适合需要隐私保护的强化学习场景,如医疗、金融领域

我们提出一种高效的加密策略生成方法,用于构建隐私保护的基于模型强化学习。首次证明相对熵正则化强化学习框架具有计算上便利的线性且无极小值(min-free)结构,支持直接高效地将全同态加密(FHE)与自举(bootstrapping)集成到策略合成中。分析了加密策略合成过程中的收敛性与误差界,考虑了由量化和自举引入的加密误差。理论分析通过数值仿真得到验证,结果表明该框架在结合FHE进行加密策略合成方面具有有效性。

原文摘要 · Abstract (English)

We propose an efficient encrypted policy synthesis to develop privacy-preserving model-based reinforcement learning. We first demonstrate that the relative-entropy-regularized reinforcement learning framework offers a computationally convenient linear and ``min-free'' structure for value iteration, enabling a direct and efficient integration of fully homomorphic encryption with bootstrapping into policy synthesis. Convergence and error bounds are analyzed as encrypted policy synthesis propagates errors under the presence of encryption-induced errors including quantization and bootstrapping. Theoretical analysis is validated by numerical simulations. Results demonstrate the effectiveness of the RERL framework in integrating FHE for encrypted policy synthesis.

强化学习加密计算隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。