arXiv:2502.17515cs.LGcs.AI2025-02被引 13

提出用户级隐私保护的强化学习框架,提升人类反馈下的隐私安全

Towards User-level Private Reinforcement Learning with Human Feedback

  • 将用户级标签差分隐私融入强化学习人类反馈流程
  • 在情感生成与摘要任务中实现更优隐私-效用平衡
  • 适用于重视用户数据隐私的AI对齐场景

基于人类反馈的强化学习(RLHF)已成为对齐大语言模型与人类偏好的重要技术。然而,如何保护用户偏好隐私成为关键挑战。现有工作多关注项级隐私保护,对更常见的用户级隐私保护效果不佳。本文提出AUP-RLHF框架,首次将用户级标签差分隐私(user-level label DP)引入RLHF。我们证明经典随机响应算法在用户级设置下会导致效用下降,并建立了用户级标签DP-RLHF的理论下界。所提AUP-RLHF算法在保证(ε, δ)用户级隐私的前提下,实现了更优的估计误差。实验表明,在情感生成和摘要任务中,AUP-RLHF优于现有基线方法,显著提升了隐私-效用权衡性能。

原文摘要 · Abstract (English)

Reinforcement Learning with Human Feedback (RLHF) has emerged as an influential technique, enabling the alignment of large language models (LLMs) with human preferences. Despite the promising potential of RLHF, how to protect user preference privacy has become a crucial issue. Most previous work has focused on using differential privacy (DP) to protect the privacy of individual data. However, they have concentrated primarily on item-level privacy protection and have unsatisfactory performance for user-level privacy, which is more common in RLHF. This study proposes a novel framework, AUP-RLHF, which integrates user-level label DP into RLHF. We first show that the classical random response algorithm, which achieves an acceptable performance in item-level privacy, leads to suboptimal utility when in the user-level settings. We then establish a lower bound for the user-level label DP-RLHF and develop the AUP-RLHF algorithm, which guarantees $(\varepsilon, δ)$ user-level privacy and achieves an improved estimation error. Experimental results show that AUP-RLHF outperforms existing baseline methods in sentiment generation and summarization tasks, achieving a better privacy-utility trade-off.

强化学习隐私保护大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。