arXiv:2603.22563stat.MLcs.LG2026-03被引 3

为人类反馈强化学习设计隐私保护框架,仅对奖励建模加噪仍能保持良好对齐效果。

Privacy-Preserving Reinforcement Learning from Human Feedback via Decoupled Reward Modeling

  • 仅对奖励学习阶段施加差分隐私,政策由私有奖励模型生成
  • 理论证明隐私带来额外可加误差项,且在不同数据量与隐私水平下表现各异
  • 在真实数据集上优于现有方法,适用于高隐私需求的RLHF场景

基于偏好微调已成为训练大语言模型的重要环节,但该阶段的数据可能包含敏感用户信息。核心问题是:如何为人类反馈强化学习(RLHF)设计一个契合其结构的差分隐私流水线?本文提出一种隐私保护框架,仅在奖励学习阶段施加差分隐私,并从得到的私有奖励模型中推导最终策略。理论上,我们分析了次优性差距,发现隐私会引入一个额外的加性项,超出常规非私有统计误差。我们还建立了极小极大下界,表明主导项随样本量和隐私级别变化,从而刻画出上界达到率最优(对数因子内)的适用区间。实验方面,合成实验验证了理论预测的缩放规律;在Anthropic HH-RLHF数据集上使用Gemma-2B-IT模型的实证显示,本方法在各类隐私预算下均优于现有的差分隐私基线方法。

原文摘要 · Abstract (English)

Preference-based fine-tuning has become an important component in training large language models, and the data used at this stage may contain sensitive user information. A central question is how to design a differentially private pipeline that is well suited to the distinct structure of reinforcement learning from human feedback. We propose a privacy-preserving framework that imposes differential privacy only on reward learning and derives the final policy from the resulting private reward model. Theoretically, we study the suboptimality gap and show that privacy contributes an additional additive term beyond the usual non-private statistical error. We also establish a minimax lower bound and show that the dominant term changes with sample size and privacy level, which in turn characterizes regimes in which the upper bound is rate-optimal up to logarithmic factors. Empirically, synthetic experiments confirm the scaling predicted by the theory, and experiments on the Anthropic HH-RLHF dataset using the Gemma-2B-IT model show stronger private alignment performance than existing differentially private baseline methods across privacy budgets.

强化学习隐私保护差分隐私语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。