arXiv:2608.26324cs.LG2026-08

给奖励模型加噪声,实现隐私保护与推理对齐的双赢。

Privacy Without Regret: Differentially Private Inference-Time Alignment

  • 用校准过的高斯或Gumbel噪声替代原始奖励分数进行选择。
  • 在隐私预算超过临界值ε*时,噪声既是隐私保障也最优对齐正则。
  • 新方法适用于强隐私场景,且不随候选数量增加而性能下降。

最佳N选一(BoN)是当前最广泛使用的推理阶段对齐策略,但存在两大问题:奖励劫持(即响应利用代理奖励模型的缺陷)以及训练该模型的人类偏好数据缺乏隐私保护。本文证明,在选择前向奖励分数添加校准噪声,可同时解决这两个问题。提出的私有最佳N选一(PrivBoN)表明,适当尺度的Gumbel噪声能同时实现ε-差分隐私和KL正则化对齐。当隐私预算超过临界阈值ε*时,隐私所需的噪声恰好是最优后悔率正则,隐私无额外对齐代价——达到黄等(2025)提出的信息论极限。由于ε*依赖未知覆盖率系数,本文进一步提出私有推理阶段悲观主义(PrivITP),结合χ²正则化拒采与两阶段高斯机制,实现事后(ε,δ)-差分隐私,且隐私成本与候选数n无关,正则化参数与隐私参数完全解耦,性能逼近极限,仅受噪声膨胀项影响。在多个语言模型、数据集和奖励模型上的实验验证了上述结论:PrivBoN和PrivITP均具缩放单调性(而传统BoN在某临界n后性能下降),且在同等隐私水平下,PrivITP表现优于或相当于是PrivBoN,尤其在强隐私条件下优势显著。

原文摘要 · Abstract (English)

Best-of-N (BoN) sampling is the simplest and most widely deployed inference-time alignment strategy, but it suffers from two distinct problems: reward hacking, in which the selected response exploits errors in the proxy reward model, and the absence of any privacy protection for the sensitive human preference data used to train that reward model. We show that a single intervention-adding calibrated noise to reward scores before selection-resolves both. Our first result, Private Best-of-N (PrivBoN), establishes that Gumbel noise at an appropriate scale simultaneously provides $ε$-differential privacy and implements KL-regularized alignment. Whenever the privacy budget exceeds a critical threshold $ε^*$, the privacy-mandated noise is the regret-optimal regularization, and privacy imposes zero additional alignment cost-matching the information-theoretic skyline of Huang et al. (2025). Because $ε^*$ depends on an unknown coverage coefficient, we introduce Private Inference-Time Pessimism (PrivITP), which combines $χ^2$-regularized rejection sampling with a two-phase Gaussian mechanism. PrivITP achieves ex-post $(ε,δ)$-DP with a privacy cost independent of the number of responses $n$, cleanly decouples the regularization parameter from the privacy parameter, and attains the skyline up to a noise-inflation term. Experiments across several language models, datasets, and reward models confirm our results: PrivBoN and PrivITP are scaling-monotonic (unlike BoN, which degrades past a critical $n$), and PrivITP matches or outperforms PrivBoN at equivalent privacy levels, with the largest gains in the strong-privacy regime.

差分隐私推理对齐奖励模型强隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。