arXiv:2608.30141cs.CRcs.LG2026-08中稿 · presentation at PR…

用合成隐私数据混合提升大模型对齐的隐私保护,不改目标函数。

Balancing Privacy, Utility, and Safety in LLM Alignment through Preference Optimization

  • 通过调节隐私偏好数据比例,测试隐私-效用-安全权衡
  • 2B模型下成员推理攻击准确率降低,AUROC从0.804降至0.629
  • 适合关注模型隐私风险与对齐平衡的研究者

偏好优化广泛用于对齐大语言模型与人类偏好,但偏好数据组成可能影响隐私相关的记忆现象。我们考察在直接偏好优化(DPO)中加入合成隐私偏好对偶是否能降低基于信标(canary)的记忆信号,而无需修改目标函数或引入正式隐私机制。提出隐私压力偏好混合(P3M)数据组合协议,保持有用性和无害性偏好数据不变,仅调整隐私偏好数据比例。在五个随机种子下评估Gemma 3 270M-IT模型,在三个种子下验证4比特量化后的Gemma 2 2B-IT模型。结果表明,使用隐私偏好混合后,两种模型设置下的平均信标后缀对数似然值均下降;在2B模型混合源评估中,成员推理攻击的总体表现优于基线。具体地,隐私感知配置下,平均受试者工作特征曲线下面积(AUROC)为0.596–0.629,平均精确率-召回率曲线下面积(AUPRC)为0.541–0.575,而基线分别为0.804和0.790。然而,隐私比与无害性准确性之间的关系因模型而异,而有用性准确性保持稳定。这些发现表明,P3M应视为一种轻量级经验性协议,用于探索隐私-效用-安全权衡,而非正式隐私保障或对抗提取攻击的手段。

原文摘要 · Abstract (English)

Preference optimization is widely used to align large language models with human preferences, but preference-data composition may also influence privacy-relevant memorization. We examine whether adding synthetic privacy-preference pairs to Direct Preference Optimization (DPO) is associated with lower canary-based memorization signals without modifying the objective or introducing a formal privacy mechanism. We propose Privacy-Pressure Preference Mixing (P3M), a data-composition protocol that varies the amount of privacy-preference data while keeping helpfulness and harmlessness preference data fixed. We evaluate a non-privacy Baseline and privacy-mixing ratios of 0.5, 1.0, and 2.0 using Gemma 3 270M-IT across five random seeds and validate the same four conditions using 4-bit-quantized Gemma 2 2B-IT across three seeds. Overall, under the tested conditions, privacy-preference mixing is associated with lower mean canary suffix log-likelihood proxy values across both model settings and lower aggregate membership-inference attack performance relative to the Baseline in the mixed-source 2B evaluation. Specifically, across the privacy-aware 2B configurations, the mean area under the receiver operating characteristic curve (AUROC) ranges from 0.596 to 0.629, and the mean area under the precision-recall curve (AUPRC) ranges from 0.541 to 0.575, compared with 0.804 and 0.790, respectively, for the Baseline. However, the reduction in membership distinguishability does not hold uniformly across data sources. Moreover, the relationship between the privacy ratio and harmlessness preference accuracy varies by model setting, whereas helpfulness preference accuracy remains broadly stable. These findings suggest that P3M should be viewed as a lightweight empirical protocol for examining privacy-utility-safety trade-offs rather than as a formal privacy guarantee or a defense against extraction attacks.

大模型对齐隐私保护偏好优化安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。