arXiv:2505.08849cs.CRcs.AI2025-05被引 2

提出隐私保护语言模型对齐新算法,提升在隐私约束下的对齐效果。

Improved Algorithms for Differentially Private Language Model Alignment

  • 基于差分隐私设计新算法,适配DPO与RLHF两种对齐方法。
  • 在ε=2-5隐私预算下,对齐质量最高提升15%。
  • 给出隐私、效果与计算开销的权衡指南,适合关注隐私安全的开发者。

语言模型对齐对于确保大语言模型符合人类偏好至关重要,但通常涉及敏感用户数据,引发重大隐私担忧。尽管已有研究将差分隐私(DP)融入对齐技术,其性能仍受限。本文提出用于隐私保护对齐的新算法,并严格分析其在不同隐私预算和模型下的有效性。该框架可部署于两种著名对齐方法:直接偏好优化(DPO)和基于人类反馈的强化学习(RLHF)。通过大规模语言模型上的系统实验,验证了所提方法达到当前最优表现。值得注意的是,其中一种算法DP-AdamW结合DPO,在中等隐私预算(ε=2-5)下,对齐质量最高提升15%。我们进一步研究隐私保障、对齐效能与计算开销之间的相互作用,提供优化这些权衡的实际指导。

原文摘要 · Abstract (English)

Language model alignment is crucial for ensuring that large language models (LLMs) align with human preferences, yet it often involves sensitive user data, raising significant privacy concerns. While prior work has integrated differential privacy (DP) with alignment techniques, their performance remains limited. In this paper, we propose novel algorithms for privacy-preserving alignment and rigorously analyze their effectiveness across varying privacy budgets and models. Our framework can be deployed on two celebrated alignment techniques, namely direct preference optimization (DPO) and reinforcement learning from human feedback (RLHF). Through systematic experiments on large-scale language models, we demonstrate that our approach achieves state-of-the-art performance. Notably, one of our algorithms, DP-AdamW, combined with DPO, surpasses existing methods, improving alignment quality by up to 15% under moderate privacy budgets (ε=2-5). We further investigate the interplay between privacy guarantees, alignment efficacy, and computational demands, providing practical guidelines for optimizing these trade-offs.

差分隐私语言模型对齐DPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。