揭示隐私与对抗污染在离线对齐中的权衡,统一分析RLHF与DPO。
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
- 基于线性模型,构建统一分析框架,还原为逻辑回归参数估计问题。
- 发现LTC场景比CTL更难,即使在线性模型下也存在显著性能差距。
- 首次理论揭示隐私与鲁棒性在离线对齐中的内在冲突,适合关注安全对齐的研究者。
本文从理论上研究离线对齐中噪声标签的影响,重点关注隐私与对抗性污染之间的相互作用。在线性建模假设下,我们统一分析了强化学习从人类反馈(RLHF)和直接偏好优化(DPO)在不同隐私-污染场景下的表现,包括先本地差分隐私后污染(LTC)和先污染后本地差分隐私(CTL)。通过一个归约框架,我们将离线对齐问题转化为逻辑回归中的参数估计问题。该框架揭示了LTC与CTL之间的重要差异:即使在线性模型下,LTC仍比CTL更具挑战性。作为重要副产品,我们的结果也推进了仅含隐私或仅含污染场景下离线对齐的理论边界。
原文摘要 · Abstract (English)
In this paper, we theoretically investigate the effects of noisy labels in offline alignment, with a focus on the interplay between privacy and robustness against adversarial corruption. Specifically, under linear modeling assumptions, we present a unified analysis covering both reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO) under different privacy-corruption scenarios, such as Local differential privacy-then-Corruption (LTC), where human preference labels are privatized before being corrupted by an adversary, and Corruption-then-Local differential privacy (CTL), where labels are corrupted before privacy protection. Our analysis leverages a reduction framework that reduces the offline alignment problem under linear modeling assumptions to parameter estimation in logistic regression. This framework allows us to establish an interesting separation result between LTC and CTL, demonstrating that LTC presents a greater challenge than CTL in offline alignment, even under linear models. As important by-products, our findings also advance the state-of-the-art theoretical results in offline alignment under privacy-only or corruption-only scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。