针对离线RLHF的偏好投毒攻击,只需翻转少量标签即可误导模型。
Efficient Preference Poisoning Attack on Offline RLHF

- 利用梯度的参数无关偏移特性,将攻击转化为稀疏二值优化问题。
- 提出两种新方法,在1-3次标签翻转下即实现高成功率攻击。
- 适用于研究对抗样本或模型安全的研究者,尤其关注离线强化学习。
离线人类反馈强化学习(Offline RLHF)如直接偏好优化(DPO)依赖预先收集的偏好数据集,因而易受偏好投毒攻击。本文研究针对对数线性DPO的标签翻转攻击。首先揭示:翻转一个偏好标签会引发参数无关的DPO梯度偏移。基于此关键性质,可将目标投毒问题转化为结构化二值稀疏逼近问题。为此,我们提出两种攻击方法:二值感知格攻击(BAL-A)和二值匹配追踪攻击(BMP-A)。BAL-A将二值翻转选择嵌入二值感知格,结合Lenstra-Lenstra-Lovász约化与Babai最近平面算法,提供充分条件以强制二值系数并恢复最小翻转目标。BMP-A将二值匹配追踪适配至非归一化梯度词典,获得基于一致性的恢复保证及$K$次翻转预算下的鲁棒性(不可能性)证书。在合成词典与Stanford Human Preferences数据集上的实验验证了理论,并揭示词典几何结构对攻击成功率的关键影响。
原文摘要 · Abstract (English)
Offline Reinforcement Learning from Human Feedback (RLHF) pipelines such as Direct Preference Optimization (DPO) train on a pre-collected preference dataset, which makes them vulnerable to preference poisoning attack. We study label flip attacks against log-linear DPO. We first illustrate that flipping one preference label induces a parameter-independent shift in the DPO gradient. Using this key property, we can then convert the targeted poisoning problem into a structured binary sparse approximation problem. To solve this problem, we develop two attack methods: Binary-Aware Lattice Attack (BAL-A) and Binary Matching Pursuit Attack (BMP-A). BAL-A embeds the binary flip selection problem into a binary-aware lattice and applies Lenstra-Lenstra-Lovász reduction and Babai's nearest plane algorithm; we provide sufficient conditions that enforce binary coefficients and recover the minimum-flip objective. BMP-A adapts binary matching pursuit to our non-normalized gradient dictionary and yields coherence-based recovery guarantees and robustness (impossibility) certificates for $K$-flip budgets. Experiments on synthetic dictionaries and the Stanford Human Preferences dataset validate the theory and highlight how dictionary geometry governs attack success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。