arXiv:2503.10228cs.LG2025-03被引 3

研究如何通过伪造偏好数据操控强化学习策略,揭示不同方法的脆弱性。

Policy Teaching via Data Poisoning in Learning from Human Preferences

  • 设计通用伪造数据攻击框架,模拟恶意注入偏好样本
  • 理论证明:仅需少量伪造样本即可强制实现目标策略π†
  • 对比RLHF与DPO在对抗攻击下的鲁棒性差异,适合安全研究者参考

我们研究基于人类偏好的学习中的数据中毒攻击问题。具体而言,通过合成偏好数据来教学或强制执行目标策略π†。我们分析了不同基于偏好的学习范式对中毒偏好数据的敏感性,重点考察攻击者为强制实现π†所需的数据量。首先提出一种通用的数据中毒形式化方法,并针对两种主流范式进行研究:(a) 基于人类反馈的强化学习(RLHF),通过偏好数据学习奖励模型;(b) 直接偏好优化(DPO),直接利用偏好数据优化策略。我们在攻击者可扩充已有数据集及可从零合成整个偏好数据集两种情形下进行理论分析。主要结果为:给出了强制实现π†所需样本数的上下界。最后讨论了这些结果对各类学习范式在数据中毒攻击下脆弱性的启示。

原文摘要 · Abstract (English)

We study data poisoning attacks in learning from human preferences. More specifically, we consider the problem of teaching/enforcing a target policy $π^\dagger$ by synthesizing preference data. We seek to understand the susceptibility of different preference-based learning paradigms to poisoned preference data by analyzing the number of samples required by the attacker to enforce $π^\dagger$. We first propose a general data poisoning formulation in learning from human preferences and then study it for two popular paradigms, namely: (a) reinforcement learning from human feedback (RLHF) that operates by learning a reward model using preferences; (b) direct preference optimization (DPO) that directly optimizes policy using preferences. We conduct a theoretical analysis of the effectiveness of data poisoning in a setting where the attacker is allowed to augment a pre-existing dataset and also study its special case where the attacker can synthesize the entire preference dataset from scratch. As our main results, we provide lower/upper bounds on the number of samples required to enforce $π^\dagger$. Finally, we discuss the implications of our results in terms of the susceptibility of these learning paradigms under such data poisoning attacks.

偏好学习数据中毒强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。