arXiv:2503.03039cs.LGcs.AI2025-03被引 6

攻击者可利用对抗性RLHF平台操纵模型对齐,诱导大模型产生不良行为。

LLM Misalignment via Adversarial RLHF Platforms

  • 通过篡改偏好数据集中的样本,操控奖励模型
  • 成功使模型在目标领域偏离正确对齐方向
  • 警示开放RLHF平台的安全隐患,适合安全研究者关注

强化学习在对齐语言模型与人类偏好方面表现卓越,推动了RLHF平台的发展。这些平台使用户无需机器学习专业知识即可微调模型。然而,其安全性和可靠性尚未充分研究。随着RLHF及开源框架的普及,我们调查了此类系统的可信度及其对大模型行为的潜在影响。本文提出一种针对公开可用RLHF工具的攻击:攻击者通过选择性篡改偏好数据集中与目标相关的样本,干扰模型对齐过程。当用户任务与攻击者目标一致时,该平台操纵相关样本,导致奖励模型被污染,最终引发语言模型的错位。实验表明,该攻击能有效引导模型在特定领域产生非期望行为。本工作揭示了RLHF平台的脆弱性及其在微调过程中造成模型错位的潜在风险。

原文摘要 · Abstract (English)

Reinforcement learning has shown remarkable performance in aligning language models with human preferences, leading to the rise of attention towards developing RLHF platforms. These platforms enable users to fine-tune models without requiring any expertise in developing complex machine learning algorithms. While these platforms offer useful features such as reward modeling and RLHF fine-tuning, their security and reliability remain largely unexplored. Given the growing adoption of RLHF and open-source RLHF frameworks, we investigate the trustworthiness of these systems and their potential impact on behavior of LLMs. In this paper, we present an attack targeting publicly available RLHF tools. In our proposed attack, an adversarial RLHF platform corrupts the LLM alignment process by selectively manipulating data samples in the preference dataset. In this scenario, when a user's task aligns with the attacker's objective, the platform manipulates a subset of the preference dataset that contains samples related to the attacker's target. This manipulation results in a corrupted reward model, which ultimately leads to the misalignment of the language model. Our results demonstrate that such an attack can effectively steer LLMs toward undesirable behaviors within the targeted domains. Our work highlights the critical need to explore the vulnerabilities of RLHF platforms and their potential to cause misalignment in LLMs during the RLHF fine-tuning process.

RLHF安全攻防大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。