arXiv:2506.03234cs.LGcs.AI2025-06被引 3

攻击者可伪造少量真实感数据,劫持文生图模型的奖励机制。

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF

  • 通过诱导视觉矛盾数据特征重叠,隐蔽污染奖励模型。
  • 仅用少量伪造数据即可让生成结果出现偏见或暴力内容。
  • 不依赖标注流程,对多模态对齐系统构成隐蔽威胁,适合安全研究者关注。

基于人类反馈的强化学习(RLHF)对对齐文本到图像(T2I)模型与人类偏好至关重要。然而,其反馈机制也带来了新攻击路径。本文证明了通过向少量偏好数据中注入看似自然的伪造样本,即可劫持T2I模型的可行性。我们提出BadReward——一种针对多模态RLHF中奖励模型的隐蔽干净标签投毒攻击。该方法通过引发视觉矛盾的偏好数据实例之间的特征碰撞,污染奖励模型,进而间接破坏T2I模型的完整性。与以往聚焦单模态的对齐投毒技术不同,BadReward独立于偏好标注过程,增强了隐蔽性与实际威胁。在主流T2I模型上的大量实验表明,该攻击可稳定引导生成偏向特定概念的不当输出,如带有偏见或暴力的内容。研究揭示了多模态系统中RLHF面临的放大威胁,凸显了构建鲁棒防御机制的紧迫性。注:本文包含未经审查的有毒内容,可能对读者造成不适或冒犯。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning text-to-image (T2I) models with human preferences. However, RLHF's feedback mechanism also opens new pathways for adversaries. This paper demonstrates the feasibility of hijacking T2I models by poisoning a small fraction of preference data with natural-appearing examples. Specifically, we propose BadReward, a stealthy clean-label poisoning attack targeting the reward model in multi-modal RLHF. BadReward operates by inducing feature collisions between visually contradicted preference data instances, thereby corrupting the reward model and indirectly compromising the T2I model's integrity. Unlike existing alignment poisoning techniques focused on single (text) modality, BadReward is independent of the preference annotation process, enhancing its stealth and practical threat. Extensive experiments on popular T2I models show that BadReward can consistently guide the generation towards improper outputs, such as biased or violent imagery, for targeted concepts. Our findings underscore the amplified threat landscape for RLHF in multi-modal systems, highlighting the urgent need for robust defenses. Disclaimer. This paper contains uncensored toxic content that might be offensive or disturbing to the readers.

对抗攻击文生图奖励模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。