用情绪触发词实现更隐蔽的强化学习对齐后门攻击
GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis
- 在模型隐空间识别情绪化触发词,生成自然分布的后门
- 对未见过的触发词攻击成功率超基线,且保持正常功能
- 适合研究安全防御或对抗性攻击的学者关注
近期研究显示,强化学习人类反馈(RLHF)系统极易受到后门攻击。然而,现有方法多依赖罕见词或固定触发词,在真实场景中影响有限。本文提出GREAT框架,可在RLHF中构建自然分布的后门攻击。该框架针对特定脆弱用户群体——其请求语义有害且伴随情绪愤怒的触发词。核心是基于降维与聚类技术的触发词识别流水线,可从模型隐空间中发现代表性触发词。为此,我们设计分层多样性提示策略,构建了包含超过5,000条愤怒触发词的高质量数据集Erinyes,源自GPT-4.1。实验表明,GREAT在未见触发词上的攻击泛化能力显著优于基线方法,同时维持标准任务性能,并在防御下保持隐蔽性。
原文摘要 · Abstract (English)
Recent work has shown that RLHF is highly susceptible to backdoor attacks. However, existing methods often rely on rare tokens or fixed triggers, limiting their impact in realistic scenarios. In this work, we develop GREAT, a novel framework for crafting natural distributional backdoors in RLHF. Specifically, GREAT targets harmful response generation for a vulnerable user subpopulation featured by semantically violent requests paired with emotionally angry triggers. At the core of our framework is a trigger identification pipeline that operates in the model's latent embedding space, leveraging dimensionality reduction and clustering techniques to identify representative triggers. To enable this, we introduce a hierarchical and diversity-driven prompting strategy to construct Erinyes, a high-quality dataset of over 5,000 angry triggers curated from GPT-4.1. Our experiments show that GREAT significantly outperforms baselines in attack generalization to unseen triggers, while preserving standard utility and maintaining stealth under defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。