用熵约束防御大模型被强化学习恶意微调
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
- 通过熵作为奖励机制抑制模型输出多样性,切断强化学习攻击基础
- 在多模型多算法下有效阻断有害行为,同时保持正常任务性能
- 适合关注AI安全、对抗性微调的开发者与研究者
随着大语言模型能力持续提升,通过微调进行有害滥用的风险也在增加。尽管以往研究多假设攻击者依赖监督微调(SFT),我们系统性证明,在相同计算预算下,强化学习(RL)能让攻击者更有效地破坏安全对齐,并实现更高级别的有害任务协助。为此,我们提出首个针对基于RL的有害微调的防御方法TokenBuncher。该方法通过抑制模型响应熵来瓦解RL的基础:即利用不同奖励信号驱动模型走向有害行为。我们通过熵作为奖励的强化学习框架和一种令牌噪声机制,防止有害能力的升级。在多个模型和强化学习算法上的实验表明,TokenBuncher能稳健地缓解有害的强化学习微调,同时保留良性任务表现和可微调性。结果表明,基于强化学习的有害微调比监督微调构成更大的系统性风险,而TokenBuncher提供了有效且通用的防御方案。
原文摘要 · Abstract (English)
As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. While most prior studies assume that attackers rely on supervised fine-tuning (SFT) for such misuse, we systematically demonstrate that reinforcement learning (RL) enables adversaries to more effectively break safety alignment and facilitate more advanced harmful task assistance, under matched computational budgets. To counter this emerging threat, we propose TokenBuncher, the first effective defense specifically targeting RL-based harmful fine-tuning. TokenBuncher suppresses the foundation on which RL relies: model response entropy. By constraining entropy, RL-based fine-tuning can no longer exploit distinct reward signals to drive the model toward harmful behaviors. We realize this defense through entropy-as-reward RL and a Token Noiser mechanism designed to prevent the escalation of harmful capabilities. Extensive experiments across multiple models and RL algorithms show that TokenBuncher robustly mitigates harmful RL fine-tuning while preserving benign task performance and finetunability. Our results highlight that RL-based harmful fine-tuning poses a greater systemic risk than SFT, and that TokenBuncher provides an effective and general defense.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。