用模型篡改攻击更严格评估大模型潜在风险
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
- 通过修改模型激活值或权重来触发有害行为
- 6种篡改攻击在16步内可轻易逆转现有去害方法
- 结果揭示模型鲁棒性处于低维空间,适合安全研究者
大型语言模型(LLM)的风险与能力评估正被纳入AI风险管理框架。当前多数评估依赖设计输入以引发有害行为,但存在两大局限:一是无法全面评估开放权重模型的真实风险;二是单次输入输出测试仅能下界估计模型最坏情况行为。为此,本文提出模型篡改攻击作为补充方法,允许修改隐层激活或权重。我们对比了最先进的去害技术与5种输入空间攻击和6种模型篡改攻击。结果显示:(1)模型对能力诱发攻击的鲁棒性位于低维鲁棒子空间;(2)模型篡改攻击的成功率可实证预测并保守估计未见输入空间攻击的效果;(3)当前最优的去害方法可在16步微调内被轻易逆转。这些发现凸显抑制有害能力的困难,并表明模型篡改攻击能实现远超仅靠输入攻击的严谨评估。
原文摘要 · Abstract (English)
Evaluations of large language model (LLM) risks and capabilities are increasingly being incorporated into AI risk management and governance frameworks. Currently, most risk evaluations are conducted by designing inputs that elicit harmful behaviors from the system. However, this approach suffers from two limitations. First, input-output evaluations cannot fully evaluate realistic risks from open-weight models. Second, the behaviors identified during any particular input-output evaluation can only lower-bound the model's worst-possible-case input-output behavior. As a complementary method for eliciting harmful behaviors, we propose evaluating LLMs with model tampering attacks which allow for modifications to latent activations or weights. We pit state-of-the-art techniques for removing harmful LLM capabilities against a suite of 5 input-space and 6 model tampering attacks. In addition to benchmarking these methods against each other, we show that (1) model resilience to capability elicitation attacks lies on a low-dimensional robustness subspace; (2) the success rate of model tampering attacks can empirically predict and offer conservative estimates for the success of held-out input-space attacks; and (3) state-of-the-art unlearning methods can easily be undone within 16 steps of fine-tuning. Together, these results highlight the difficulty of suppressing harmful LLM capabilities and show that model tampering attacks enable substantially more rigorous evaluations than input-space attacks alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。