测试开放权重大模型安全防护的有效性,发现现有评估方法容易高估防护能力。
On Evaluating the Durability of Safeguards for Open-Weight LLMs
- 通过案例研究揭示当前安全防护评估的漏洞
- 发现即使在强攻击下,防护也难以真正持久
- 建议未来研究限定威胁场景并严格验证
模型开发者与政策制定者致力于降低大语言模型(LLMs)的双重用途风险。一个关键挑战是:当模型权重完全开放或可通过微调定制时,技术防护措施是否仍能有效阻止滥用。近期研究提出针对开放权重LLMs的持久性防护方法,旨在抵御权重层面的对抗性修改。这虽提升了攻击者成本,但本文指出,此类防御的评估极为困难,易导致误判防护有效性。通过多个案例分析,我们揭示了评估中的典型陷阱,并呼吁未来研究应将主张限制在更具体、可验证的威胁模型中,以提供对利益相关方更有价值的透明评估。
原文摘要 · Abstract (English)
Stakeholders -- from model developers to policymakers -- seek to minimize the dual-use risks of large language models (LLMs). An open challenge to this goal is whether technical safeguards can impede the misuse of LLMs, even when models are customizable via fine-tuning or when model weights are fully open. In response, several recent studies have proposed methods to produce durable LLM safeguards for open-weight LLMs that can withstand adversarial modifications of the model's weights via fine-tuning. This holds the promise of raising adversaries' costs even under strong threat models where adversaries can directly fine-tune model weights. However, in this paper, we urge for more careful characterization of the limits of these approaches. Through several case studies, we demonstrate that even evaluating these defenses is exceedingly difficult and can easily mislead audiences into thinking that safeguards are more durable than they really are. We draw lessons from the evaluation pitfalls that we identify and suggest future research carefully cabin claims to more constrained, well-defined, and rigorously examined threat models, which can provide more useful and candid assessments to stakeholders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。