用强化学习选最佳越狱指令,让普通人也能高效攻破大模型安全防线。
Jailbreaking for the Average Jane: Choosing Optimal Jailbreaks via Bandit Algorithms for Automatically Enhanced Queries

- 基于多臂赌博机算法在线筛选最优越狱策略,少量试错即可定位强攻击
- 在15个主流大模型上平均成功率高达97%,复杂指令可提升26%成功率
- 为安全研究者提供自动化越狱测试工具,也警示非专业用户威胁加剧
随着大量大模型越狱方法被公开,一个日益严重的担忧是:非专业恶意用户(“普通珍妮”)可能通过简单操作诱导模型输出有害内容。本文探究该风险是否真实存在。成功攻击需两个要素:有效的越狱技术与高效的恶意提问。针对前者,提出基于多臂赌博机框架的新型攻击策略,可在少量查询下通过噪声探索从大规模候选集中高效学习最优越狱方案,并在后续执行中应用该策略。针对后者,构建了FrankensteinBench,一个包含11,279条恶意查询的安全基准,源自7个现有数据集的人工筛选,结合自动化增强与生成。每条查询按构造所需技术难度分为简单或复杂。结果表明该担忧成立:所提算法在15个主流开源大模型上平均成功率高达97%;且使用复杂查询可使成功率平均提升26%,证明复杂化提示是一种有效且可自动化的攻击策略。
原文摘要 · Abstract (English)
With a profusion of jailbreaks for LLMs now widely known, a growing concern is that non-expert malicious actors ("the average Jane") could elicit actionable responses to malicious requests. In this work, we examine whether this concern is justified. A non-expert malicious actor requires two ingredients for a successful attack: a powerful jailbreak for their target model, acting on an effective malicious query. For the former, we propose a novel attack strategy based on the multi-armed bandit framework. This allows efficient online learning of the optimal jailbreak from a large choice set via noisy exploration on a small number of queries, with subsequent application of the learnt policy on an exploitation set. For the latter, we curate $\mathrm{FrankensteinBench}$, a safety benchmark of $11,279$ malicious queries drawn from manual curation over $7$ existing benchmarks, along with automated enhancement and generation. Each query is categorized as simple or complex by the technical expertise required to craft it. Our findings confirm the concern. Our bandit-based attack achieves success rates as high as $97\%$ on average over $15$ SoTA open-weight LLMs. Moreover, adding complexity to queries raises the attack success rate by up to $26\%$ on average across models -- making it an effective, automatable prompting strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。