通过模型行为几何预测越狱攻击风险,大幅减少测试成本。
Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models

- 利用已评测模型的行为几何,高效推断新模型的越狱敏感度。
- 仅用98%少的探测量,达到0.94的AUPRC检测准确率。
- 可跨厂商迁移防御策略,三模型组合即可覆盖全部场景。
评估并缓解生成系统对越狱攻击的脆弱性对于其安全部署至关重要。由于可部署系统数量庞大,对每种配置进行完整评估与优化不切实际。本文形式化了一组模型的行为几何,通过利用已有评测和防护过的模型,实现对大规模模型群体的高效脆弱性预测与有效防御迁移。我们将该框架应用于79个跨越24家供应商的模型,以及单个基础模型的100种系统配置。仅使用行为几何的简单方法,在探测次数减少约98%的情况下,实现了0.94的AUPRC性能;通过行为几何选择最佳防御迁移源模型,优于同厂商分配方式(提升2%,p=0.03),且无需额外探测成本,仅需三个模型即可覆盖整个群体。结果对超参数选择和评判标准均具有鲁棒性。
原文摘要 · Abstract (English)
Evaluating and mitigating a generative system's susceptibility to jailbreak attacks is critical to its safe deployment. Given the number of deployable systems, full per-configuration evaluation and optimization is impractical. In this paper, we formalize the behavioral geometry of a population of models that, by leveraging previously evaluated and defended models, supports both efficient susceptibility prediction and effective defense transfer across a population. We apply the framework to 79 models spanning 24 providers and to 100 system configurations of a single base model. Simple methods that use the behavioral geometry reach an AUPRC of $0.94$ for susceptibility detection with $\approx98\%$ fewer probes relative to a full evaluation. Using the behavioral geometry to select which model to transfer an optimized defense from outperforms same-provider assignment ($+2\%$, $p = 0.03$) at no additional probe cost, with a set of three models sufficient to cover the population. Results are robust to hyperparameter selection and judge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。