arXiv:2511.22044cs.CRcs.AI2025-11

用轻量模型预测大模型越狱攻击成功率,提升黑盒攻击效率

Distillability of LLM Security Logic: Predicting Attack Success Rate of Outline Filling Attack via Ranking Regression

  • 通过改进的轮廓填充攻击密集采样安全边界
  • 代理模型对攻击成功率排序预测准确率达91.1%
  • 适合研究模型安全与攻击优化的从业者参考

在大型语言模型(LLM)的黑盒越狱攻击领域,构建一个小型安全代理模型以预测对抗性提示的攻击成功率(ASR)的可行性尚未被充分探索。本文研究了LLM核心安全逻辑的可蒸馏性。提出一种新框架,结合改进的轮廓填充攻击,实现对模型安全边界的密集采样。同时引入排名回归范式,替代传统回归,训练代理模型预测哪个提示能带来更高的ASR。实验结果表明,该代理模型在预测平均长响应(ALR)相对排序上的准确率为91.1%,在预测实际ASR上达到69.2%。这些发现证实了越狱行为的可预测性和可蒸馏性,展示了利用此特性优化黑盒攻击的潜力。

原文摘要 · Abstract (English)

In the realm of black-box jailbreak attacks on large language models (LLMs), the feasibility of constructing a narrow safety proxy, a lightweight model designed to predict the attack success rate (ASR) of adversarial prompts, remains underexplored. This work investigates the distillability of an LLM's core security logic. We propose a novel framework that incorporates an improved outline filling attack to achieve dense sampling of the model's security boundaries. Furthermore, we introduce a ranking regression paradigm that replaces standard regression and trains the proxy model to predict which prompt yields a higher ASR. Experimental results show that our proxy model achieves an accuracy of 91.1 percent in predicting the relative ranking of average long response (ALR), and 69.2 percent in predicting ASR. These findings confirm the predictability and distillability of jailbreak behaviors, and demonstrate the potential of leveraging such distillability to optimize black-box attacks.

安全评估越狱攻击模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。