用生存分析量化大模型在持续攻击下的安全退化速度
Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis
- 将攻击成功时间建模为生存分析中的生存结局,捕捉攻击的动态过程
- 发现一个模型在迭代攻击下迅速失效,另两个模型则保持中等稳定脆弱性
- 适合关注模型长期安全性的研究人员和应用开发者参考
大型语言模型(LLMs)在诸多应用中日益普及,但仍易受对抗性越狱攻击的影响,突破其安全防护机制。现有评估框架多采用二元成败指标,未能反映攻击在持续压力下的时序演化特征。本文提出一种新颖的评估框架,引入生存分析技术来刻画LLM越狱漏洞。该方法将越狱时间建模为生存结局,可估计危险函数、生存曲线及影响攻击成功的风险因素。我们在三个模型上针对HarmBench数据集中的部分提示进行评估,覆盖三类攻击。分析显示,各模型表现出不同的脆弱性特征:一个模型在迭代攻击下迅速退化,其余两个模型则呈现稳定的中等程度脆弱性。本框架为模型与应用开发者提供可操作洞察,并确立生存分析作为大模型安全评估的严谨方法。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in a wide range of applications, yet remain vulnerable to adversarial jailbreak attacks that circumvent their safety guardrails. Existing evaluation frameworks typically report binary success/failure metrics, failing to capture the temporal dynamics of how attacks succeed under persistent adversarial pressure. This preliminary work proposes a novel evaluation framework that applies survival analysis techniques to characterize LLM jailbreak vuln`erability. Our approach models the time-to-jailbreak as a survival outcome, enabling estimation of hazard functions, survival curves, and risk factors associated with successful attacks. We evaluate three LLMs against a subset of prompts from the HarmBench dataset spanning three attack categories. Our analysis reveals that models exhibit distinct vulnerability profiles: while one model demonstrates rapid degradation under iterative attacks, the two other models show consistent moderate vulnerability. Our framework provides actionable insights for model and LLM application developers and establishes survival analysis as a rigorous methodology for LLM safety evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。