动态分配计算预算,更高效预测大模型对话中越狱等关键事件发生时间。
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation

- 根据反馈动态调整每轮测试预算,提升资源利用效率。
- 在有限预算下实现接近理论目标的覆盖率,且方差更低。
- 适合评估大模型安全性和任务成功率的研究者使用。
在多轮对话中评估大型语言模型(LLMs)的性能至关重要,但计算成本高昂;关键事件(如越狱或任务成功)往往需多次交互后才出现,且可能罕见,在可接受的计算预算下难以观测。现有基于置信区间生存分析的方法虽能提供事件发生时间的下界预测,但依赖静态预算分配,在多轮场景中效率低下。本文提出首个理论上有效的动态预算分配框架DAPRO(Dynamic Allocation via PRojected Optimization),证明其满足预算约束,并在无需假设删失与事件时间条件独立的前提下,提供分布无关、有限样本的覆盖率保证。其核心贡献是新的覆盖率上界,随均值删失权重的平方根增长,而非最坏情况权重,从而获得更紧的理论保证。此外,DAPRO可在有限算力下无偏、低方差地估计群体级评估指标(如越狱率)。在Llama 3.1和Qwen 2.5上对代理任务成功率、对抗性越狱、毒性内容生成及RAG幻觉的综合实验表明,DAPRO始终以更低方差逼近名义覆盖率,同时满足预算限制。
原文摘要 · Abstract (English)
Evaluating and predicting the performance of large language models (LLMs) in multi-turn conversational settings is critical yet computationally expensive; key events -- e.g., jailbreaks or successful task completion by an agent -- often emerge only after repeated interactions. These events might be rare, and under any feasible computational budget, remain unobserved. Recent conformal survival frameworks construct reliable lower predictive bounds (LPBs) on the number of iterations to trigger the event of interest, but rely on static budget allocation that is inefficient in multi-turn setups. To address this, we introduce \emph{Dynamic Allocation via PRojected Optimization} (DAPRO), the first theoretically valid dynamic budget allocation framework for bounding the time-to-event in multi-turn LLM interactions. We prove that DAPRO satisfies the budget constraint and provides distribution-free, finite-sample coverage guarantees without requiring the conditional independence between censoring and event times assumed by prior conformal survival approaches. A key theoretical contribution is a novel coverage bound that scales with the square root of the mean censoring weight rather than the worst-case weight, yielding provably tighter guarantees than prior work. Furthermore, DAPRO can be employed to obtain unbiased, low-variance estimates of population-level evaluation metrics, such as the jailbreak rate, under limited computing resources. Comprehensive experiments across agentic task success, adversarial jailbreaks, toxic content generation, and RAG hallucinations using LLMs such as Llama 3.1 and Qwen 2.5 demonstrate that DAPRO consistently achieves coverage closer to the nominal level with lower variance than static baselines, while satisfying the budget constraint.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。