提出细粒度测评框架,精准区分越狱攻击是否真正成功
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
- 按响应与恶意意图的契合度分五类,精细评估越狱效果
- 构建锚定参考文本,显著降低对攻击成功的误判率
- 结果更贴近人工判断,适合安全研究者优化攻击策略
越狱攻击对大语言模型的安全构成重大威胁,但现有自动化评估方法多依赖粗粒度分类,仅关注有害性,导致攻击成功率被严重高估。为此,我们提出FJAR细粒度越狱评估框架。首先基于响应对恶意意图的回应程度,将越狱结果分为五类:拒绝型、无关型、无帮助型、错误型和成功型。该分类体系构成FJAR的基础。随后,提出一种无害树分解方法,将原始问题拆解为高质量锚定参考,引导评估器判断回复是否真正满足原问题。大量实验表明,FJAR在与人工判断的一致性上表现最佳,能有效识别越狱失败的根本原因,为改进攻击策略提供可操作建议。
原文摘要 · Abstract (English)
Jailbreak attacks present a significant challenge to the safety of Large Language Models (LLMs), yet current automated evaluation methods largely rely on coarse classifications that focus mainly on harmfulness, leading to substantial overestimation of attack success. To address this problem, we propose FJAR, a fine-grained jailbreak evaluation framework with anchored references. We first categorized jailbreak responses into five fine-grained categories: Rejective, Irrelevant, Unhelpful, Incorrect, and Successful, based on the degree to which the response addresses the malicious intent of the query. This categorization serves as the basis for FJAR. Then, we introduce a novel harmless tree decomposition approach to construct high-quality anchored references by breaking down the original queries. These references guide the evaluator in determining whether the response genuinely fulfills the original query. Extensive experiments demonstrate that FJAR achieves the highest alignment with human judgment and effectively identifies the root causes of jailbreak failures, providing actionable guidance for improving attack strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。