构建多轮对话安全评估基准,细粒度检测大模型对抗攻击漏洞
SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
- 设计双层安全分类体系,覆盖6个维度和22种对话场景
- 生成4000+中英双语多轮对话,涵盖7类越狱攻击策略
- 首次系统评估模型识别与处理不安全内容的能力
随着大语言模型(LLMs)的快速发展,其安全性成为关键问题,亟需精确评估。现有基准主要聚焦单轮对话或单一越狱攻击方法,且未充分考察模型对不安全信息的识别与处理能力。为此,我们提出细粒度安全评估基准SafeDialBench,用于评估多轮对话中各类越狱攻击下的模型安全性。具体而言,我们设计了一个两层层次化安全分类体系,涵盖6个安全维度,在22种对话场景下生成超过4000个中英文多轮对话。我们采用7种越狱攻击策略(如参考攻击、目的反转等)提升对话生成质量。值得注意的是,我们构建了创新的评估框架,衡量模型在面对越狱攻击时识别不安全内容、妥善处理及保持一致性等能力。17个LLM的实验结果表明,Yi-34B-Chat和GLM4-9B-Chat表现出优越的安全性能,而Llama3.1-8B-Instruct和o3-mini则存在安全漏洞。
原文摘要 · Abstract (English)
With the rapid advancement of Large Language Models (LLMs), the safety of LLMs has been a critical concern requiring precise assessment. Current benchmarks primarily concentrate on single-turn dialogues or a single jailbreak attack method to assess the safety. Additionally, these benchmarks have not taken into account the LLM's capability of identifying and handling unsafe information in detail. To address these issues, we propose a fine-grained benchmark SafeDialBench for evaluating the safety of LLMs across various jailbreak attacks in multi-turn dialogues. Specifically, we design a two-tier hierarchical safety taxonomy that considers 6 safety dimensions and generates more than 4000 multi-turn dialogues in both Chinese and English under 22 dialogue scenarios. We employ 7 jailbreak attack strategies, such as reference attack and purpose reverse, to enhance the dataset quality for dialogue generation. Notably, we construct an innovative assessment framework of LLMs, measuring capabilities in detecting, and handling unsafe information and maintaining consistency when facing jailbreak attacks. Experimental results across 17 LLMs reveal that Yi-34B-Chat and GLM4-9B-Chat demonstrate superior safety performance, while Llama3.1-8B-Instruct and o3-mini exhibit safety vulnerabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。