arXiv:2411.06835cs.CLcs.CR2024-11被引 9

构建多层级危害评估基准,测试大模型在量化下的安全漏洞

HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment

  • 设计多级危害评估数据集,精细划分输出危害程度
  • 量化使模型对迁移攻击更鲁棒,但易受直接攻击影响
  • 针对Vicuna 13B模型评测主流越狱技术,揭示压缩风险

随着Transformer架构的兴起,大型语言模型在自然语言处理领域取得了突破性进展。然而,其快速发展也带来了安全挑战。模型计算能力与推理能力的指数级增长加剧了对其安全性的担忧。本文旨在填补现有研究中关于越狱技术与模型脆弱性评估的空白。我们构建了一个新型数据集,用于评估模型输出在多个危害等级上的表现,并聚焦于细粒度的危害级别分析。基于此框架,我们对当前最先进的越狱攻击进行了全面基准测试,特别针对Vicuna 13B v1.5模型。此外,我们研究了量化技术(如AWQ和GPTQ)对模型对齐性和鲁棒性的影响,发现量化在提升对迁移攻击的鲁棒性的同时,可能增加对直接攻击的脆弱性。本研究旨在揭示有害输入查询对越狱技术复杂性的影响,深化对大模型漏洞的理解,并改进在压缩策略下评估模型面对有害内容时鲁棒性的方法。

原文摘要 · Abstract (English)

With the introduction of the transformers architecture, LLMs have revolutionized the NLP field with ever more powerful models. Nevertheless, their development came up with several challenges. The exponential growth in computational power and reasoning capabilities of language models has heightened concerns about their security. As models become more powerful, ensuring their safety has become a crucial focus in research. This paper aims to address gaps in the current literature on jailbreaking techniques and the evaluation of LLM vulnerabilities. Our contributions include the creation of a novel dataset designed to assess the harmfulness of model outputs across multiple harm levels, as well as a focus on fine-grained harm-level analysis. Using this framework, we provide a comprehensive benchmark of state-of-the-art jailbreaking attacks, specifically targeting the Vicuna 13B v1.5 model. Additionally, we examine how quantization techniques, such as AWQ and GPTQ, influence the alignment and robustness of models, revealing trade-offs between enhanced robustness with regards to transfer attacks and potential increases in vulnerability on direct ones. This study aims to demonstrate the influence of harmful input queries on the complexity of jailbreaking techniques, as well as to deepen our understanding of LLM vulnerabilities and improve methods for assessing model robustness when confronted with harmful content, particularly in the context of compression strategies.

模型安全越狱攻击量化影响基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。