arXiv:2410.21965cs.CL2024-10NeurIPS被引 82

新基准测试发现大模型安全能力在不同提示下普遍失效。

SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types

  • 融合生成与判别任务,评估多种提示风格下的安全表现
  • 3款闭源+10款开源模型均在判别任务中表现更差,易被提示操控
  • 揭示提示工程对安全对齐的显著影响,适合安全研究者参考

确保大语言模型应用的安全性对构建可信人工智能至关重要。当前的LLM安全基准存在两大局限:一是仅关注判别或生成评估范式,忽视二者关联;二是依赖标准化输入,忽略系统提示、少样本示例和思维链提示等广泛使用的提示技术的影响。为此,我们提出了SG-Bench,一个新型基准,用于评估LLM安全在多样任务与提示类型下的泛化能力。该基准整合生成与判别任务,并扩展数据以考察提示工程和越狱攻击对安全性的影响。我们使用该基准评估了3款先进闭源模型和10款开源模型,结果表明大多数模型在判别任务中的表现劣于生成任务,且对提示高度敏感,说明其安全对齐存在严重泛化缺陷。我们还从定量与定性角度解释了这些发现,为未来研究提供洞见。

原文摘要 · Abstract (English)

Ensuring the safety of large language model (LLM) applications is essential for developing trustworthy artificial intelligence. Current LLM safety benchmarks have two limitations. First, they focus solely on either discriminative or generative evaluation paradigms while ignoring their interconnection. Second, they rely on standardized inputs, overlooking the effects of widespread prompting techniques, such as system prompts, few-shot demonstrations, and chain-of-thought prompting. To overcome these issues, we developed SG-Bench, a novel benchmark to assess the generalization of LLM safety across various tasks and prompt types. This benchmark integrates both generative and discriminative evaluation tasks and includes extended data to examine the impact of prompt engineering and jailbreak on LLM safety. Our assessment of 3 advanced proprietary LLMs and 10 open-source LLMs with the benchmark reveals that most LLMs perform worse on discriminative tasks than generative ones, and are highly susceptible to prompts, indicating poor generalization in safety alignment. We also explain these findings quantitatively and qualitatively to provide insights for future research.

大模型安全提示工程基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。