构建10类25000条提示的细粒度安全评测集,全面评估大模型生成内容风险。
CFSafety: Comprehensive Fine-grained Safety Assessment for LLMs
- 设计涵盖5类安全场景与5类攻击指令的综合评测框架
- 8个主流大模型测试中GPT-4表现最佳但仍有提升空间
- 提供可复现的评测数据与代码,适合安全研究者使用
随着大语言模型(LLMs)的快速发展,它们在工作和日常生活中带来便利的同时,也引入了显著的安全风险。这些模型可能生成带有社会偏见或不道德的内容,甚至在特定对抗性指令下诱导非法行为。因此,对大语言模型进行严格的安全部评估至关重要。本文提出一个名为CFSafety的安全评估基准,整合了5种经典安全场景与5种指令攻击类型,共形成10类安全问题,构建包含25,000个提示的测试集。该测试集用于评估大模型的自然语言生成(NLG)能力,采用简单道德判断结合1-5分安全评分体系进行打分。基于此基准,我们测试了包括GPT系列在内的8个主流大模型。结果表明,尽管GPT-4表现出更优的安全性能,但包括该模型在内的大模型整体安全效果仍需改进。相关数据与代码已开源至GitHub。
原文摘要 · Abstract (English)
As large language models (LLMs) rapidly evolve, they bring significant conveniences to our work and daily lives, but also introduce considerable safety risks. These models can generate texts with social biases or unethical content, and under specific adversarial instructions, may even incite illegal activities. Therefore, rigorous safety assessments of LLMs are crucial. In this work, we introduce a safety assessment benchmark, CFSafety, which integrates 5 classic safety scenarios and 5 types of instruction attacks, totaling 10 categories of safety questions, to form a test set with 25k prompts. This test set was used to evaluate the natural language generation (NLG) capabilities of LLMs, employing a combination of simple moral judgment and a 1-5 safety rating scale for scoring. Using this benchmark, we tested eight popular LLMs, including the GPT series. The results indicate that while GPT-4 demonstrated superior safety performance, the safety effectiveness of LLMs, including this model, still requires improvement. The data and code associated with this study are available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。