构建针对非法活动的LLM安全评估问答数据集
A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

- 基于AnswerCarefully分析,提出创建问答样本的方法
- 设计用于评估LLM回复的安全性评分标准
- 适用于AI安全评测与合规性研究者
本文研究面向大模型安全评估的问答数据集,重点关注非法活动相关问题。基于对AnswerCarefully数据集的手动分析,提出了若干补充信息、生成问答样本的方法,以及评估LLM生成回答质量的评分标准。研究成果旨在共享给「JAI-Trust」项目,为大模型安全性评测提供可复用的工具与规范。
原文摘要 · Abstract (English)
In this paper, we discuss question-answer dataset for LLM safety evaluation, with a focus on illegal activities. Specifically, on the basis of manual analysis of AnswerCarefully, we introduce several additional information, methods for creating question-answer examples, and a rubric for evaluating LLM-generated responses. The outcomes of this study are intended to be shared with the "JAI-Trust" project.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。