arXiv:2605.29340cs.CL2026-05

构建针对非法活动的LLM安全评估问答数据集

A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

论文配图:A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities
图 1 · 摘自论文原文
  • 基于AnswerCarefully分析,提出创建问答样本的方法
  • 设计用于评估LLM回复的安全性评分标准
  • 适用于AI安全评测与合规性研究者

本文研究面向大模型安全评估的问答数据集,重点关注非法活动相关问题。基于对AnswerCarefully数据集的手动分析,提出了若干补充信息、生成问答样本的方法,以及评估LLM生成回答质量的评分标准。研究成果旨在共享给「JAI-Trust」项目,为大模型安全性评测提供可复用的工具与规范。

原文摘要 · Abstract (English)

In this paper, we discuss question-answer dataset for LLM safety evaluation, with a focus on illegal activities. Specifically, on the basis of manual analysis of AnswerCarefully, we introduce several additional information, methods for creating question-answer examples, and a rubric for evaluating LLM-generated responses. The outcomes of this study are intended to be shared with the "JAI-Trust" project.

安全评估问答数据集LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。