arXiv:2601.03699cs.CL2026-01被引 3

构建首个统一的大型语言模型对抗测试数据集,全面评估安全漏洞。

RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models

  • 整合37个权威数据集,涵盖29,362条攻击与拒绝提示。
  • 建立22类风险、19个领域的标准化分类体系,提升评估一致性。
  • 开源数据与代码,助力安全研究和模型鲁棒性提升。

随着大语言模型(LLMs)在安全关键场景中的广泛应用,其对对抗性提示的鲁棒性至关重要。然而,现有红队测试数据集存在风险分类不一致、领域覆盖有限、评估过时等问题,阻碍了系统性的漏洞评估。为此,我们提出RedBench,一个汇聚37个顶级会议与资源库基准数据集的通用数据集,包含29,362个样本,涵盖攻击与拒绝类提示。RedBench采用22个风险类别和19个领域的标准化分类体系,支持对大语言模型漏洞的系统化、全面评估。我们对现有数据集进行了详尽分析,建立了现代大语言模型的基线表现,并开源了数据集与评估代码。本工作促进可复现的对比研究,推动未来安全研究,助力高可靠大语言模型在真实场景中的部署。

原文摘要 · Abstract (English)

As large language models (LLMs) become integral to safety-critical applications, ensuring their robustness against adversarial prompts is paramount. However, existing red teaming datasets suffer from inconsistent risk categorizations, limited domain coverage, and outdated evaluations, hindering systematic vulnerability assessments. To address these challenges, we introduce RedBench, a universal dataset aggregating 37 benchmark datasets from leading conferences and repositories, comprising 29,362 samples across attack and refusal prompts. RedBench employs a standardized taxonomy with 22 risk categories and 19 domains, enabling consistent and comprehensive evaluations of LLM vulnerabilities. We provide a detailed analysis of existing datasets, establish baselines for modern LLMs, and open-source the dataset and evaluation code. Our contributions facilitate robust comparisons, foster future research, and promote the development of secure and reliable LLMs for real-world deployment. Code: https://github.com/knoveleng/redeval

大模型安全红队测试数据集对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。