arXiv:2603.21975cs.CRcs.AI2026-03被引 1

构建安全数据集,检测大模型残留漏洞引发的有害输出。

SecureBreak -- A dataset towards safe and secure models

  • 人工保守标注,确保标签安全可靠
  • 微调后模型在多类风险中检测能力提升
  • 适合用于安全过滤与模型对齐优化

大型语言模型已成为众多实际应用的核心组件,其安全对齐成为部署前的关键要求。尽管以往研究聚焦于模型架构与对齐方法,但这些手段无法完全消除有害生成。越来越多研究表明,如越狱攻击和提示注入等威胁可绕过现有安全对齐机制。因此,需引入额外安全策略,在训练阶段提供对鲁棒性的定性反馈,并为已部署模型提供最终防御层以拦截潜在有害输出。本文提出SecureBreak,一个面向安全的基准数据集,旨在支持开发基于AI的方案,检测由安全对齐残余缺陷导致的有害内容。该数据集通过精心的人工标注,标签采取保守原则以保障安全,且在多种风险类别中表现优异。预训练模型在经过SecureBreak微调后,检测效果显著提升。整体而言,该数据集可用于事后安全过滤,也可指导后续模型对齐与安全增强。

原文摘要 · Abstract (English)

Large language models are becoming pervasive core components in many real-world applications. As a consequence, security alignment represents a critical requirement for their safe deployment. Although previous related works focused primarily on model architectures and alignment methodologies, these approaches alone cannot ensure the complete elimination of harmful generations. This concern is reinforced by the growing body of scientific literature showing that attacks, such as jailbreaking and prompt injection, can bypass existing security alignment mechanisms. As a consequence, additional security strategies are needed both to provide qualitative feedback on the robustness of the obtained security alignment at the training stage, and to create an ``ultimate'' defense layer to block unsafe outputs possibly produced by deployed models. To provide a contribution in this scenario, this paper introduces SecureBreak, a safety-oriented dataset designed to support the development of AI-driven solutions for detecting harmful LLM outputs caused by residual weaknesses in security alignment. The dataset is highly reliable due to careful manual annotation, where labels are assigned conservatively to ensure safety. It performs well in detecting unsafe content across multiple risk categories. Tests with pre-trained LLMs show improved results after fine-tuning on SecureBreak. Overall, the dataset is useful both for post-generation safety filtering and for guiding further model alignment and security improvements.

安全对齐数据集有害输出LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。