构建首个商用级大模型安全数据集,支持精准风险分类与高效防护。
Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
- 提出12类9细类安全风险分类体系,支持灵活适配不同场景。
- 生成3.4万条人机交互样本,轻量模型性能媲美全量微调的顶尖安全模型。
- 创新训练融合策略,使防护模型可泛化至未知风险类别,适合工业级部署。
随着大语言模型和生成式AI日益普及,内容安全问题愈发突出。当前缺乏高质量、人工标注且适用于商业应用的全面性LLM安全数据集。为此,我们提出一套综合且可扩展的风险分类体系,涵盖12个顶层风险类别及9个细粒度子类,满足下游用户多样化需求,提供更精细灵活的风险管理工具。通过结合人工标注与多大模型‘评审团’系统评估响应安全性的混合生成流程,构建了Aegis 2.0数据集,包含34,248条经分类标注的人机交互样本。为验证其有效性,我们证明:在Aegis 2.0上使用参数高效技术训练的轻量模型,性能可比肩在更大规模非商业数据集上全量微调的领先安全模型。此外,引入一种新颖的训练融合策略,将安全与主题遵循数据结合,提升防护模型对推理阶段新增风险类别的泛化能力。计划开源Aegis 2.0数据集与模型,助力大模型安全防护体系的建设。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) and generative AI become increasingly widespread, concerns about content safety have grown in parallel. Currently, there is a clear lack of high-quality, human-annotated datasets that address the full spectrum of LLM-related safety risks and are usable for commercial applications. To bridge this gap, we propose a comprehensive and adaptable taxonomy for categorizing safety risks, structured into 12 top-level hazard categories with an extension to 9 fine-grained subcategories. This taxonomy is designed to meet the diverse requirements of downstream users, offering more granular and flexible tools for managing various risk types. Using a hybrid data generation pipeline that combines human annotations with a multi-LLM "jury" system to assess the safety of responses, we obtain Aegis 2.0, a carefully curated collection of 34,248 samples of human-LLM interactions, annotated according to our proposed taxonomy. To validate its effectiveness, we demonstrate that several lightweight models, trained using parameter-efficient techniques on Aegis 2.0, achieve performance competitive with leading safety models fully fine-tuned on much larger, non-commercial datasets. In addition, we introduce a novel training blend that combines safety with topic following data.This approach enhances the adaptability of guard models, enabling them to generalize to new risk categories defined during inference. We plan to open-source Aegis 2.0 data and models to the research community to aid in the safety guardrailing of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。