打造可适应新安全类别的智能过滤系统,提升大模型输出安全性。
Taxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Models
- 用指令微调构建跨领域安全分类的过滤模型
- 在未见过的安全类别上仍保持高准确率
- 适合需要强内容安全防护的平台使用
大语言模型在后训练阶段虽已对齐安全准则,但仍可能生成不当内容。为此,本文提出 Roblox Guard 1.0,一个基于 Llama-3.1-8B-Instruct 的指令微调模型,通过输入输出双端协同监控提升系统安全性。该模型采用合成与开源安全数据混合训练,并引入思维链(CoT)推理与输入反转增强上下文理解能力,可在未见安全分类上实现良好泛化。为支持系统评估,研究还发布了 RobloxGuard-Eval 基准,包含可扩展的安全分类体系,用于评测大模型护栏与内容过滤框架的有效性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are typically aligned for safety during the post-training phase; however, they may still generate inappropriate outputs that could potentially pose risks to users. This challenge underscores the need for robust safeguards that operate across both model inputs and outputs. In this work, we introduce Roblox Guard 1.0, a state-of-the-art instruction fine-tuned LLM designed to enhance the safety of LLM systems through comprehensive input-output moderation, using a pipeline of LLMs to enhance moderation capability. Built on the Llama-3.1-8B-Instruct backbone, our model is instruction fine-tuned to generalize across previously unseen safety taxonomies and demonstrates strong performance on out-of-domain safety benchmarks. The instruction fine-tuning process uses a mix of synthetic and open-source safety datasets, augmented with chain-of-thought (CoT) rationales and input inversion to enhance contextual understanding and decision making. To support systematic evaluation, we also release RobloxGuard-Eval, a new benchmark featuring an extensible safety taxonomy to assess the effectiveness of LLM guardrails and moderation frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。