arXiv:2607.18268cs.AI2026-07

用小模型+合成数据打造专用安全防护,提升大模型应用安全性

Fence: Specialized SLM Guardrails for LLM Applications

  • 用类GAN方法生成高质量合成数据,训练专用小模型作安全守卫
  • 在幻觉、话题偏移等场景下,性能优于传统提示词防护方式
  • 适合需要定制化安全机制的LLM落地应用,降低标注成本

现实世界中使用闭源大语言模型的应用需要超越基础内容过滤的高级安全措施。毒性与偏见等内容审核有较标准定义,而幻觉、话题漂移和行为偏差等应用场景特定的防护机制更难建模且因场景而异。此外,数据稀缺和标注成本使定制化防护的构建与测试变得困难。本文提出利用在合成数据上训练的小语言模型(SLMs)作为大模型应用的专用安全守卫。我们设计了一种受生成对抗网络(GAN)启发的新型合成数据生成方法,生成高质量样本以训练SLMs,使其编码特定使用场景的安全规则,从而充当专业化防护机制。实验表明,基于高质量合成数据训练的SLM守卫,在多个任务上表现优于基于提示的LLM防护方案。

原文摘要 · Abstract (English)

Real-world applications that use closed-source large language models (LLMs) need advanced safety measures that go beyond the basic content filters. Content moderation filters such as toxicity and bias have relatively standard definitions where as application specific guardrails like hallucination, topic drift and behaviour deviation are more difficult to model and can vary by use case. Additionally, data scarcity and annotation costs, make the process of creating and testing specialized guardrails challenging. In this work, we propose using Small Language Models (SLMs) trained on synthetic data as specialized guardrails for LLM applications. We introduce a novel synthetic data generation method inspired by the design of Generative Adversarial Networks (GANs) to generate high quality synthetic data samples which can be used to train SLMs to encode use case specific guardrail information and hence function as specialized guardrails. Our experiments demonstrate that SLM guardrails trained on high quality synthetic data show performance gains over prompt based LLM guardrails.

小模型安全防护合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。