arXiv:2507.08284cs.LGcs.AI2025-07被引 1

小模型通过合成数据与强化学习训练,实现高效内容安全防护。

Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training

  • 用人工种子数据生成高保真合成样本,提升多样性与上下文丰富度。
  • 结合强化学习指导的对抗训练,使小模型在内容审核上超越大模型。
  • 适合需要低资源、强鲁棒性的AI内容安全部署场景。

本文提出一种轻量级但高效的语言模型安全防护框架,证明小型语言模型可在内容审核任务中达到甚至超越大型模型的表现。该方法通过高保真合成数据生成与对抗训练实现:以人工标注的种子数据为基础,经查询增强与改写生成多样且语境丰富的示例,并经过多轮筛选确保数据质量;受生成对抗网络(GAN)启发,采用强化学习引导生成器构造具有挑战性的合成样本,用于微调安全分类器,从而提升对有害内容的检测与抑制能力。同时,借鉴高效大模型训练策略,利用小模型能力优化大模型性能。通过迭代式对抗训练与高质量合成数据生成,本框架使小型语言模型(SLMs)可作为稳健的安全防护屏障,显著降低计算开销,增强对对抗攻击的抵抗力,为AI系统的内容审核提供可扩展、高效率的解决方案。

原文摘要 · Abstract (English)

We introduce a lightweight yet highly effective safety guardrail framework for language models, demonstrating that small-scale language models can achieve, and even surpass, the performance of larger counterparts in content moderation tasks. This is accomplished through high-fidelity synthetic data generation and adversarial training. The synthetic data generation process begins with human-curated seed data, which undergoes query augmentation and paraphrasing to create diverse and contextually rich examples. This augmented data is then subjected to multiple rounds of curation, ensuring high fidelity and relevance. Inspired by recent advances in the Generative Adversarial Network (GAN) architecture, our adversarial training employs reinforcement learning to guide a generator that produces challenging synthetic examples. These examples are used to fine-tune the safety classifier, enhancing its ability to detect and mitigate harmful content. Additionally, we incorporate strategies from recent research on efficient LLM training, leveraging the capabilities of smaller models to improve the performance of larger generative models. With iterative adversarial training and the generation of diverse, high-quality synthetic data, our framework enables small language models (SLMs) to serve as robust safety guardrails. This approach not only reduces computational overhead but also enhances resilience against adversarial attacks, offering a scalable and efficient solution for content moderation in AI systems.

安全防护小模型合成数据对抗训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。