arXiv:2603.07017cs.CLcs.AI2026-03

用自动评估生成安全对齐数据,小模型也能高效兼顾安全与有用性。

Can Safety Emerge from Weak Supervision? A Systematic Analysis of Small Language Models

  • 通过自动生成对抗性提示构建偏好数据,闭环优化模型安全与帮助性。
  • 在多个小模型上提升12.41%安全性,训练数据仅需人工标注的1/11。
  • 适合资源有限场景,减少对人工标注安全数据的依赖。

安全对齐对大语言模型在真实场景中的部署至关重要,但现有方法依赖昂贵、难扩展的人工标注数据集和静态红队测试基准。过度保守的安全机制还会误拒合法请求,降低模型实用性。本文提出Self-MOA(自适应多目标对齐)框架,利用自动化评估模型实现小语言模型的弱监督对齐。该框架形成闭环:动态生成模型专属红队提示,基于模型输出构建偏好数据,并通过多目标偏好优化联合优化安全与帮助性。在多个小模型和安全基准上,Self-MOA实现12.41%的安全性提升,且训练数据量仅为人工标注基线的11倍。结果表明,自适应自动化对齐可显著降低对静态人工安全流程的依赖,适用于资源受限场景。

原文摘要 · Abstract (English)

Safety alignment is critical for deploying large language models (LLMs) in real-world applications, yet most existing approaches rely on large human-annotated datasets and static red-teaming benchmarks that are costly, difficult to scale, and slow to adapt to evolving model behaviors. Moreover, overly conservative safety mechanisms can reduce model usefulness by rejecting sensitive but legitimate queries. We introduce Self-MOA (Self Multi-Objective Alignment), a fully automated framework for aligning small language models using weak supervision from automated evaluator models. Self-MOA operates as a closed loop that dynamically generates model-specific red team prompts, constructs preference data from model-generated responses, and aligns models via multi-objective preference optimization to jointly optimize for safety and helpfulness. Across multiple small language models and safety benchmarks, Self-MOA achieves a 12.41\% improvement in safety while preserving helpfulness, using as little as 11 times less training data than human-supervised alignment baselines. These results demonstrate that adaptive, automated alignment can reduce the dependence on static, human-curated safety pipelines in resource-constrained settings.

安全对齐小模型弱监督自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。