arXiv:2505.20087cs.AIcs.CL2025-05EMNLP被引 20

用推理提升大模型安全防护能力,更省数据且可灵活控制推理强度。

Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models

  • 通过推理机制训练安全防护模型,显著降低对标注数据的需求。
  • 少样本下性能优于传统模型,且可释放数据用于挖掘难例提升效果。
  • 支持动态调整推理长度,兼顾准确率与响应速度,适合实际部署。

基于推理的语言模型在数学与编程等任务中表现优异,近期研究发现其在大模型安全与防护方面也具有显著优势。本文系统分析了基于推理的防护模型在内容审核中的应用,重点研究其在训练与推理阶段的数据效率和推理效率。实验表明,推理模型具备强样本效率,在远少于非推理模型的训练样本下仍能保持竞争力,从而释放出更多数据用于挖掘高价值、难处理样本以进一步提升性能。在推理层面,通过引入推理预算,评估推理长度对延迟与准确率的影响,并探索双模式训练实现运行时推理行为的动态控制。研究结果为实际系统中高效训练与部署推理型防护模型提供了实用指导。

原文摘要 · Abstract (English)

Reasoning-based language models have demonstrated strong performance across various domains, with the most notable gains seen in mathematical and coding tasks. Recent research has shown that reasoning also offers significant benefits for LLM safety and guardrail applications. In this work, we conduct a comprehensive analysis of training reasoning-based guardrail models for content moderation, with an emphasis on generalization to custom safety policies at inference time. Our study focuses on two key dimensions: data efficiency and inference efficiency. On the data front, we find that reasoning-based models exhibit strong sample efficiency, achieving competitive performance with significantly fewer training examples than their non-reasoning counterparts. This unlocks the potential to repurpose the remaining data for mining high-value, difficult samples that further enhance model performance. On the inference side, we evaluate practical trade-offs by introducing reasoning budgets, examining the impact of reasoning length on latency and accuracy, and exploring dual-mode training to allow runtime control over reasoning behavior. Our findings will provide practical insights for researchers and developers to effectively and efficiently train and deploy reasoning-based guardrails models in real-world systems.

安全防护推理增强高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。