让AI模型慢思考,通过结构化批评提升安全检测的准确性和解释性。
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
- 用高能力模型生成带结构化批评的安全标签,训练更谨慎的防护模型。
- 在多个基准上平均F1和AUPRC领先,比LLaMA Guard 3准确率高16.1%。
- 适合需要高安全性和可解释性的实际部署场景,如内容审核与医疗咨询。
确保大语言模型(LLMs)的安全性在真实应用中至关重要。现有防护机制依赖规则过滤或单次分类,难以处理细微的安全违规。为此,我们提出ThinkGuard,一种通过生成结构化批评并附带安全标签来提炼高容量模型知识的批判增强型防护模型。在包含批评的数据上微调后,该模型显著提升了防护的谨慎性与可解释性。在多个安全基准上的评估显示,ThinkGuard在平均F1和AUPRC上均优于所有基线。相比LLaMA Guard 3,其准确率提升16.1%,宏平均F1提升27.0%。此外,它超越仅使用标签微调的模型,证明结构化批评能同时提升分类精度与复杂安全推理能力,且保持计算高效。
原文摘要 · Abstract (English)
Ensuring the safety of large language models (LLMs) is critical as they are deployed in real-world applications. Existing guardrails rely on rule-based filtering or single-pass classification, limiting their ability to handle nuanced safety violations. To address this, we propose ThinkGuard, a critique-augmented guardrail model that distills knowledge from high-capacity LLMs by generating structured critiques alongside safety labels. Fine-tuned on critique-augmented data, the captured deliberative thinking ability drastically enhances the guardrail's cautiousness and interpretability. Evaluated on multiple safety benchmarks, ThinkGuard achieves the highest average F1 and AUPRC, outperforming all baselines. Compared to LLaMA Guard 3, ThinkGuard improves accuracy by 16.1% and macro F1 by 27.0%. Moreover, it surpasses label-only fine-tuned models, confirming that structured critiques enhance both classification precision and nuanced safety reasoning while maintaining computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。