arXiv:2507.21929cs.AI2025-07被引 2

构建中文大模型安全防护系统,提升内容安全性与评估标准。

Libra: Large Chinese-based Safeguard for AI Content

  • 分两阶段训练:合成数据预训练+真实数据微调,减少人工标注依赖。
  • 在5700+样本的专用测试集上达到86.79%准确率,优于多个开源模型。
  • 首次提供中文内容安全评测基准,适合关注中文AI治理的研究者。

大型语言模型在文本理解与生成方面表现优异,但在高风险应用中引发显著安全与伦理问题。为缓解此类风险,本文提出Libra-Guard,一种专为中文大模型设计的先进安全防护系统。该系统采用两级课程训练流程:先在合成样本上进行守卫预训练,再基于高质量真实数据进行微调,显著降低对人工标注的依赖。为实现严谨的安全评估,我们还推出了首个针对中文内容安全防护系统的评测基准Libra-Test,涵盖七类关键危害场景,包含超过5,700个由领域专家标注的样本。实验表明,Libra-Guard在测试集上取得86.79%的准确率,超越Qwen2.5-14B-Instruct(74.33%)和ShieldLM-Qwen-14B-Chat(65.69%),接近Claude-3.5-Sonnet与GPT-4o等闭源模型水平。本工作为推进中文大模型的安全治理提供了坚实框架,是构建更安全、可靠中文AI系统的重要一步。

原文摘要 · Abstract (English)

Large language models (LLMs) excel in text understanding and generation but raise significant safety and ethical concerns in high-stakes applications. To mitigate these risks, we present Libra-Guard, a cutting-edge safeguard system designed to enhance the safety of Chinese-based LLMs. Leveraging a two-stage curriculum training pipeline, Libra-Guard enhances data efficiency by employing guard pretraining on synthetic samples, followed by fine-tuning on high-quality, real-world data, thereby significantly reducing reliance on manual annotations. To enable rigorous safety evaluations, we also introduce Libra-Test, the first benchmark specifically designed to evaluate the effectiveness of safeguard systems for Chinese content. It covers seven critical harm scenarios and includes over 5,700 samples annotated by domain experts. Experiments show that Libra-Guard achieves 86.79% accuracy, outperforming Qwen2.5-14B-Instruct (74.33%) and ShieldLM-Qwen-14B-Chat (65.69%), and nearing closed-source models like Claude-3.5-Sonnet and GPT-4o. These contributions establish a robust framework for advancing the safety governance of Chinese LLMs and represent a tentative step toward developing safer, more reliable Chinese AI systems.

大模型安全中文AI评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。