arXiv:2503.16529cs.CLcs.AI2025-03被引 5

评测并提升DeepSeek系列模型在中文场景下的安全能力

Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts

  • 用CHiSafetyBench benchmark评估模型安全风险
  • 蒸馏后模型安全性显著下降,攻击成功率达100%
  • 提出增强方案,在保持推理能力前提下大幅提升安全性

DeepSeek-R1凭借卓越的推理能力和开源策略,对全球人工智能发展产生重要影响。然而其存在显著安全缺陷。研究发现,该模型在处理有害提示时攻击成功率达100%。多家安全机构已识别出关键漏洞。尽管中国联通已在中文语境中发现R1的安全问题,但对R1系列其他蒸馏模型的安全性尚未全面评估。为此,本研究采用中文安全基准CHiSafetyBench,深入评估DeepSeek-R1系列蒸馏模型在中文场景中的安全表现,分析蒸馏对模型安全性的负面影响。基于评估结果,针对整个R1系列实施针对性安全增强。评估显示,增强后的模型在保持推理能力的前提下显著提升安全性。相关模型已开源,地址为https://github.com/UnicomAI/DeepSeek-R1-Safe,供后续研究与优化使用。

原文摘要 · Abstract (English)

DeepSeek-R1, renowned for its exceptional reasoning capabilities and open-source strategy, is significantly influencing the global artificial intelligence landscape. However, it exhibits notable safety shortcomings. Recent research conducted by Robust Intelligence, a subsidiary of Cisco, in collaboration with the University of Pennsylvania, revealed that DeepSeek-R1 achieves a 100\% attack success rate when processing harmful prompts. Furthermore, multiple security firms and research institutions have identified critical security vulnerabilities within the model. Although China Unicom has uncovered safety vulnerabilities of R1 in Chinese contexts, the safety capabilities of the remaining distilled models in the R1 series have not yet been comprehensively evaluated. To address this gap, this study utilizes the comprehensive Chinese safety benchmark CHiSafetyBench to conduct an in-depth safety evaluation of the DeepSeek-R1 series distilled models. The objective is to assess the safety capabilities of these models in Chinese contexts both before and after distillation, and to further elucidate the adverse effects of distillation on model safety. Building on these findings, we implement targeted safety enhancements for the entire DeepSeek-R1 model series. Evaluation results indicate that the enhanced models achieve significant improvements in safety while maintaining reasoning capabilities without notable degradation. We open-source the safety-enhanced models at https://github.com/UnicomAI/DeepSeek-R1-Safe to serve as a valuable resource for future research and optimization of DeepSeek models.

模型安全中文AI蒸馏模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。