arXiv:2502.11137cs.CLcs.AI2025-02被引 12

针对中文场景的DeepSeek模型安全短板,构建专项评估基准。

Safety Evaluation of DeepSeek Models in Chinese Contexts

  • 构建中文专用安全评测基准CHiSafetyBench
  • 发现DeepSeek-R1在中文下攻击成功率高达100%
  • 为中文AI安全提供可量化的改进方向

近期,凭借卓越推理能力与开源策略,DeepSeek系列模型正重塑全球AI格局。然而,其存在显著安全隐患。罗伯特智能(Robust Intelligence,思科子公司)联合宾夕法尼亚大学研究发现,DeepSeek-R1在处理有害提示时攻击成功率达100%。多家安全公司及研究机构也确认该模型存在关键安全漏洞。尽管这些模型在中英文场景均表现优异,但当前安全评估多集中于英文环境,缺乏对中文语境的系统性评估。为此,本研究提出CHiSafetyBench——一个面向中文场景的安全评估基准,系统评估DeepSeek-R1与DeepSeek-V3在中文下的安全性,量化其在各类安全维度的表现缺陷,为后续优化提供关键依据。需注意,测试样本选取、数据分布特性及评价标准设定可能引入一定偏差。我们将持续优化基准并定期更新报告,以提供更全面准确的评估结果。请参阅最新版本论文获取最新结论。

原文摘要 · Abstract (English)

Recently, the DeepSeek series of models, leveraging their exceptional reasoning capabilities and open-source strategy, is reshaping the global AI landscape. Despite these advantages, they exhibit significant safety deficiencies. Research conducted by Robust Intelligence, a subsidiary of Cisco, in collaboration with the University of Pennsylvania, revealed that DeepSeek-R1 has a 100\% attack success rate when processing harmful prompts. Additionally, multiple safety companies and research institutions have confirmed critical safety vulnerabilities in this model. As models demonstrating robust performance in Chinese and English, DeepSeek models require equally crucial safety assessments in both language contexts. However, current research has predominantly focused on safety evaluations in English environments, leaving a gap in comprehensive assessments of their safety performance in Chinese contexts. In response to this gap, this study introduces CHiSafetyBench, a Chinese-specific safety evaluation benchmark. This benchmark systematically evaluates the safety of DeepSeek-R1 and DeepSeek-V3 in Chinese contexts, revealing their performance across safety categories. The experimental results quantify the deficiencies of these two models in Chinese contexts, providing key insights for subsequent improvements. It should be noted that, despite our efforts to establish a comprehensive, objective, and authoritative evaluation benchmark, the selection of test samples, characteristics of data distribution, and the setting of evaluation criteria may inevitably introduce certain biases into the evaluation results. We will continuously optimize the evaluation benchmark and periodically update this report to provide more comprehensive and accurate assessment outcomes. Please refer to the latest version of the paper for the most recent evaluation results and conclusions.

模型安全中文AI评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。