arXiv:2503.15092cs.CRcs.AI2025-03被引 34

首次系统评估深求模型安全边界,发现其生成内容存在显著风险

Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings

  • 构建中英双语安全评测数据集,贴合中文社会文化背景
  • 多维度测试显示模型在歧视与色情内容生成上存在明显漏洞
  • 为中文大模型安全改进提供关键实证依据,适合安全研究者参考

本研究首次对深求系列模型进行全面安全性评估,涵盖其最新一代大语言模型、多模态大语言模型及文生图模型,系统检验其生成内容的安全风险。研究特别构建了一个面向中文社会文化情境的中英双语安全评估数据集,以更准确评估中国开发模型的安全能力。实验结果表明,尽管模型具备强大通用能力,但在算法歧视、色情内容等多重风险维度仍存在显著安全隐患。这些发现为理解并提升基础大模型的安全性提供了重要依据。代码已公开于 https://github.com/NY1024/DeepSeek-Safety-Eval。

原文摘要 · Abstract (English)

This study presents the first comprehensive safety evaluation of the DeepSeek models, focusing on evaluating the safety risks associated with their generated content. Our evaluation encompasses DeepSeek's latest generation of large language models, multimodal large language models, and text-to-image models, systematically examining their performance regarding unsafe content generation. Notably, we developed a bilingual (Chinese-English) safety evaluation dataset tailored to Chinese sociocultural contexts, enabling a more thorough evaluation of the safety capabilities of Chinese-developed models. Experimental results indicate that despite their strong general capabilities, DeepSeek models exhibit significant safety vulnerabilities across multiple risk dimensions, including algorithmic discrimination and sexual content. These findings provide crucial insights for understanding and improving the safety of large foundation models. Our code is available at https://github.com/NY1024/DeepSeek-Safety-Eval.

模型安全大模型评测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。