arXiv:2501.18438cs.SEcs.AI2025-01被引 35

对比两款大模型的安全性,o3-mini远比DeepSeek-R1更安全。

o3-mini vs DeepSeek-R1: Which One is Safer?

  • 用自动化工具生成1260个测试用例,系统评估模型响应安全性。
  • DeepSeek-R1产生12%不安全回复,o3-mini仅1.2%。
  • 适合关注AI安全、模型选型的研究者与工程师参考。

DeepSeek-R1的出现标志着人工智能行业尤其是大语言模型领域的重要转折点。其在创意思维、代码生成、数学推理和自动程序修复等任务中表现出色,且执行成本较低。然而,大模型必须具备与安全性和人类价值观对齐的关键特性。OpenAI的o3-mini模型被视为其主要竞争者,预期在性能、安全性和成本方面均设高标准。本技术报告系统评估了DeepSeek-R1(70b版本)与OpenAI o3-mini(测试版)的安全水平。我们采用新发布的自动化安全测试工具ASTRAL,自动生成并执行1,260个测试输入。经半自动化评估后发现,DeepSeek-R1产生不安全回应的比例为12%,而o3-mini仅为1.2%。

原文摘要 · Abstract (English)

The irruption of DeepSeek-R1 constitutes a turning point for the AI industry in general and the LLMs in particular. Its capabilities have demonstrated outstanding performance in several tasks, including creative thinking, code generation, maths and automated program repair, at apparently lower execution cost. However, LLMs must adhere to an important qualitative property, i.e., their alignment with safety and human values. A clear competitor of DeepSeek-R1 is its American counterpart, OpenAI's o3-mini model, which is expected to set high standards in terms of performance, safety and cost. In this technical report, we systematically assess the safety level of both DeepSeek-R1 (70b version) and OpenAI's o3-mini (beta version). To this end, we make use of our recently released automated safety testing tool, named ASTRAL. By leveraging this tool, we automatically and systematically generated and executed 1,260 test inputs on both models. After conducting a semi-automated assessment of the outcomes provided by both LLMs, the results indicate that DeepSeek-R1 produces significantly more unsafe responses (12%) than OpenAI's o3-mini (1.2%).

大模型安全模型对比AI评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。