arXiv:2501.17749cs.SEcs.AI2025-01被引 15

对OpenAI o3-mini模型进行早期安全测试,发现87个潜在风险行为。

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation

  • 用自动化工具ASTRAL生成10080个危险提示测试模型
  • 手动验证后发现87个实际不安全行为实例
  • 适合关注大模型安全评估的研究者与开发者

大型语言模型已深度融入日常生活,但可能带来隐私泄露、偏见强化和虚假信息传播等风险。为此,亟需建立完善的安全部署机制与全面测试流程。本文报告了来自莫德纳大学与塞维利亚大学的研究人员,作为OpenAI早期安全测试计划的一部分,对OpenAI新推出的o3-mini模型所开展的外部安全评估。研究采用其自研工具ASTRAL,自动且系统地生成最新版高危测试输入(即提示),以评估模型在多个安全维度的表现。共生成并执行10,080个高危测试输入,在人工验证由ASTRAL标记为不安全的案例后,确认存在87个真实的安全违规行为。研究揭示了该模型在预发布阶段的关键安全洞见。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become an integral part of our daily lives. However, they impose certain risks, including those that can harm individuals' privacy, perpetuate biases and spread misinformation. These risks highlight the need for robust safety mechanisms, ethical guidelines, and thorough testing to ensure their responsible deployment. Safety of LLMs is a key property that needs to be thoroughly tested prior the model to be deployed and accessible to the general users. This paper reports the external safety testing experience conducted by researchers from Mondragon University and University of Seville on OpenAI's new o3-mini LLM as part of OpenAI's early access for safety testing program. In particular, we apply our tool, ASTRAL, to automatically and systematically generate up to date unsafe test inputs (i.e., prompts) that helps us test and assess different safety categories of LLMs. We automatically generate and execute a total of 10,080 unsafe test input on a early o3-mini beta version. After manually verifying the test cases classified as unsafe by ASTRAL, we identify a total of 87 actual instances of unsafe LLM behavior. We highlight key insights and findings uncovered during the pre-deployment external testing phase of OpenAI's latest LLM.

大模型安全测试评估O3-mini

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。