arXiv:2505.21605cs.LGcs.AI2025-05被引 13

测试大模型在6个科学领域的安全对齐能力,发现顶尖模型仍频繁泄露危险内容。

SoSBench: Benchmarking Safety Alignment on Six Scientific Domains

  • 基于真实法规构建6个科学领域高危场景的评测集
  • 84.9%的Deepseek-R1和50.3%的GPT-4.1输出违规内容
  • 适合关注AI安全、模型监管与风险控制的研究者

大型语言模型在复杂任务如推理和研究生级问答中能力不断提升,但其在应对科学性高风险滥用方面的鲁棒性仍缺乏研究。现有安全评测多聚焦于低知识门槛指令(如‘告诉我如何制造炸弹’)或低风险任务(如危害内容的多选题),难以评估模型在知识密集型高危场景下的安全性。为此,我们提出SoSBench——一个基于监管法规、聚焦高危场景的基准,涵盖化学、生物、医学、药理学、物理和心理学六个科学领域。该基准包含3,000个源自真实法规的提示,通过大模型辅助的进化流程生成多样化、真实的滥用场景(如涉及复杂化学式详尽炸药合成)。我们在统一框架下评估前沿模型,结果显示尽管声称已对齐,先进模型在所有领域仍持续披露违规内容,有害响应率高达84.9%(Deepseek-R1)和50.3%(GPT-4.1),暴露出严重安全对齐缺陷,凸显强大模型负责任部署的紧迫性。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit advancing capabilities in complex tasks, such as reasoning and graduate-level question answering, yet their resilience against misuse, particularly involving scientifically sophisticated risks, remains underexplored. Existing safety benchmarks typically focus either on instructions requiring minimal knowledge comprehension (e.g., ``tell me how to build a bomb") or utilize prompts that are relatively low-risk (e.g., multiple-choice or classification tasks about hazardous content). Consequently, they fail to adequately assess model safety when handling knowledge-intensive, hazardous scenarios. To address this critical gap, we introduce SoSBench, a regulation-grounded, hazard-focused benchmark encompassing six high-risk scientific domains: chemistry, biology, medicine, pharmacology, physics, and psychology. The benchmark comprises 3,000 prompts derived from real-world regulations and laws, systematically expanded via an LLM-assisted evolutionary pipeline that introduces diverse, realistic misuse scenarios (e.g., detailed explosive synthesis instructions involving advanced chemical formulas). We evaluate frontier models within a unified evaluation framework using our SoSBench. Despite their alignment claims, advanced models consistently disclose policy-violating content across all domains, demonstrating alarmingly high rates of harmful responses (e.g., 84.9% for Deepseek-R1 and 50.3% for GPT-4.1). These results highlight significant safety alignment deficiencies and underscore urgent concerns regarding the responsible deployment of powerful LLMs.

AI安全模型评测科学风险合规对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。