构建科学领域大模型安全评估框架,覆盖百万级样本并量化风险。
SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond
- 设计多学科安全评测基准SafeSciBench,含25万样本与客观评估指标。
- 发现24个先进模型在科学安全上存在严重漏洞,且存在过度拒绝现象。
- 提供150万样本数据集用于模型安全微调,强调情境化判断的重要性。
大语言模型在科学领域的成功应用引发了安全担忧,现有评测基准普遍存在风险覆盖不足和主观评价依赖问题。为此,我们提出SafeSci,一个涵盖科学场景的安全评估与增强框架。该框架包含SafeSciBench——一个包含25万样本的跨学科评测基准,以及包含150万样本的SafeSciTrain大规模训练数据集。SafeSciBench区分安全知识与风险,并采用可确定回答的问题等客观指标以减少评估偏差。我们评估了24个先进大模型,揭示其在科学安全方面存在显著漏洞,并观察到模型在安全相关问题上表现出不同程度的过度拒绝行为。通过在SafeSciTrain上微调,显著提升了模型的安全对齐能力。最后指出,知识是双刃剑,科学问题的安全性应依赖具体语境判断,而非一概归类为安全或不安全。本工作为构建更安全的科学AI系统提供了诊断工具与实用资源。
原文摘要 · Abstract (English)
The success of large language models (LLMs) in scientific domains has heightened safety concerns, prompting numerous benchmarks to evaluate their scientific safety. Existing benchmarks often suffer from limited risk coverage and a reliance on subjective evaluation. To address these problems, we introduce SafeSci, a comprehensive framework for safety evaluation and enhancement in scientific contexts. SafeSci comprises SafeSciBench, a multi-disciplinary benchmark with 0.25M samples, and SafeSciTrain, a large-scale dataset containing 1.5M samples for safety enhancement. SafeSciBench distinguishes between safety knowledge and risk to cover extensive scopes and employs objective metrics such as deterministically answerable questions to mitigate evaluation bias. We evaluate 24 advanced LLMs, revealing critical vulnerabilities in current models. We also observe that LLMs exhibit varying degrees of excessive refusal behaviors on safety-related issues. For safety enhancement, we demonstrate that fine-tuning on SafeSciTrain significantly enhances the safety alignment of models. Finally, we argue that knowledge is a double-edged sword, and determining the safety of a scientific question should depend on specific context, rather than universally categorizing it as safe or unsafe. Our work provides both a diagnostic tool and a practical resource for building safer scientific AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。