构建科学任务安全评估基准,覆盖分子、蛋白等多语言场景
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
- 设计跨科学语言的多模态安全评测框架
- 引入'越狱'测试强化对恶意意图的防御能力
- 适用于科研用大模型的安全性验证与风险控制
大型语言模型在生物学、化学、医学和物理等学科的科学研究中具有变革性影响。然而,确保这些模型在科学任务中的安全性对齐仍是一个研究不足的领域,现有基准主要关注文本内容,忽略了分子、蛋白质和基因组语言等关键科学表示形式。此外,大模型在科学任务中的安全机制尚未得到充分研究。为此,我们提出 SciSafeEval,一个涵盖多种科学任务的综合性安全对齐评估基准。该基准覆盖文本、分子、蛋白质和基因组等多种科学语言,横跨广泛科学领域。我们在零样本、少样本和思维链设置下评估大模型,并引入“越狱”增强功能,挑战配备安全防护机制的大模型,严格测试其抵御恶意意图的能力。该基准在规模和覆盖范围上均超越现有安全数据集,为评估大模型在科学场景下的安全性和性能提供了坚实平台。本工作旨在推动大模型在科研中的负责任发展与部署,促进其与安全及伦理标准对齐。
原文摘要 · Abstract (English)
Large language models (LLMs) have a transformative impact on a variety of scientific tasks across disciplines including biology, chemistry, medicine, and physics. However, ensuring the safety alignment of these models in scientific research remains an underexplored area, with existing benchmarks primarily focusing on textual content and overlooking key scientific representations such as molecular, protein, and genomic languages. Moreover, the safety mechanisms of LLMs in scientific tasks are insufficiently studied. To address these limitations, we introduce SciSafeEval, a comprehensive benchmark designed to evaluate the safety alignment of LLMs across a range of scientific tasks. SciSafeEval spans multiple scientific languages-including textual, molecular, protein, and genomic-and covers a wide range of scientific domains. We evaluate LLMs in zero-shot, few-shot and chain-of-thought settings, and introduce a "jailbreak" enhancement feature that challenges LLMs equipped with safety guardrails, rigorously testing their defenses against malicious intention. Our benchmark surpasses existing safety datasets in both scale and scope, providing a robust platform for assessing the safety and performance of LLMs in scientific contexts. This work aims to facilitate the responsible development and deployment of LLMs, promoting alignment with safety and ethical standards in scientific research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。