为科学领域AI安全设计多维度评估基准,识别高风险场景
SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

- 从风险维度与学科交叉视角构建评测体系
- 覆盖7大学科、31个子领域和10类风险
- 适合评估科学大模型在高风险任务中的安全性
大型语言模型正深度融入人工智能赋能科学(AI4Science)流程,涵盖科学问答、文献分析、实验规划与自主发现等。这一进展迫切需要能评估模型科学能力与风险意识的安全基准。现有数据集虽覆盖多个学科与任务形式,但缺乏对底层风险维度的明确刻画。本文提出「SciRisk-Bench」,一个从显式风险维度与科学学科双重角度评估AI4Science安全性的基准,涵盖7个学科、31个子学科和10种风险维度。实验部分对主流大模型与科学专用模型在不同风险维度、学科及子领域上的表现进行评估,实现对科学模型安全短板的细粒度诊断。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature analysis to laboratory planning and autonomous discovery. This progress creates an urgent need for safety benchmarks that evaluate not only scientific competence, but also whether models recognize and avoid risks in high-stakes scientific contexts. Existing AI4Science safety datasets cover several disciplines and task formats, leaving the underlying risk dimensions underspecified. We introduce \textbf{SciRisk-Bench}, a benchmark designed to evaluate AI4Science safety from two complementary perspectives: explicit risk dimensions and scientific disciplines. SciRisk-Bench covers 7 disciplines, 31 subdisciplines and 10 risk dimensions. In the experimental section, we evaluate both mainstream LLMs and science-oriented LLMs across risk dimensions, disciplines, and sub-disciplines, enabling fine-grained diagnosis of where scientific models remain unsafe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。