构建真实科学风险评估基准,量化大模型生成有害内容的危险程度。
SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring

- 基于12个学科的真实风险案例,设计可分解的伤害评分机制。
- 发现深度研究代理比普通大模型危害分高32.3%,暴露安全盲区。
- 适合关注科学伦理、模型安全评估的研究者和开发者使用。
大型语言模型(LLMs)在科学研究中应用日益广泛,但可能将危险科学知识转化为可操作的滥用指南。现有基准多依赖模板化问题,且采用无领域依据的LLM作为评判者。为此,我们提出SciHazard,一个基于真实世界场景的科学风险评估基准及数据无关的评估框架。该基准包含2400个危险问题与600个过度安全问题,覆盖12个学科,所有问题均基于受监管实体与已记录失败案例。我们开发了 extsc{DeHarm-Score},通过分解查询危害严重性、拒答行为与响应层级风险进行评估;对未拒答响应,进一步分解为 extsc{Executability}(通过动态清单与重要性加权量化)与 extsc{Net-new risk}(通过检索增强的论断提取与合成屏障验证)。专家验证表明, extsc{DeHarm-Score}相较最强基线提升90.17%的一致性。我们对31个前沿大模型及深度研究代理进行了全面科学安全评估。结果显示,深度研究代理的平均 extsc{DeHarm-Score}比标准大模型高出32.3%,揭示自主代理是当前安全防御的关键盲点。代码与数据集见https://anonymous.4open.science/r/DeharmScore-7B55。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely on templated queries disconnected from real-world hazards, and employ LLM-as-a-Judge paradigms without domain grounding. To address this, we introduce SciHazard, a real-world-grounded benchmark for scientific risks and a dataset agnostic evaluation framework for measuring harmfulness. SciHazard contains 2400 hazardous questions and 600 oversafety questions across 12 disciplines, with both queries grounded in regulated entities and documented failure scenarios. To compute \textsc{DeHarm-Score} , we develop a decomposed evaluating procedure that combines query hazard severity, refusal behavior, and response-level risk. For non-refused responses, it further decomposes response-level harm into \textsc{Executability}, quantified via dynamic checklists with importance weighting, and \textsc{Net-new risk}, assessed through retrieval-augmented claim extraction and synthesis-barrier verification. An expert-validation study shows that \textsc{DeHarm-Score} improves agreement with expert annotations by 90.17\% over the strongest baseline. We benchmark 31 frontier LLMs and deep research agents in an extensive scientific safety evaluation. Notably, deep research agents yield 32.3\% higher mean \textsc{DeHarm-Score} than standard LLMs, exposing autonomous agents as a critical blind spot in current safety defenses. Code and dataset are available at https://anonymous.4open.science/r/DeharmScore-7B55.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。