提出评估大模型是否提升普通人制造生化核武器能力的新框架。
A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models

- 设计可拆分的阈值超越评估框架,区分生成与改进两类风险。
- 实证发现仅放射性威胁存在显著风险提升,其他领域未达阈值。
- 适合政策制定者和模型开发者用于安全评估与部署决策。
随着前沿语言模型的发展,政策制定者与开发者亟需评估模型访问是否显著提升非专业人士策划高后果化学、生物、放射或核(CBRN)滥用的能力。现有评估在非专业人士定义、威胁范围、基线设置、评分标准和决策规则上差异显著,导致结果难以比较。本文提出阈值超越准则(TEC)框架,将提升效应研究分解为独立可执行的三部分:确定非专业人士参与资格、定义研究中的CBRN威胁范围、统计估计实质性提升。我们基于该框架开展大规模实证研究,区分两种提升形式:生成型(模型辅助从零创建计划)与修订型(模型协助优化已有计划)。研究生成了多个领域的攻击计划,并通过领域专家评审评估其生成与修订提升水平。结果显示,尽管部分模型辅助计划达到专家级指导水平,但实质性提升仅在放射性领域被确认。该发现用于指导缓解策略与部署治理,而非描述已部署模型的行为。文章最后总结未来评估的方法论启示:强调预设标准、明确基线、分离生成与修订评估、区分初步筛查信号与最终风险判定。
原文摘要 · Abstract (English)
As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misuse relative to public tools alone. Existing CBRN evaluations differ in non-expert definitions, threat scope, baselines, scoring rubrics, and decision rules, making results difficult to compare across studies. We introduce a Threshold Exceedance Criteria (TEC) framework that decomposes an uplift study into independently executable components: determining non-expert participant eligibility, defining the CBRN threat scope for the study, and statistically estimating material uplift. We then operationalize the TEC framework in a large-scale empirical study using a design that determines two forms of uplift: generative (where a model assists plan creation from scratch) and revisionist (where a model assists refinement of an existing plan). The study produced attack plans across the CBRN domains, which we evaluated through subject-matter-expert review to estimate generative and revisionist uplift. Applying the framework, our empirical study revealed domain heterogeneity: under this controlled pre-release evaluation, model-assisted plans sometimes received expert-equivalent instructional ratings, but confirmed material uplift was limited to the radiological domain. These findings informed mitigation and deployment-governance decisions rather than characterizing deployed model behavior. We conclude with methodological lessons for future CBRN uplift evaluations, emphasizing prespecified criteria, explicit baselines, separation of generative and revisionist estimates, and careful distinction between preliminary screening signals and confirmed risk determinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。