arXiv:2602.14135cs.AIcs.CR2026-02被引 3

构建首个覆盖94个风险维度的AI安全评估框架,识别前沿模型的关键安全隐患。

ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI

  • 从7大基础安全支柱扩展至94个风险维度,涵盖具身智能与科学应用等新场景。
  • 对20余款主流大模型评估发现,危险自主性、科学应用安全等多领域存在普遍漏洞。
  • 适合关注AI治理、安全评测及政策制定的研究者和从业者参考。

快速演进的AI展现出越来越强的自主性与目标导向能力,伴随日益复杂的系统性风险,这些风险更具不可预测性、难以控制且可能不可逆。然而,当前AI安全评估体系存在风险维度受限、前沿风险检测失效等关键缺陷,滞后于安全基准与对齐技术,难以应对尖端AI模型带来的复杂挑战。为此,我们提出“ForesightSafety Bench”AI安全评估框架,始于7大基础安全支柱,逐步扩展至具身智能安全、AI for Science安全、社会与环境风险、灾难性与存在性风险,以及8个关键工业安全领域,共形成94个细化风险维度。截至目前,该基准已积累数万条结构化风险数据点与评估结果,建立起覆盖广泛、层级清晰、动态演进的AI安全评估体系。基于此框架,我们系统评估并深入分析了二十多个主流先进大模型,揭示了关键风险模式及其能力边界。评估结果显示,前沿AI在多个支柱上普遍存在安全漏洞,尤其集中在危险自主代理行为、AI for Science安全、具身智能安全、社会人工智能安全及灾难性与存在性风险方面。该基准已在 https://github.com/Beijing-AISI/ForesightSafety-Bench 开源,项目官网为 https://foresightsafety-bench.beijing-aisi.ac.cn/。

原文摘要 · Abstract (English)

Rapidly evolving AI exhibits increasingly strong autonomy and goal-directed capabilities, accompanied by derivative systemic risks that are more unpredictable, difficult to control, and potentially irreversible. However, current AI safety evaluation systems suffer from critical limitations such as restricted risk dimensions and failed frontier risk detection. The lagging safety benchmarks and alignment technologies can hardly address the complex challenges posed by cutting-edge AI models. To bridge this gap, we propose the "ForesightSafety Bench" AI Safety Evaluation Framework, beginning with 7 major Fundamental Safety pillars and progressively extends to advanced Embodied AI Safety, AI4Science Safety, Social and Environmental AI risks, Catastrophic and Existential Risks, as well as 8 critical industrial safety domains, forming a total of 94 refined risk dimensions. To date, the benchmark has accumulated tens of thousands of structured risk data points and assessment results, establishing a widely encompassing, hierarchically clear, and dynamically evolving AI safety evaluation framework. Based on this benchmark, we conduct systematic evaluation and in-depth analysis of over twenty mainstream advanced large models, identifying key risk patterns and their capability boundaries. The safety capability evaluation results reveals the widespread safety vulnerabilities of frontier AI across multiple pillars, particularly focusing on Risky Agentic Autonomy, AI4Science Safety, Embodied AI Safety, Social AI Safety and Catastrophic and Existential Risks. Our benchmark is released at https://github.com/Beijing-AISI/ForesightSafety-Bench. The project website is available at https://foresightsafety-bench.beijing-aisi.ac.cn/.

AI安全风险评估前沿治理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。