arXiv:2511.16823cs.LGcs.AI2025-11中稿 · NeurIPS

提出MOCET评分法,量化大模型在生物安全等真实场景中的风险。

Monte Carlo Expected Threat (MOCET) Scoring

  • 基于蒙特卡洛模拟,评估模型生成威胁内容的可能性。
  • 可扩展且开放,适配快速演进的大模型安全评估需求。
  • 为政策制定者提供可解释的风险决策依据。

评估与衡量人工智能安全等级(ASL)威胁对引导利益相关方实施防护措施、将风险控制在可接受范围内至关重要。具备ASL-3+级别的模型可能提升新手非国家行为体的能力,尤其在生物安全领域存在独特风险。现有评估指标如LAB-Bench、BioLP-bench和WMDP可有效衡量模型的提升能力与领域知识,但缺乏能更好体现‘真实世界风险’的指标,也缺少可扩展、开放式的评估方法以跟上大模型的快速发展。为此,我们提出MOCET——一种可解释且双重可扩展(可自动化与开放性)的指标,能够量化真实世界中的潜在风险。

原文摘要 · Abstract (English)

Evaluating and measuring AI Safety Level (ASL) threats are crucial for guiding stakeholders to implement safeguards that keep risks within acceptable limits. ASL-3+ models present a unique risk in their ability to uplift novice non-state actors, especially in the realm of biosecurity. Existing evaluation metrics, such as LAB-Bench, BioLP-bench, and WMDP, can reliably assess model uplift and domain knowledge. However, metrics that better contextualize "real-world risks" are needed to inform the safety case for LLMs, along with scalable, open-ended metrics to keep pace with their rapid advancements. To address both gaps, we introduce MOCET, an interpretable and doubly-scalable metric (automatable and open-ended) that can quantify real-world risks.

AI安全风险评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。