提出MOCET评分法,量化大模型在生物安全等真实场景中的风险。
Monte Carlo Expected Threat (MOCET) Scoring
- 基于蒙特卡洛模拟,评估模型生成威胁内容的可能性。
- 可扩展且开放,适配快速演进的大模型安全评估需求。
- 为政策制定者提供可解释的风险决策依据。
评估与衡量人工智能安全等级(ASL)威胁对引导利益相关方实施防护措施、将风险控制在可接受范围内至关重要。具备ASL-3+级别的模型可能提升新手非国家行为体的能力,尤其在生物安全领域存在独特风险。现有评估指标如LAB-Bench、BioLP-bench和WMDP可有效衡量模型的提升能力与领域知识,但缺乏能更好体现‘真实世界风险’的指标,也缺少可扩展、开放式的评估方法以跟上大模型的快速发展。为此,我们提出MOCET——一种可解释且双重可扩展(可自动化与开放性)的指标,能够量化真实世界中的潜在风险。
原文摘要 · Abstract (English)
Evaluating and measuring AI Safety Level (ASL) threats are crucial for guiding stakeholders to implement safeguards that keep risks within acceptable limits. ASL-3+ models present a unique risk in their ability to uplift novice non-state actors, especially in the realm of biosecurity. Existing evaluation metrics, such as LAB-Bench, BioLP-bench, and WMDP, can reliably assess model uplift and domain knowledge. However, metrics that better contextualize "real-world risks" are needed to inform the safety case for LLMs, along with scalable, open-ended metrics to keep pace with their rapid advancements. To address both gaps, we introduce MOCET, an interpretable and doubly-scalable metric (automatable and open-ended) that can quantify real-world risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。