动态更新的AI安全基准,随政策变化自动生成新测试题。
AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models
- 自动追踪各国法规,按四层风险分类并新增细粒度类别。
- 覆盖335个风险点,新类别源自7国31项新规,平均难度提升0.06分。
- 适合关注AI合规性、安全评估及多语言风险的研究者使用。
基础模型安全基准往往在发布后迅速过时:随着模型能力提升和政府出台新法规,原有风险分类和攻击提示逐渐失效。本文提出AIR-BENCH Live,作为AIR-BENCH 2024的自演化升级版本。其自动化更新流水线持续监控政府监管政策,并将新法规归类至现有四层风险体系,或提出新的细粒度类别。随后,基于多智能体与角色驱动的提示生成算法,自动创建真实、多语言的攻击提示,仅需少量人工审核,支持对新型越狱技术的持续适应。该算法用于重构旧提示并生成新类别的提示。当前版本将基准风险项从314项扩展至335项,新增21个类别源自7个司法管辖区的31项全新政策条款。评估14个近期模型发现,安全表现差异显著(0.17至0.89),现代提示平均比2024版难0.06分,尤其对高合规模型影响最大,且多数模型在非英语提示下安全性下降。通过持续吸收新法规并重生成提示,AIR-BENCH Live旨在与快速演进的领域同步发展。
原文摘要 · Abstract (English)
Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation, their risk taxonomies become incomprehensive and their attack prompts become ineffective. We present AIR-BENCH Live, a self-evolving successor to AIR-BENCH 2024. An automated update pipeline monitors government regulation and classifies new policies against the current four-tier risk taxonomy, either matching them to existing categories or proposing new granular categories. Then, a multi-agent, persona-driven prompt generation algorithm generates realistic, multilingual prompts with minimal human review, leaving room for improvement with modern jail breaking techniques. This algorithm is used to overhaul legacy prompts and generate prompts for new categories. In our current version, the pipeline has expanded the benchmark from 314 to 335 granular risks, with the 21 new categories drawing from 31 truly novel policy clauses across seven jurisdictions. Evaluating 14 recent models, we find a wide safety spread (from 0.17 to 0.89 among the models judged on their own behavior), that the modernized prompts are on average 0.06 points harder than the 2024 set, with the largest drops concentrated among the most compliant models, and that most models are modestly less safe on non-English prompts. By continuously absorbing new regulation and regenerating prompts, AIR-BENCH Live is designed to evolve alongside a fast-moving field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。