构建可量化评估的AI滥用防护安全论证,助力实时风险响应
An Example Safety Case for Safeguards Against Misuse
- 通过红队测试估算规避防护措施所需努力,建立量化评估基础
- 利用上浮模型计算防护屏障对滥用行为的抑制效果,实现风险连续监测
- 提供可复用的安全论证框架,适合安全团队与开发者参考
当前AI滥用防护措施的评估多为零散证据,难以支持实际决策。为弥合这一差距,本文构建了一个端到端的安全论证(“安全案例”),证明滥用防护能将AI助手带来的风险降至低水平。首先,假设开发方对防护措施进行红队测试,估算规避所需努力;随后,将该估算结果输入定量“上浮模型”(https://www.aimisusemodel.com/),以确定防护措施对滥用行为的威慑程度。该流程在部署期间提供持续的风险信号,帮助开发方快速应对新兴威胁。最后,说明如何将各环节整合为简洁的安全案例。本研究提供了一条严谨论证AI滥用风险可控的具体路径,虽非唯一路径,但具有实践参考价值。
原文摘要 · Abstract (English)
Existing evaluations of AI misuse safeguards provide a patchwork of evidence that is often difficult to connect to real-world decisions. To bridge this gap, we describe an end-to-end argument (a "safety case") that misuse safeguards reduce the risk posed by an AI assistant to low levels. We first describe how a hypothetical developer red teams safeguards, estimating the effort required to evade them. Then, the developer plugs this estimate into a quantitative "uplift model" to determine how much barriers introduced by safeguards dissuade misuse (https://www.aimisusemodel.com/). This procedure provides a continuous signal of risk during deployment that helps the developer rapidly respond to emerging threats. Finally, we describe how to tie these components together into a simple safety case. Our work provides one concrete path -- though not the only path -- to rigorously justifying AI misuse risks are low.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。