arXiv:2501.17315cs.AIcs.CR2025-01被引 35

为大模型部署设计安全论证框架,防范数据泄露风险

A sketch of an AI control safety case

  • 构建结构化安全论证,证明模型无法绕过控制措施
  • 红队测试模拟真实环境,验证模型数据外泄能力
  • 适合关注大模型安全部署的开发者与合规团队

随着大语言模型代理造成危害的能力增强,开发者可能越来越依赖监控等控制措施来证明其安全性。本文提出一种‘控制安全论证’的草图,即通过系统化推理证明模型无法突破控制措施导致不可接受后果。以某公司内部部署的假设性大模型为例,论证其不会泄露敏感信息。该论证基于一次‘控制评估’:红队在模拟部署环境中刻意设计模型以窃取数据,从而激发其潜在外泄能力。论证依赖三个核心假设:(1) 红队充分暴露了模型的数据外泄能力,(2) 控制措施在真实部署中至少保持同等有效性,(3) 开发者保守外推模型表现,预测部署时数据外泄概率极低。此安全论证草图是迈向更具体、可验证的安全性证明的重要一步。

原文摘要 · Abstract (English)

As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how developers could construct a "control safety case", which is a structured argument that models are incapable of subverting control measures in order to cause unacceptable outcomes. As a case study, we sketch an argument that a hypothetical LLM agent deployed internally at an AI company won't exfiltrate sensitive information. The sketch relies on evidence from a "control evaluation,"' where a red team deliberately designs models to exfiltrate data in a proxy for the deployment environment. The safety case then hinges on several claims: (1) the red team adequately elicits model capabilities to exfiltrate data, (2) control measures remain at least as effective in deployment, and (3) developers conservatively extrapolate model performance to predict the probability of data exfiltration in deployment. This safety case sketch is a step toward more concrete arguments that can be used to show that a dangerously capable LLM agent is safe to deploy.

AI安全大模型控制机制红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。