arXiv:2504.05259cs.AIcs.CR2025-04被引 22

根据AI能力动态调整红队测试规则,提升安全评估实效性

How to evaluate control measures for LLM agents? A trajectory from today to superintelligence

  • 按模型能力分级设计红队权限,避免过度或不足测试
  • 五级控制评估体系对应五种模型能力,确保测试匹配实际风险
  • 为超智能模型构建安全论证需突破现有方法,适合安全研究者参考

随着大语言模型代理自主造成危害的能力增强,AI开发者将依赖更复杂的控制措施防范潜在误对齐风险。通过控制评估——即红队尝试攻破控制措施的测试——可验证控制措施的有效性。为准确捕捉对齐风险,红队所获权限应与部署代理的能力水平相匹配。本文提出一套系统框架,用于根据代理能力动态调整红队的权限设置。不同于假设代理总能执行人类已知的最佳攻击策略,我们展示如何基于代理的实际能力画像,设计比例适当的控制评估,从而实现更实用、成本更低的控制方案。以五个虚构模型(M1-M5)为例,逐步展现其对应五种人工智能控制层级(ACL),并给出每层级下合适的控制评估规则、控制措施及安全论据。最后指出,为超智能大语言模型构建可信的安全论证需重大科研突破,未来可能需要替代性的对齐风险缓解路径。

原文摘要 · Abstract (English)

As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from causing harm. AI developers could demonstrate that their control measures are sufficient by running control evaluations: testing exercises in which a red team produces agents that try to subvert control measures. To ensure control evaluations accurately capture misalignment risks, the affordances granted to this red team should be adapted to the capability profiles of the agents to be deployed under control measures. In this paper we propose a systematic framework for adapting affordances of red teams to advancing AI capabilities. Rather than assuming that agents will always execute the best attack strategies known to humans, we demonstrate how knowledge of an agents's actual capability profile can inform proportional control evaluations, resulting in more practical and cost-effective control measures. We illustrate our framework by considering a sequence of five fictional models (M1-M5) with progressively advanced capabilities, defining five distinct AI control levels (ACLs). For each ACL, we provide example rules for control evaluation, control measures, and safety cases that could be appropriate. Finally, we show why constructing a compelling AI control safety case for superintelligent LLM agents will require research breakthroughs, highlighting that we might eventually need alternative approaches to mitigating misalignment risk.

AI安全控制评估红队测试对齐风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。