arXiv:2501.07238cs.AI2025-01被引 33

基于100个AI产品红队测试,提炼出8条安全实战关键经验。

Lessons From Red Teaming 100 Generative AI Products

  • 通过真实产品测试构建威胁模型,聚焦系统能力与应用场景。
  • 自动化可拓展风险覆盖范围,但人力判断不可替代。
  • 适合关注AI安全落地的开发者与安全团队参考。

近年来,人工智能红队测试已成为探测生成式AI系统安全与可信性的实践方法。由于该领域尚处初期,如何开展红队操作仍存在诸多开放问题。基于微软内部对超过100个生成式AI产品的红队测试经验,本文提出我们的内部威胁模型本体,并总结出八条核心经验:1. 明确系统能力及其应用场景;2. 无需计算梯度即可破坏AI系统;3. 红队测试不等同于安全基准评估;4. 自动化有助于覆盖更广风险面;5. 人为因素在红队中至关重要;6. 责任型AI危害普遍存在但难量化;7. 大语言模型放大既有安全风险并引入新风险;8. 保障AI系统的任务永无止境。结合实际案例,本文提供与现实风险对齐的实用建议,并澄清常被误解的方面,同时提出该领域需思考的开放问题。

原文摘要 · Abstract (English)

In recent years, AI red teaming has emerged as a practice for probing the safety and security of generative AI systems. Due to the nascency of the field, there are many open questions about how red teaming operations should be conducted. Based on our experience red teaming over 100 generative AI products at Microsoft, we present our internal threat model ontology and eight main lessons we have learned: 1. Understand what the system can do and where it is applied 2. You don't have to compute gradients to break an AI system 3. AI red teaming is not safety benchmarking 4. Automation can help cover more of the risk landscape 5. The human element of AI red teaming is crucial 6. Responsible AI harms are pervasive but difficult to measure 7. LLMs amplify existing security risks and introduce new ones 8. The work of securing AI systems will never be complete By sharing these insights alongside case studies from our operations, we offer practical recommendations aimed at aligning red teaming efforts with real world risks. We also highlight aspects of AI red teaming that we believe are often misunderstood and discuss open questions for the field to consider.

AI安全红队测试大模型风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。