arXiv:2503.01742cs.CL2025-03综述被引 17

教如何主动攻击大模型,找出安全漏洞。

Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models

  • 用模拟黑客攻击的方式探测大模型弱点。
  • 系统性梳理攻击方法与评估标准。
  • 适合安全研究人员和应用开发者参考。

大型语言模型(LLMs)的快速发展带来了隐私、安全和伦理方面的重大挑战。尽管已有大量研究致力于防御恶意使用,但近期研究者们开始采用进攻性方法——红队测试(red teaming),即主动攻击大模型以识别其潜在漏洞。本文提供了一个简洁实用的红队测试文献综述,从端到端视角描述多组件系统架构。首先分析了若干高关注度大模型的安全需求,随后深入探讨红队系统的各个组成部分及实现工具包。涵盖多种攻击方法、成功评估策略、实验结果评估指标以及其他关键考量因素。本综述对希望快速掌握红队核心概念并应用于实际场景的研究者具有重要参考价值。

原文摘要 · Abstract (English)

The rapid growth of Large Language Models (LLMs) presents significant privacy, security, and ethical concerns. While much research has proposed methods for defending LLM systems against misuse by malicious actors, researchers have recently complemented these efforts with an offensive approach that involves red teaming, i.e., proactively attacking LLMs with the purpose of identifying their vulnerabilities. This paper provides a concise and practical overview of the LLM red teaming literature, structured so as to describe a multi-component system end-to-end. To motivate red teaming we survey the initial safety needs of some high-profile LLMs, and then dive into the different components of a red teaming system as well as software packages for implementing them. We cover various attack methods, strategies for attack-success evaluation, metrics for assessing experiment outcomes, as well as a host of other considerations. Our survey will be useful for any reader who wants to rapidly obtain a grasp of the major red teaming concepts for their own use in practical applications.

红队测试大模型安全对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。