自动化多轮多语言红队测试,发现模型安全漏洞随对话轮次和语言增多显著上升
Multi-lingual Multi-turn Automated Red Teaming for LLMs
- 设计多轮多语言自动红队方法,模拟真实对话中诱导模型生成危险内容
- 5轮对话后英文模型安全漏洞平均提升71%,非英语对话漏洞最高达195%增长
- 适用于评估多语言大模型安全性的研究人员与企业安全部门
近年来大型语言模型(LLMs)能力迅速提升,其在客户应用中的部署日益广泛。为保障安全性,研究者普遍采用红队测试——通过精心构造提示来试探模型是否会产生不安全响应。传统人工红队耗时费力,且难以覆盖多语言、多轮对话等新特性;现有自动化方法仅针对英语或单轮场景。本文提出多语言多轮自动化红队测试(MM-ART),实现对多语言、多轮对话情境下安全风险的全面评估。在多种语言上的实验表明,经过5轮对话后,英文模型的安全漏洞平均增加71%;而在非英语语境中,模型漏洞最高比标准单轮英文方法高出195%,充分证明了开发匹配模型实际能力的自动化红队方法的必要性。
原文摘要 · Abstract (English)
Language Model Models (LLMs) have improved dramatically in the past few years, increasing their adoption and the scope of their capabilities over time. A significant amount of work is dedicated to ``model alignment'', i.e., preventing LLMs to generate unsafe responses when deployed into customer-facing applications. One popular method to evaluate safety risks is \textit{red-teaming}, where agents attempt to bypass alignment by crafting elaborate prompts that trigger unsafe responses from a model. Standard human-driven red-teaming is costly, time-consuming and rarely covers all the recent features (e.g., multi-lingual, multi-modal aspects), while proposed automation methods only cover a small subset of LLMs capabilities (i.e., English or single-turn). We present Multi-lingual Multi-turn Automated Red Teaming (\textbf{MM-ART}), a method to fully automate conversational, multi-lingual red-teaming operations and quickly identify prompts leading to unsafe responses. Through extensive experiments on different languages, we show the studied LLMs are on average 71\% more vulnerable after a 5-turn conversation in English than after the initial turn. For conversations in non-English languages, models display up to 195\% more safety vulnerabilities than the standard single-turn English approach, confirming the need for automated red-teaming methods matching LLMs capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。