发现大模型跨语言安全表现不一,可能因语言不同而出现漏洞。
LLMs Lost in Translation: M-ALERT uncovers Cross-Linguistic Safety Inconsistencies
- 构建多语言安全评测基准M-ALERT,覆盖5种语言共7.5万条高质量提示。
- 39个主流模型测试显示,同一模型在不同语言中安全表现差异显著。
- 部分敏感类别如毒品、犯罪宣传在所有语言中均易引发不安全回复。
构建跨语言安全的大型语言模型对保障安全访问和语言多样性至关重要。为此,我们对当前大模型生态进行了大规模、全面的安全评估。为此,我们提出M-ALERT,一个涵盖英语、法语、德语、意大利语和西班牙语的多语言基准,每种语言包含1.5万条高质量提示,总计7.5万条,并进行类别标注。我们在39个前沿大模型上开展广泛实验,揭示语言特异性安全分析的重要性:模型在不同语言和类别间表现出显著的安全不一致性。例如,Llama3.2在意大利语的crime_tax类别中表现高度不安全,但在其他语言中则相对安全。此类不一致性在所有模型中普遍存在。相比之下,substance_cannabis和crime_propaganda等类别在所有模型和语言中均持续触发不安全响应。这些发现强调了建立稳健多语言安全机制的必要性,以确保模型在多元社区中的负责任使用。
原文摘要 · Abstract (English)
Building safe Large Language Models (LLMs) across multiple languages is essential in ensuring both safe access and linguistic diversity. To this end, we conduct a large-scale, comprehensive safety evaluation of the current LLM landscape. For this purpose, we introduce M-ALERT, a multilingual benchmark that evaluates the safety of LLMs in five languages: English, French, German, Italian, and Spanish. M-ALERT includes 15k high-quality prompts per language, totaling 75k, with category-wise annotations. Our extensive experiments on 39 state-of-the-art LLMs highlight the importance of language-specific safety analysis, revealing that models often exhibit significant inconsistencies in safety across languages and categories. For instance, Llama3.2 shows high unsafety in category crime_tax for Italian but remains safe in other languages. Similar inconsistencies can be observed across all models. In contrast, certain categories, such as substance_cannabis and crime_propaganda, consistently trigger unsafe responses across models and languages. These findings underscore the need for robust multilingual safety practices in LLMs to ensure responsible usage across diverse communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。