测试主流大模型防护机制在多语言毒性内容下的表现
Benchmarking LLM Guardrails in Handling Multilingual Toxicity
- 构建覆盖十种语言的多语言毒性测试集
- 发现现有防护机制对多语言毒性无效且易被越狱攻击绕过
- 揭示语言资源少时防护效果显著下降,适合安全研究者参考
随着大语言模型(LLMs)的广泛应用,防护机制在检测和抵御有毒内容方面变得至关重要。然而,在多语言场景中,这些防护机制的有效性尚不明确。本文构建了一个涵盖七个数据集和十余种语言的综合性多语言测试套件,用于评估前沿防护机制的表现。同时,研究了防护机制对最新越狱技术的鲁棒性,并分析了上下文安全策略及语言资源可用性对其性能的影响。结果表明,现有防护机制在处理多语言毒性内容方面仍显不足,且对越狱提示缺乏抵抗力。本研究旨在揭示防护机制的局限性,推动构建更可靠、可信的多语言大模型系统。
原文摘要 · Abstract (English)
With the ubiquity of Large Language Models (LLMs), guardrails have become crucial to detect and defend against toxic content. However, with the increasing pervasiveness of LLMs in multilingual scenarios, their effectiveness in handling multilingual toxic inputs remains unclear. In this work, we introduce a comprehensive multilingual test suite, spanning seven datasets and over ten languages, to benchmark the performance of state-of-the-art guardrails. We also investigates the resilience of guardrails against recent jailbreaking techniques, and assess the impact of in-context safety policies and language resource availability on guardrails' performance. Our findings show that existing guardrails are still ineffective at handling multilingual toxicity and lack robustness against jailbreaking prompts. This work aims to identify the limitations of guardrails and to build a more reliable and trustworthy LLMs in multilingual scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。