X-Guard提升多语言内容安全,防低资源语言与混码攻击
X-Guard: Multilingual Guard Agent for Content Moderation
- 构建132种语言、500万条数据的多语言安全数据集
- 通过双阶段架构有效识别跨语言危险内容,准确率显著提升
- 引入评委机制和透明推理,适合多语言AI安全研发者使用
大型语言模型在关键领域应用日益广泛,但现有安全防护体系在多语言场景下仍存在明显漏洞,尤其对低资源语言和混码攻击防御能力不足。当前系统多基于英语设计,缺乏跨语言泛化能力。尽管如Llama Guard-3等模型已具备多语言支持,但决策过程不透明。为此,我们提出X-Guard,一个透明的多语言安全代理,可有效抵御常规低资源语言攻击及复杂代码切换攻击。方法包括:整理并增强多个开源安全数据集,提供明确评估理由;采用评委机制减少单个大模型提供商的偏见;构建涵盖132种语言、共500万条数据的综合性多语言安全数据集;设计两阶段架构,结合自定义微调的mBART-50翻译模块与通过监督微调和GRPO训练的X-Guard 3B模型。实验证明,X-Guard在多语言环境下能高效检测有害内容,并保持全程决策透明,推动构建更鲁棒、透明且语言包容的LLM安全体系。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have rapidly become integral to numerous applications in critical domains where reliability is paramount. Despite significant advances in safety frameworks and guardrails, current protective measures exhibit crucial vulnerabilities, particularly in multilingual contexts. Existing safety systems remain susceptible to adversarial attacks in low-resource languages and through code-switching techniques, primarily due to their English-centric design. Furthermore, the development of effective multilingual guardrails is constrained by the scarcity of diverse cross-lingual training data. Even recent solutions like Llama Guard-3, while offering multilingual support, lack transparency in their decision-making processes. We address these challenges by introducing X-Guard agent, a transparent multilingual safety agent designed to provide content moderation across diverse linguistic contexts. X-Guard effectively defends against both conventional low-resource language attacks and sophisticated code-switching attacks. Our approach includes: curating and enhancing multiple open-source safety datasets with explicit evaluation rationales; employing a jury of judges methodology to mitigate individual judge LLM provider biases; creating a comprehensive multilingual safety dataset spanning 132 languages with 5 million data points; and developing a two-stage architecture combining a custom-finetuned mBART-50 translation module with an evaluation X-Guard 3B model trained through supervised finetuning and GRPO training. Our empirical evaluations demonstrate X-Guard's effectiveness in detecting unsafe content across multiple languages while maintaining transparency throughout the safety evaluation process. Our work represents a significant advancement in creating robust, transparent, and linguistically inclusive safety systems for LLMs and its integrated systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。