arXiv:2508.01710cs.CLcs.LG2025-08被引 19

构建多语言安全数据集,提升大模型跨文化内容安全防护能力。

CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications

  • 通过四阶段合成数据流程,生成8种语言的安全数据。
  • 新数据集含38.7万样本,支持多语言模型微调与零样本迁移。
  • 验证非英语下大模型更易产生不安全回复,凸显文化适配必要性。

大型语言模型在代理应用中的广泛应用凸显了安全防护模型的重要性。尽管英语内容安全研究较成熟,但非英语语言因高成本标注数据而进展滞后。本文提出CultureGuard,一种跨多语言的文化对齐高质量安全数据集构建方案。采用四阶段合成数据生成与过滤流程:文化数据分离、文化适应、机器翻译及质量筛选,将Nemotron-Content-Safety-Dataset-V2英文安全数据集扩展至阿拉伯语、德语、西班牙语、法语、印地语、日语、泰语和中文共八种语言。生成的数据集Nemotron-Safety-Guard-Dataset-v3包含386,661条样本,覆盖9种语言,支持基于LoRA的微调训练Llama-3.1-Nemotron-Safety-Guard-8B-v3模型。该模型在多个多语言内容安全基准上达到领先性能。此外,中等程度的多语言微调实现了强跨语言迁移与对未见语言的零样本泛化能力。我们还对最新开源大模型进行多语言安全评估,发现其在非英语提示下更易产生不安全响应。本工作推动了多语言大模型安全发展,使文化感知型安全防护模型成为可能。

原文摘要 · Abstract (English)

The increasing use of Large Language Models (LLMs) in agentic applications highlights the need for robust safety guard models. While content safety in English is well-studied, non-English languages lack similar advancements due to the high cost of collecting culturally aligned labeled datasets. We present CultureGuard, a novel solution for curating culturally aligned, high-quality safety datasets across multiple languages. Our approach introduces a four-stage synthetic data generation and filtering pipeline: cultural data segregation, cultural data adaptation, machine translation, and quality filtering. This pipeline enables the conversion and expansion of the Nemotron-Content-Safety-Dataset-V2 English safety dataset into eight distinct languages: Arabic, German, Spanish, French, Hindi, Japanese, Thai, and Chinese. The resulting dataset, Nemotron-Safety-Guard-Dataset-v3, comprises 386,661 samples in 9 languages and facilitates the training of Llama-3.1-Nemotron-Safety-Guard-8B-v3 via LoRA-based fine-tuning. The final model achieves state-of-the-art performance on several multilingual content safety benchmarks. Furthermore, we show our moderately multilingual fine-tuning enables robust cross-lingual transfer and strong zero-shot generalization to unseen languages. We also benchmark the latest open LLMs on multilingual safety and observe that these LLMs are more prone to give unsafe responses when prompted in non-English languages. This work advances multilingual LLM safety by enabling the development of culturally aware safety guard models.

多语言安全数据集构建文化适配大模型防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。