为印度语系语言打造安全防护模型,提升大模型本地化内容审核能力。
IndicGuard: A Multilingual Safety Guard Model and Dataset for Indic Languages

- 基于10种印度语构建多语言安全数据集,涵盖本地敏感议题与越狱攻击
- 微调40亿参数模型,在多轮对话中实现高一致性的内容过滤效果
- 可有效泛化至低资源印地语种,适合关注南亚地区AI安全的研究者
随着大语言模型在多元语言环境中的广泛应用,确保其安全性和与区域规范价值的一致性成为关键挑战。现有安全机制主要针对英语优化,难以捕捉印度地区的独特社会文化敏感性和局部危害类别。为此,我们提出IndicGuard,一个面向印度语言的多语言安全防护模型与数据集。我们构建了一个包含十种主要印度语言的高容量、文化细腻的安全数据集,系统化地涵盖区域危害、敏感社政背景及对抗性越狱样本。基于该语料,我们对Gemma-3-4B-IT进行指令微调,训练出一个40亿参数的多语言安全防护模型,用于实时内容审核与政策合规检查。实证评估表明,IndicGuard显著提升了大模型对本地漏洞的鲁棒性,在不同对话轮次间保持高一致性。关键的是,它在所有测试语言上均优于现有基线模型CultureGuard。最后,我们证明该模型能有效泛化至未参与训练的低资源印度语言,验证了框架的结构鲁棒性与跨语言迁移能力。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) achieve widespread integration across diverse linguistic landscapes, ensuring their safety and alignment with regional normative values remains a critical challenge. Current safety mechanisms are predominantly optimized for English-centric frameworks, often failing to capture the unique socio-cultural sensitivities and localized categories of harm inherent to the Indic region. To address this gap, we introduce IndicGuard, a multilingual safety guard model and dataset for Indic languages. We construct a high-volume, culturally nuanced safety dataset encompassing ten major Indic languages, systematically curated to capture regional harms, sensitive socio-political contexts, and adversarial jailbreaks. Leveraging this corpus, we fine-tune a 4B-parameter instruction-tuned model based on Gemma-3-4B-IT to serve as a multilingual safety guardrail for real-time content moderation and policy compliance checking. Our empirical evaluations demonstrate that IndicGuard significantly enhances LLM robustness against localized vulnerabilities, achieving high moderation consistency across different conversational turns. Crucially, IndicGuard consistently outperforms the existing baseline model, CultureGuard, across evaluated languages. Finally, we demonstrate that our model effectively generalizes to low-resource Indic languages excluded from training, substantiating the structural robustness and cross-lingual transfer capabilities of the framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。