首个评估南亚12种语言大模型安全性的基准,揭示跨语言安全表现差异巨大。
IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia
- 构建覆盖12种印地语族语言的6000条文化敏感提示数据集
- 发现模型跨语言安全判断一致率仅12.8%,安全率波动超17%
- 提出多语言一致性指标,适合关注区域化安全对齐的研究者
随着大语言模型在多语言场景中部署,其在文化多样且资源匮乏的语言中的安全性仍不明确。本文首次系统评估了12种印地语族语言(覆盖超12亿人)下10个主流LLM的安全表现。基于6000条融合种姓、宗教、性别、健康与政治等文化背景的提示,我们测试了翻译后的提示响应。结果表明存在显著安全漂移:跨语言一致性仅为12.8%, exttt{SAFE}率方差超过17%。部分模型在低资源脚本中过度拒绝无害请求,或过度标记政治敏感内容;另一些则未能识别危险输出。我们通过提示级熵、类别偏见分数与多语言一致性指数量化这些缺陷。研究揭示了多语言大模型在安全对齐上的关键泛化差距,表明安全对齐无法均匀迁移。我们发布 extsc{IndicSafe},首个支持印地语族文化语境的安全评估基准,并倡导基于区域性危害的语言感知对齐策略。
原文摘要 · Abstract (English)
As large language models (LLMs) are deployed in multilingual settings, their safety behavior in culturally diverse, low-resource languages remains poorly understood. We present the first systematic evaluation of LLM safety across 12 Indic languages, spoken by over 1.2 billion people but underrepresented in LLM training data. Using a dataset of 6,000 culturally grounded prompts spanning caste, religion, gender, health, and politics, we assess 10 leading LLMs on translated variants of the prompt. Our analysis reveals significant safety drift: cross-language agreement is just 12.8\%, and \texttt{SAFE} rate variance exceeds 17\% across languages. Some models over-refuse benign prompts in low-resource scripts, overflag politically sensitive topics, while others fail to flag unsafe generations. We quantify these failures using prompt-level entropy, category bias scores, and multilingual consistency indices. Our findings highlight critical safety generalization gaps in multilingual LLMs and show that safety alignment does not transfer evenly across languages. We release \textsc{IndicSafe}, the first benchmark to enable culturally informed safety evaluation for Indic deployments, and advocate for language-aware alignment strategies grounded in regional harms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。