为印度多语言大模型打造安全评估基准,填补非英语语境下AI安全测试空白。
SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

- 构建覆盖10种印度主流语言的真实场景安全提示数据集
- 发现主流多语言模型在印地语系语言中存在过拒、漏检隐性偏见等问题
- 适合关注AI伦理、本地化部署及跨文化安全评估的研究者
现有大语言模型(LLMs)的安全评估数据集主要聚焦英语和西方语境,常忽略其他语言的语义多样性与文化特异性安全风险。为弥补这一缺口,我们提出SurakshaEval,一个由人工编写、涵盖真实世界场景的多语言安全评估基准,专为10种主要印度语言——阿萨姆语、孟加拉语、古吉拉特语、印地语、卡纳达语、马拉雅拉姆语、马拉地语、旁遮普语、泰米尔语和泰卢固语,以及英语设计。该数据集包含全印度通用提示及具有区域和语言特性的提示,以捕捉本地社会文化敏感点。我们在SurakshaEval上对多种先进多语言模型进行基准测试,建立安全性能基线,并识别出重复出现的失败模式,包括过度拒绝、对隐性偏见检测不足,以及在区域敏感情境下的上下文理解缺失。结果显示,即使强健的多语言模型在使用印地语系原生文字时也难以可靠满足复杂安全需求。这凸显了需构建融合区域数据与结构化评估协议的安全框架,以推动符合多元社会价值、安全且合乎伦理的AI系统发展。代码与数据已公开于 https://github.com/debobanerjee/SurakshaEval。注意:本文包含可能引发不适或不安全内容。
原文摘要 · Abstract (English)
Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages. To address this gap, we introduce SurakshaEval, a novel safety benchmark composed of human-written prompts spanning real-world scenarios, explicitly designed for ten major Indian languages - Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu, along with English. SurakshaEval includes both generic prompts common across India and region- and language-specific prompts that capture localized sociocultural sensitivities. We benchmark a broad range of state-of-the-art LLMs on SurakshaEval, establish baseline safety performance, and identify recurring failure modes, including over-refusal, missed detection of implicit bias, and insufficient contextual awareness in regionally sensitive settings. Our results show that even strong multilingual LLMs struggle to reliably meet nuanced safety requirements when operating in Indic languages, particularly in native scripts. These findings highlight the urgent need for safety evaluation frameworks that incorporate region-specific data and structured assessment protocols, enabling the development and deployment of AI systems that operate securely, ethically, and in alignment with diverse societal values. Our code and data are available at https://github.com/debobanerjee/SurakshaEval. Warning: This paper contains text that may be offensive or unsafe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。