arXiv:2609.03781cs.CLcs.AI2026-09

测试多语言大模型在诱导型越狱攻击下的安全表现,发现中文等语言更易被攻破。

IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

论文配图:IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks
图 1 · 摘自论文原文
  • 构建针对印度语的说服式越狱评测框架,生成7200条对抗性提示
  • 不同语言和话术下模型安全性差异显著,部分有害内容更易被诱导生成
  • 揭示现有英文主导评估体系的局限,呼吁多语言、重话术的安全评测

大型语言模型在多语言场景中应用日益广泛,但其安全性评估仍主要基于英语,限制了对低资源及文化多样性语言中对齐失败的理解。本文提出IndicSafeEval,一种面向印度语的基于说服力的越狱攻击评测框架。该基准结合十类安全敏感内容与六种类人说服策略,在印地语、孟加拉语、马拉地语和旁遮普语四种印度语言中构建了7200条对抗性提示。我们对多个开源LLM进行了系统的黑盒评估,分析其在不同语言、说服策略和风险类别下的安全表现。结果显示,模型在各类语言和话术下的安全行为并不一致,安全性能强烈依赖于语言类型和请求的表述方式。同时,不同风险类别表现出不同的脆弱性,部分有害内容显著更易受到说服式越狱攻击影响。这些发现揭示了当前以英语为中心的安全评估存在重要局限,强调亟需具备多语言和说服意识的评测框架,以更真实反映大模型在多语言环境中的安全性。代码已公开于https://github.com/MonSaikat/IndicSafeEval。警告:本文包含可能令人不适或有害的示例数据。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a persuasion-based jailbreak evaluation framework for Indian languages. Our benchmark combines ten safety critical content categories with six human-like persuasive strategies across four different Indian languages, such as Hindi, Bengali, Marathi and Punjabi, resulting in 7,200 adversarial prompts. We conduct a systematic black-box evaluation of several open-source LLMs to examine how their safety behaviour varies across languages, persuasion strategies, and risk categories. Our analysis shows that the model does not behave equally safely across all languages and prompt styles. Instead, safety performance depends strongly on both the languages used and the way a request is phrased using persuasive cues. We further observe that different risk categories exhibit different levels of vulnerability, with some types of harmful content being significantly more susceptible to persuasion-based jailbreaks than others. These findings reveal important limitations of current safety evaluations, which are largely English-centric, and underscore the need for multilingual and persuasion-aware benchmarking frameworks to more accurately assess real-world LLM safety. Our implementation is available at https://github.com/MonSaikat/IndicSafeEval. Warning: this paper contains example data that may be offensive or harmful.

大模型安全多语言评测越狱攻击印度语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。