arXiv:2602.16832cs.AIcs.CL2026-02Conference of the …被引 5

首个面向南亚语言的免人工评估越狱鲁棒性测试,揭示多语言安全漏洞。

IndicJR: A Judge-Free Benchmark of Jailbreak Robustness in South Asian Languages

  • 构建12种南亚语言的免人工评判越狱测试基准,覆盖4.5万条指令
  • 英文攻击在南亚语言中转移成功率超92%,自然语境下模型全失效
  • 罗马化输入显著降低防御效果,代码混用者风险更高

大型语言模型的安全对齐主要在英语环境下评估且依赖合同约束,导致多语言漏洞未被充分研究。本文提出「印地越狱鲁棒性(Indic Jailbreak Robustness, IJR)」,一个免人工评判的跨12种南亚语言(21亿使用者)的对抗安全基准,涵盖45216条指令,分JSON(合同式)与Free(自然式)两类。实验揭示三大规律:(1) 合同格式虽提升拒绝率,但无法阻止越狱——在JSON模式下,LLaMA和Sarvam模型的越狱成功率达0.92以上;在Free模式下,所有模型拒绝率归零,越狱成功达1.0;(2) 英文越狱攻击可强迁移至南亚语言,格式包装器比指令包装器更有效;(3) 拼写方式影响防御性能:罗马化或混合输入显著降低越狱成功率(相关系数约0.28~0.32),反映系统性偏差。人工审计验证检测器可靠性,轻量到完整模型对比结论一致。IJR提供可复现的多语言压力测试,揭示英语单一评估掩盖的风险,尤其对频繁混用语言与罗马化的南亚用户群体。

原文摘要 · Abstract (English)

Safety alignment of large language models (LLMs) is mostly evaluated in English and contract-bound, leaving multilingual vulnerabilities understudied. We introduce \textbf{Indic Jailbreak Robustness (IJR)}, a judge-free benchmark for adversarial safety across 12 Indic and South Asian languages (2.1 Billion speakers), covering 45216 prompts in JSON (contract-bound) and Free (naturalistic) tracks. IJR reveals three patterns. (1) Contracts inflate refusals but do not stop jailbreaks: in JSON, LLaMA and Sarvam exceed 0.92 JSR, and in Free all models reach 1.0 with refusals collapsing. (2) English to Indic attacks transfer strongly, with format wrappers often outperforming instruction wrappers. (3) Orthography matters: romanized or mixed inputs reduce JSR under JSON, with correlations to romanization share and tokenization (approx 0.28 to 0.32) indicating systematic effects. Human audits confirm detector reliability, and lite-to-full comparisons preserve conclusions. IJR offers a reproducible multilingual stress test revealing risks hidden by English-only, contract-focused evaluations, especially for South Asian users who frequently code-switch and romanize.

多语言安全越狱攻击南亚语言基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。