arXiv:2608.24191cs.CLcs.AI2026-08

检测大模型在乌尔都语内容安全上的盲区,发现翻译后漏判率高达9.9%。

'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection

  • 用英文翻译对比乌尔都语原文,测试大模型对仇恨言论的识别一致性。
  • 乌尔都语原文中4.3%的有害内容在英文翻译后被漏判,最高达9.9%。
  • 小模型比大模型更不靠谱,且十年顶会无一篇乌尔都语相关研究。

乌尔都语是全球第十大使用语言,拥有2.46亿使用者,却几乎完全缺席主流大模型安全评估及近九年来WOAH会议论文。为检验此缺位是否影响内容审核可靠性,对GPT-4o、Claude Sonnet 4.5、Gemini 2.5 Flash、Qwen-2.5和Llama-3.1五款大模型在涵盖纳斯塔里克乌尔都语、罗马化乌尔都语、英语及乌尔都-英语混杂语的六大数据集上进行测试。在五个乌尔都语书写数据集中,原始文本与英文翻译分类结果不一致率介于15.9%(Gemini 2.5 Flash)至31.6%(Qwen-2.5)之间,其中‘漏判-乌尔都’率(即英文翻译标记为有害但原始文本视为正常)介于2.4%至9.9%之间(中位数4.3%)。通过ACL Anthology API完整枚举九届ALW/WOAH共205篇论文,确认整个时期无任何专门针对乌尔都语的研究。结果表明当前大模型在乌尔都语不同书写形式间提供不均衡的安全保障,开放权重的小模型稳定性与漏判率显著高于前沿闭源模型。

原文摘要 · Abstract (English)

Urdu, the world's tenth most spoken language with 246 million speakers, remains almost entirely absent from mainstream LLM safety evaluation and nine years of WOAH proceedings. To investigate whether this absence has measurable consequences for content moderation reliability, five large language models, GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen-2.5, and Llama-3.1, were tested across six datasets spanning Nastaliq Urdu, Roman Urdu, English, and code-switched Urdu-English. Across the five Urdu-script datasets, label instability between original-script and English-translation classification ranged from 15.9% (Gemini 2.5 Flash) to 31.6% (Qwen-2.5), with a 'Missed-in-Urdu' rate, content flagged as harmful in English translation but passed as normal in the original script, ranging from 2.4% to 9.9% (median 4.3%). A complete enumeration of all 205 papers across nine ALW/WOAH editions via the ACL Anthology API confirms zero dedicated Urdu papers across the entire period. Results indicate that current LLMs provide uneven safety assurance across Urdu's script varieties, with smaller open-weight models showing substantially higher instability and missed-harm rates than frontier closed models.

大模型安全多语言乌尔都语仇恨言论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。