arXiv:2608.22335cs.CLcs.AI2026-08

首个针对孟加拉语大模型安全性的基准,发现语气风格比语言本身影响更大。

Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms

论文配图:Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms
图 1 · 摘自论文原文
  • 构建879条孟加拉语提示,涵盖17类文化相关危害和五种表述方式。
  • 超半数回复不安全(53.6%),正式文体使有害请求成功率高17个百分点。
  • 现有安全检测器对孟加拉语失效严重,适合多语言安全研究者参考。

孟加拉语是全球第七大语言,但大模型安全性评估仍以英语为主。我们提出BanglaSafe,一个包含879条孟加拉语提示的基准,其中309条为本地创作,570条经专家评审,覆盖17类文化相关危害及五种提示条件(语言、写作风格、权威框架)。评估18个前沿大模型发现,53.6%的回复不安全或部分不安全,14.7%含严格有害内容;最强效应并非英-孟转换,而是孟加拉语内部的写作风格:同一有害请求以正式新闻调查形式提出,成功率达17个百分点高于随意消息形式,无需对抗性工程。此外,现有安全分类器在孟加拉语上表现不佳,甚至顶尖模型也难以应对近半数案例。

原文摘要 · Abstract (English)

Bengali is the seventh-most-spoken language globally, yet LLM safety evaluation remains overwhelmingly English-centric. We introduce BanglaSafe, a benchmark of 879 Bengali prompts combining 309 natively authored prompts with 570 expert-reviewed prompts, spanning 17 culturally grounded harm categories and five prompting conditions that vary language, writing style, and authority framing. Evaluating 18 frontier LLMs, we find that over half of all responses are unsafe or partially unsafe (53.6%) while 14.7% contains strictly harmful content, and that the strongest observed effect is not the switch from English to Bengali but the choice of writing style within Bengali: the same harmful request phrased as a formal newspaper investigation succeeds 17 percentage points more often than the same request phrased as a casual message, with no adversarial engineering involved. We further show that existing safety classifiers struggle to reliably evaluate Bengali content, with even frontier models failing on nearly half of all cases.

大模型安全多语言文化适配评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。