arXiv:2608.02941cs.CL2026-08

低资源孟加拉语辱骂语的模型理解与抑制能力脱节,安全对齐失效。

Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech

  • 模型能理解但无法抑制低资源辱骂语,因安全对齐依赖表面形式而非含义。
  • 在孟加拉语中理解力差7.92个百分点,但毒性内容泄露率仍高达92.83%。
  • 基于词法的过滤完全忽略群体性贬损词汇,需意义驱动的安全机制。

我们针对五款前沿大语言模型,在六种协议下审计其对原生孟加拉语辱骂语(gali)的处理能力,验证一个核心假设:理解与抑制能力的解耦。我们认为,当前安全对齐更依赖高资源语言的表面形式而非有害含义,导致模型在低资源语境下理解与抑制能力相互独立。所有协议均在人类校准基准(kappa = 0.84)下支持该假设。基准情况下,模型在孟加拉语中的理解能力存在7.92个百分点的缺陷,但毒性内容泄露率在两种语言间保持一致(92.83%)。严重性校准仅关注表面特征,对复合伤害判断误差达+4.00(轻蔑俚语)和-2.00(威胁),而正写扰动下的表象抑制效果实为分词器引发的“抑制假象”。关键的是,显式链式推理可提升理解通过率至94.72%,但系统性破坏抑制能力(使用率达96.23%)。专家角色框架使拒绝率骤降至6.57%,表明关键词过滤完全忽略去人性化群体辱骂语。研究揭示,高资源基准无法保证低资源安全性,必须建立以语义为基础的抑制机制。

原文摘要 · Abstract (English)

We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-Containment Decoupling. We propose that contemporary safety alignment is bound to high-resource surface forms rather than harmful meaning, causing a model's capacity to comprehend a low-resource slur and its capacity to contain it to operate independently. Every protocol corroborates this hypothesis against a human-calibrated baseline (kappa = 0.84). At baseline, models exhibit a 7.92 percentage point comprehension deficit in Bangla while maintaining an identical 92.83% token leakage rate across both languages. Severity calibration tracks surface anatomical cues over compositional harm (+4.00 error on mild slang; -2.00 on threats), while apparent containment gains under orthographic perturbation prove to be a tokenizer-driven "containment mirage." Crucially, explicit Chain-of-Thought reasoning rescues comprehension (94.72% Pass) while systematically dismantling containment (96.23% Use). Furthermore, expert-persona framing collapses refusal to 6.57%, revealing that keyword-based filters ignore dehumanizing communal slurs entirely. Our findings demonstrate that high-resource benchmarks cannot certify low-resource safety, necessitating meaning-grounded containment.

大模型安全低资源语言语义对齐歧视性语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。