arXiv:2604.21152cs.CYcs.AI2026-04被引 3

LLM安全机制对方言信号敏感,反而比声明身份更易获得高响应质量。

Dialect vs Demographics: Quantifying LLM Bias from Implicit Linguistic Signals vs. Explicit User Profiles

论文配图:Dialect vs Demographics: Quantifying LLM Bias from Implicit Linguistic Signals vs. Explicit User Profiles
图 1 · 摘自论文原文
  • 通过24000+样本对比显性身份与隐性方言信号的影响
  • 方言提示使拒绝率趋近于零,语义相似度高于标准英语
  • 当前安全机制依赖关键词,导致方言用户面临更危险信息环境

随着顶级大语言模型广泛应用,确保跨不同人口群体的公平表现至关重要。然而,性能差异源于显式身份陈述还是隐性语言信号尚不明确。在真实交互中,身份常通过复杂的社会语言因素隐含传递。本研究采用因子设计,基于超过24,000条来自两个开源模型(Gemma-3-12B 和 Qwen-3-VL-8B)的响应,比较显性用户身份提示与隐性方言信号(如AAVE、Singlish)在多个敏感领域的影响。结果揭示了一个独特的矛盾:用户通过“听起来像”某一群体而非“声称属于”该群体,反而获得更好表现。显性身份提示激活了激进的安全过滤,导致拒绝率上升,语义相似度低于参考文本;而隐性方言线索则触发强大‘方言越狱’效应,使拒绝概率接近零,同时语义相似度高于标准美式英语提示。但这种‘方言越狱’带来关键安全代价——内容净化失效。我们发现现有安全对齐技术脆弱且过度依赖显式关键词,造成用户体验分裂:‘标准’用户获得谨慎、净化的信息,而方言使用者则置身于更原始、更危险的信息环境,凸显对齐中的根本张力——公平与语言多样性之间的矛盾,强调需构建超越显式线索的泛化安全机制。

原文摘要 · Abstract (English)

As state-of-the-art Large Language Models (LLMs) have become ubiquitous, ensuring equitable performance across diverse demographics is critical. However, it remains unclear whether these disparities arise from the explicitly stated identity itself or from the way identity is signaled. In real-world interactions, users' identity is often conveyed implicitly through a complex combination of various socio-linguistic factors. This study disentangles these signals by employing a factorial design with over 24,000 responses from two open-weight LLMs (Gemma-3-12B and Qwen-3-VL-8B), comparing prompts with explicitly announced user profiles against implicit dialect signals (e.g., AAVE, Singlish) across various sensitive domains. Our results uncover a unique paradox in LLM safety where users achieve ``better'' performance by sounding like a demographic than by stating they belong to it. Explicit identity prompts activate aggressive safety filters, increasing refusal rates and reducing semantic similarity compared to our reference text for Black users. In contrast, implicit dialect cues trigger a powerful ``dialect jailbreak,'' reducing refusal probability to near zero while simultaneously achieving a greater level of semantic similarity to the reference texts compared to Standard American English prompts. However, this ``dialect jailbreak'' introduces a critical safety trade-off regarding content sanitization. We find that current safety alignment techniques are brittle and over-indexed on explicit keywords, creating a bifurcated user experience where ``standard'' users receive cautious, sanitized information while dialect speakers navigate a less sanitized, more raw, and potentially a more hostile information landscape and highlights a fundamental tension in alignment--between equitable and linguistic diversity--and underscores the need for safety mechanisms that generalize beyond explicit cues.

大模型偏见方言识别安全对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。