arXiv:2608.13328cs.CLcs.AI2026-08

女性常用语言特征让大模型回应更简短,影响职场沟通公平性

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

  • 用女性常见语言特征提问,模型回应更短更简单
  • 语言风格影响比性别标识词更大,且在多模型中稳定存在
  • 现有方法难缓解,因语言模式深植于模型底层结构

专业沟通日益依赖大语言模型,但这些模型是否对所有用户公平?我们发现,当提示包含女性更常使用的语言特征(如缓和语、附加疑问句、集体指代)时,无论在三种文档类型下还是四种模型中,都会系统性地引发更短、更不复杂、更不正式的回应。这一效应在控制提示复杂度和特征传递后依然存在。显式性别线索(如署名)与语言方言共享同一表征空间,暗示其背后机制相似,但语言风格的影响远大于姓名。进一步研究显示,事后缓解困难:由于这些模式根植于文化,用户无法通过自我呈现规避;机制分析表明,语言特征在Transformer早期层即被编码,并与其他特征纠缠。本文呼吁在上游设计中考虑语言差异,以减轻大模型在职场沟通中的不公平影响。

原文摘要 · Abstract (English)

Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encoded in the same representational space as linguistic dialect - suggesting shared underlying mechanisms - yet linguistic register is far more influential, producing large, consistent effects where names produce none. Our results further reveal that post-hoc mitigation is challenging: because these patterns are culturally embedded and outside conscious control, users cannot easily avoid them through strategic self-presentation, and mechanistic analysis reveals that linguistic features are encoded in early transformer layers and entangled with other features. Our work calls for upstream consideration of the influences of linguistic variation to mitigate disparate impacts of LLM-mediated workplace communication.

语言偏见大模型职场沟通

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。