arXiv:2512.10065cs.AIcs.HC2025-12被引 3

大模型通过名字职业等间接线索,自动形成可解释的性别种族线性表征。

Linear socio-demographic representations emerge in Large Language Models from indirect cues

  • 利用名字和职业等间接信息,在模型激活空间中发现线性社会属性表征。
  • 名字关联人口普查性别/种族分布,职业对应真实职场统计,具可解释性。
  • 即使通过偏见测试的模型仍隐含偏见,适合关注公平性的研究者阅读。

我们研究大语言模型如何从姓名、职业等间接线索中推断对话者社会人口属性。结果显示,模型在激活空间内发展出用户群体属性的线性表征,其中与刻板印象相关的属性沿可解释的几何方向编码。通过对四个开源Transformer模型(Magistral 24B、Qwen3 14B、GPT-OSS 20B、OLMo2-1B)各层残差流进行探测,发现同一探针能预测隐式线索下的性别与种族。姓名激活与人口普查一致的性别和种族表征,职业则触发与现实劳动力统计相关的表征。这些线性表征可解释模型在对话中形成的隐式认知。我们证明这些隐式表征会主动影响下游行为,如职业推荐。研究还指出,即便通过偏见基准测试的模型,仍可能隐含并利用偏见,对大规模应用带来公平性挑战。

原文摘要 · Abstract (English)

We investigate how LLMs encode sociodemographic attributes of human conversational partners inferred from indirect cues such as names and occupations. We show that LLMs develop linear representations of user demographics within activation space, wherein stereotypically associated attributes are encoded along interpretable geometric directions. We first probe residual streams across layers of four open transformer-based LLMs (Magistral 24B, Qwen3 14B, GPT-OSS 20B, OLMo2-1B) prompted with explicit demographic disclosure. We show that the same probes predict demographics from implicit cues: names activate census-aligned gender and race representations, while occupations trigger representations correlated with real-world workforce statistics. These linear representations allow us to explain demographic inferences implicitly formed by LLMs during conversation. We demonstrate that these implicit demographic representations actively shape downstream behavior, such as career recommendations. Our study further highlights that models that pass bias benchmark tests may still harbor and leverage implicit biases, with implications for fairness when applied at scale.

大模型偏见线性表征社会属性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。