arXiv:2501.19337cs.CLcs.CV2025-01

名字影响模型生成,黑人姓名引发更多元输出

Token-Level Entropy Reveals Demographic Disparities in Large Language Models

  • 用名字触发模型预测分布,测量熵值揭示偏见
  • 黑人名对应更高熵和更多样输出,男女效应相反
  • 方法设计决定能否发现偏见及方向,非固定存在

仅通过名字就能显著改变语言模型的下一个词分布,而无需采样任何词。我们在5,760个句子补全提示中,使用仅含首名暗示种族与性别的任务,对六种开源模型家族进行测试。黑人关联姓名对应的首个词熵更高、后续生成更多样化,该趋势在所有六种指令微调模型、六种基础检查点中一致出现;在五种模型的原生聊天格式下,输出多样性也呈现相同趋势——这与显式群体标签下的同质化偏差(Lee et al., 2024)相反。该差异在分词与频率控制后仍存在,且在频率匹配的姓名子集中依然显著。每提示层面效应较小(d = 0.06-0.16),但符号统一(模板级配对效应量 d = 0.66-1.08)。性别效应方向相反,且与种族效应叠加。聊天格式下首词熵急剧下降,显式群体标签探测大多无结果或反向;方差匹配分析表明,输出多样性差异源于名称条件下的延续多样性——这是固定标签无法表达的维度。探测方法不仅决定是否发现偏见,还决定其方向。

原文摘要 · Abstract (English)

A name alone measurably reshapes a language model's next-token distribution before a single token is sampled. We measure full-vocabulary Shannon entropy of the next-token distribution across six open-weight model families on 5,760 sentence-completion prompts in which race and gender are signaled only by a first name. Black-associated names co-occur with higher first-token entropy and more diverse continuations than White-associated names -- directionally consistent in all six instruction-tuned models under shared raw-text input, all six base checkpoints, and, for output diversity, five of six models under native chat formatting -- opposite to the homogeneity bias documented under explicit group labels (Lee et al., 2024). The gap persists under tokenization and frequency controls and on a frequency-matched name subset; per-prompt effects are small (d = 0.06-0.16) but uniformly signed (template-level paired d = 0.66-1.08). Gender points the other way, additively with race. First-token entropy attenuates sharply under chat-formatted input, and explicit group-label probing is mostly null or reversed; a variance-matched comparison locates the output-diversity disparity in heterogeneity across name-conditioned continuations -- a dimension a fixed group label cannot express. Probing methodology shapes not only whether a disparity is detected but which direction it takes.

大模型偏见熵分析社会公平命名效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。