揭示大模型中人口身份表征的可分离特性,破解模拟群体差异的迷思。
Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

- 通过注意力头读出比标准残差读出更准确,五类属性相关性达rho=0.63
- 单一注意力头在六类属性中均稳定忠实,但种族类仍脆弱且易受提示干扰
- 模型虽能忠实编码身份,却未真正使用这些信息,因果使用与忠实度不一致
大型语言模型广泛用于模拟调查受访者,但其回答同质化且不忠实于真实群体差异。我们探究了人口群体身份在语言模型中的位置、几何结构对真实群体意见分布的忠实程度,以及模型是否实际使用所编码的信息。基于169个人口细分单元的皮尤研究中心真实数据,对Mistral-7B中1,089个读出位置进行表示相似性分析,并在六类属性上实施因果干预。结果表明:(1) 标准最后令牌残差读出低估了模型表现,注意力头读出在五类属性中占主导地位,校正选择后的相关性最高达rho=0.63(约达到测量可靠性上限的70%),且在词汇相似性控制下仍成立;(2) 单一注意力头(L11 H16)在六类属性中均显著忠实,而种族相关类型始终较弱且对提示敏感;该现象在第二模型家族三个检查点中重复出现,仅十亿训练词元的微调几乎未改变结构;(3) 因果使用并不随忠实度变化:最清晰的因果路径出现在最不忠实的类型中(p=0.002,聚类稳健,固定深度),而最忠实类型无显著单层效应,替换整个身份仅使预测误差变动不足2%;(4) 对该头的128维探测器使距离调查真值缩短21%-31%,但无法恢复各问题的群体排序,效果不优于模型自身回答。可读性、忠实性、因果使用是同一模型中可分离的三种属性,三者混淆正是“语言模型能否模拟人群”争议持续的原因。
原文摘要 · Abstract (English)
Large language models are widely used to simulate survey respondents, yet their answers are homogeneous and unfaithful to real inter-group differences. We ask where demographic group identity lives inside an LLM, how faithfully its geometry mirrors real inter-group opinion structure, and whether it uses what it encodes. Using representational similarity analysis against Pew ground truth over 169 demographic cells, we score 1,089 read-out locations in Mistral-7B and intervene causally across six attribute types. Four results. (1) The standard last-token residual read-out understates the model: attention-head read-outs dominate it in five of six types, with selection-corrected fidelity up to rho=0.63 -- roughly 70% of the measurement-reliability ceiling -- surviving a lexical-similarity control. (2) A single head (L11 H16) is significantly faithful in all six types as a fixed location, while race-based types stay weak and prompt-fragile. Both phenomena replicate -- the analogous head significant in five of six types, weakest on the same race type -- across three checkpoints of a second model family, where ten billion training tokens barely move the map. (3) Causal use does not follow fidelity: the clearest causal pathway sits in one of the least faithful types (p=0.002, cluster-robust, fixed depth), the most faithful type shows no correction-surviving single-layer effect, and replacing the entire identity moves predictions by under 2% of their error. (4) A 128-dimensional probe of the single head lands 21-31% closer to survey truth than the model's own answers -- yet recovers almost none of the per-question group ordering, no better than the answers themselves. Readable, faithfully arranged, and causally used are three dissociable properties of the same model; treating them as one claim is what keeps the "can LLMs simulate populations" debate unresolved.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。