arXiv:2608.18768cs.CL2026-08

揭示大模型中人口身份表征的可分离特性,破解模拟群体差异的迷思。

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

论文配图:Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model
图 1 · 摘自论文原文
  • 通过注意力头读出比标准残差读出更准确,五类属性相关性达rho=0.63
  • 单一注意力头在六类属性中均稳定忠实,但种族类仍脆弱且易受提示干扰
  • 模型虽能忠实编码身份,却未真正使用这些信息,因果使用与忠实度不一致

大型语言模型广泛用于模拟调查受访者,但其回答同质化且不忠实于真实群体差异。我们探究了人口群体身份在语言模型中的位置、几何结构对真实群体意见分布的忠实程度,以及模型是否实际使用所编码的信息。基于169个人口细分单元的皮尤研究中心真实数据,对Mistral-7B中1,089个读出位置进行表示相似性分析,并在六类属性上实施因果干预。结果表明:(1) 标准最后令牌残差读出低估了模型表现,注意力头读出在五类属性中占主导地位,校正选择后的相关性最高达rho=0.63(约达到测量可靠性上限的70%),且在词汇相似性控制下仍成立;(2) 单一注意力头(L11 H16)在六类属性中均显著忠实,而种族相关类型始终较弱且对提示敏感;该现象在第二模型家族三个检查点中重复出现,仅十亿训练词元的微调几乎未改变结构;(3) 因果使用并不随忠实度变化:最清晰的因果路径出现在最不忠实的类型中(p=0.002,聚类稳健,固定深度),而最忠实类型无显著单层效应,替换整个身份仅使预测误差变动不足2%;(4) 对该头的128维探测器使距离调查真值缩短21%-31%,但无法恢复各问题的群体排序,效果不优于模型自身回答。可读性、忠实性、因果使用是同一模型中可分离的三种属性,三者混淆正是“语言模型能否模拟人群”争议持续的原因。

原文摘要 · Abstract (English)

Large language models are widely used to simulate survey respondents, yet their answers are homogeneous and unfaithful to real inter-group differences. We ask where demographic group identity lives inside an LLM, how faithfully its geometry mirrors real inter-group opinion structure, and whether it uses what it encodes. Using representational similarity analysis against Pew ground truth over 169 demographic cells, we score 1,089 read-out locations in Mistral-7B and intervene causally across six attribute types. Four results. (1) The standard last-token residual read-out understates the model: attention-head read-outs dominate it in five of six types, with selection-corrected fidelity up to rho=0.63 -- roughly 70% of the measurement-reliability ceiling -- surviving a lexical-similarity control. (2) A single head (L11 H16) is significantly faithful in all six types as a fixed location, while race-based types stay weak and prompt-fragile. Both phenomena replicate -- the analogous head significant in five of six types, weakest on the same race type -- across three checkpoints of a second model family, where ten billion training tokens barely move the map. (3) Causal use does not follow fidelity: the clearest causal pathway sits in one of the least faithful types (p=0.002, cluster-robust, fixed depth), the most faithful type shows no correction-surviving single-layer effect, and replacing the entire identity moves predictions by under 2% of their error. (4) A 128-dimensional probe of the single head lands 21-31% closer to survey truth than the model's own answers -- yet recovers almost none of the per-question group ordering, no better than the answers themselves. Readable, faithfully arranged, and causally used are three dissociable properties of the same model; treating them as one claim is what keeps the "can LLMs simulate populations" debate unresolved.

语言模型身份表征因果推理群体模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。