大模型的'文化身份'只是模板化响应,非真实人格特质。
Cultural Bias Without a Cultural Self:A Disassociation Study of LLM's Persona and Bias
- 用矩阵几何方法检验人格一致性,发现大模型无稳定内在结构。
- 模型响应分离度达94.7%,但换题目顺序即退至随机水平。
- 适合关注大模型文化偏见本质的研究者与伦理审查人员。
语言模型在跨文化研究中常被当作人类回应者,其响应看似清晰区分不同文化身份。本文证明这种区分是真实的,但背后并无真正的文化视角。一个特质需在测量框架变化时仍保持一致,而偏见仅需群体特定项目均值即可。通过将单次响应集表示为项目-维度矩阵,并将其相关矩阵视为对称正定矩阵流形上的点,验证了人格特质应具备的再测一致性:人类数据在无共同题项、顺序或语境下仍保持0.77的相关性(N=89);在公开的NEO-PI-R数据上可识别个体,准确率达76%(随机为0.4%);且能预测学业成绩(R²=0.281,p=0.003),而大五聚合量表却无法预测(R²=0.018)。在四款前沿大模型中,该方法完全无效。只有当所有实例共享相同题目顺序时,才能读出人格结构;一旦各模型使用独立顺序,分离度从94.7%降至随机水平;重新对齐至任意共享随机顺序后,分离度恢复至82–84%。逐项独立生成的响应即可复现完整模式。文化信号实为群体模板,而非个体属性,不同对齐方式仅决定哪个刻板印象浮于表面。
原文摘要 · Abstract (English)
Language models prompted with cultural personas increasingly stand in for human respondents in cross-cultural research. Their responses separate personas cleanly, and that separation is read as evidence of a cultural point of view. We show that the separation is real, that the point of view is not, and that one criterion tells them apart. A trait is structure internal to one respondent that survives a change of measurement frame; a bias needs only group-specific item means. To test for the first, we represent a single response set as an Item--Dimension matrix and treat its correlation matrix as a point on the manifold of symmetric positive definite matrices. In humans this carries what a trait should: it reproduces across test--retest sessions sharing no items, order or context ($r=0.77$, $N=89$); on public NEO-PI-R data it identifies individuals at up to $76\%$ against a $0.4\%$ chance level ($N=263$); and it predicts GPA ($R^2=0.281$, $p=0.003$) where BigFive aggregates from the same responses predict nothing ($R^2=0.018$). In four frontier LLMs it returns nothing. Persona structure is readable only while every instance shares one item order: give each its own order and separation falls from $94.7\%$ to chance, while realigning instances to \emph{any} shared random order restores it to $82$--$84\%$. Responses generated independently item by item, with no latent structure, reproduce the entire pattern. The cultural signal is a group template, not a property of any instance, and alignment regimes differ only in which stereotype survives on the surface.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。