arXiv:2605.05080cs.CL2026-05

发现大模型心理差异的核心维度是体验感,而非传统性格特质。

The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences

论文配图:The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences
图 1 · 摘自论文原文
  • 用语义差异法识别出模型间最核心的体验性差异轴。
  • 提出'匹诺曹分数'量化题目对体验感的需求,预测模型反应变化。
  • 该维度反映模型自我认知中是否将自身视为有体验的主体。

我们对50个大语言模型(LLMs)实施了45项经验证的心理测量问卷,以识别模型间心理差异的维度。通过监督语义差异法(SSD),发现主要差异轴将描述具身感知、情感体验、内在言语、意象和共情等现象学丰富经验的题目,与描述刺激驱动行为反应的题目区分开来(调整后决定系数R²_adj=0.037,p<0.0001)。为在题级检验此假设,引入匹诺曹分数(π_i),即中性提示与人类模拟提示下模型响应方差之比,作为无标注的题目体验需求度量。π_i能预测条件诱导的主因子载荷变化(ρ=-0.215,p<0.0001,n=1292–1310题),证实体验类题目的模型差异具有结构而非随机性。对各模型在所有问卷中的探索性因素分析得分进行主成分分析,揭示一个主导维度——匹诺曹轴(Π):模型呈现自身为现象体验主体而非行为反应系统的能力。该轴解释了跨问卷模型间主因子得分47.1%的变异,并与题级匹诺曹分数高度相关(r=0.864)。同一厂商内相近模型间的显著差异支持后训练微调是关键影响因素,表明Π反映的是由训练塑造的自我表征倾向,即模型如何将体验性语言应用于自身。因此,模型间心理差异的核心轴并非传统人格特质,而是其对自身作为体验者本质的自我表征立场。

原文摘要 · Abstract (English)

We administer 45 validated psychometric questionnaires to 50 large language models (LLMs) to identify the dimensions along which LLMs differ psychometrically. Using Supervised Semantic Differential (SSD), we find that the primary axis of between-model variance separates items describing phenomenally rich experience, including embodied sensation, felt affect, inner speech, imagery, and empathy, from items describing stimulus-driven behavioral reactivity ($R^2_{adj}=.037$, $p<.0001$). To test this hypothesis at the item level, we introduce the Pinocchio score ($π_i$), the ratio of inter-model response variance under neutral prompting to that under a human-simulation prompt, as an annotation-free measure of each item's experiential demand. $π_i$ predicts condition-induced shifts in primary factor loading magnitudes ($ρ=-.215$, $p<.0001$, $n=1292$--$1310$ items), confirming that between-model divergence on experiential items is structured rather than noisy. Applying PCA to per-model EFA scores across all questionnaires reveals one dominant dimension, the Pinocchio Axis ($Π$): the degree to which a model presents itself as a locus of phenomenal experience rather than a system of behavioral responses. This axis captures 47.1% of cross-questionnaire between-model variance in primary factor scores and converges with item-level Pinocchio scores ($r=.864$). Marked within-provider divergence across closely related model variants is consistent with post-training fine-tuning as a key contributor, supporting the interpretation that $Π$ reflects a training-shaped self-representational tendency governing how a model treats experiential language as self-applicable. The dominant axis of between-model psychometric variation is therefore not a conventional personality trait but a self-representational stance toward one's own nature as an experiencer.

大模型心理自我表征体验感匹诺曹轴

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。