人类想象网络具有一致结构,大模型无法复现这种心理组织模式。
Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate
- 用心理网络分析对比人类与大模型的想象结构
- 人类跨群体中心性相关性达0.31-0.93,模型几乎不匹配
- 模型缺乏多簇结构,反映其缺乏具身经验积累
视觉想象生动性是稳定的个体特质,但想象场景在人类与合成大语言模型(LLM)群体间是否具有相似关系结构尚不清楚。我们对两个经验证的问卷——视觉意象生动性量表(VVIQ-2)和普利茅斯感官意象问卷(PSIQ)——的生动性评分进行了心理网络分析,覆盖三个地理语言各异的人类样本(佛罗里达、波兰、伦敦;总人数N=2,743)及六种大语言模型(Gemma3-12B/27B及其量化版本、Llama3.3-70B、Llama4-16x17B)。通过正则化偏相关图构建想象网络,使用皮尔逊相关系数和调整兰德指数(ARI)比较不同群体间的节点中心性与社区结构。人类网络在预期影响、强度和接近度中心性上表现出稳健的跨群体相关性(r = 0.31–0.93),社区检测成功识别出与VVIQ-2场景语境(ARI = 0.27–0.40)及PSIQ感官模态(ARI = 0.87–1.0)一致的聚类。介数中心性在所有群体中均不稳定,与其对个人经验史的敏感性一致。而大模型未能复制人类网络结构:模型与人类中心性相关性弱且大多经校正后不显著,多数模型配置呈现退化的单簇拓扑(中位ARI = 0)。该失败在不同模型架构、参数规模(12B–272B)和对话条件下均一致。我们认为,这些发现可能源于人类想象网络反映了通过具身经验积累的记忆组织结构,而仅靠语言训练无法还原,无论模型规模或多轮对话记忆如何。
原文摘要 · Abstract (English)
Mental imagery vividness is a stable individual trait, yet whether imagined scenarios share relational structure across human and synthetic large language model (LLM) populations remains unknown. We applied psychological network analysis to vividness ratings from two validated questionnaires: the Vividness of Visual Imagery Questionnaire (VVIQ-2) and the Plymouth Sensory Imagery Questionnaire (PSIQ), across geographically and linguistically distinct human samples (Florida, Poland, and London; total N = 2,743) and six large language models (LLMs; Gemma3-12B/27B, their quantization-aware counterparts, Llama3.3-70B, and Llama4-16x17B). Imagination networks were constructed as regularized partial correlation graphs, with node centrality and community structure compared across populations using Pearson correlations and the Adjusted Rand Index (ARI). Human networks showed robust cross-population centrality correlations for expected influence, strength, and closeness (r = 0.31-0.93), and community detection recovered clusters aligned with VVIQ-2 scene contexts (ARI = 0.27-0.40) and PSIQ sensory modalities (ARI = 0.87-1.0). Betweenness centrality was unstable across all populations, consistent with its sensitivity to individual experiential history. LLMs failed to replicate human network structure: LLM-human centrality correlations were weak and largely non-significant after correction, and most LLM configurations produced degenerate single-cluster topologies (median ARI = 0). This failure was consistent across model architectures, parameter scales (12B-272B), and conversational conditions. We posit that these findings may be driven by human imagination networks reflecting memory organization accumulated through embodied experience, a representational structure that linguistic training alone does not reproduce regardless of model scale and conversational memory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。