用心理投射测试评估大模型的人格特质,发现其懂人际关系但不懂控制攻击性。
Projective Psychological Assessment of Large Multimodal Models Using Thematic Apperception Tests
- 让大模型看图编故事并自我评估,用心理学量表分析其人格特征。
- 所有模型都理解人际互动和自我概念,但均无法识别与调节攻击性。
- 越新越大的模型表现越好,说明能力随规模提升而增强。
主题统觉测验(TAT)是一种基于心理测量学的多维度评估框架,可系统区分人格功能中的认知表征与情感关系成分。本研究探讨大型多模态模型(LMMs)是否可通过非语言模态进行人格特质评估,采用社会认知与客体关系量表-全局版(SCORS-G)。LMMs扮演两种角色:作为生成故事的主体模型(SMs),根据TAT图像生成叙事;作为评估者模型(EMs),使用SCORS-G框架分析这些叙事。评估者表现出极强的理解与分析能力,其判断与人类专家高度一致。结果显示,所有模型均良好理解人际动态并具备自我概念,但持续无法感知与调节攻击性。性能在不同模型家族间系统性差异,更大、更先进的模型在各SCORS-G维度上均显著优于小型和早期模型。
原文摘要 · Abstract (English)
Thematic Apperception Test (TAT) is a psychometrically grounded, multidimensional assessment framework that systematically differentiates between cognitive-representational and affective-relational components of personality-like functioning. This test is a projective psychological framework designed to uncover unconscious aspects of personality. This study examines whether the personality traits of Large Multimodal Models (LMMs) can be assessed through non-language-based modalities, using the Social Cognition and Object Relations Scale - Global (SCORS-G). LMMs are employed in two distinct roles: as subject models (SMs), which generate stories in response to TAT images, and as evaluator models (EMs), who assess these narratives using the SCORS-G framework. Evaluators demonstrated an excellent ability to understand and analyze TAT responses. Their interpretations are highly consistent with those of human experts. Assessment results highlight that all models understand interpersonal dynamics very well and have a good grasp of the concept of self. However, they consistently fail to perceive and regulate aggression. Performance varied systematically across model families, with larger and more recent models consistently outperforming smaller and earlier ones across SCORS-G dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。