arXiv:2510.09739cs.LGcs.AI2025-10

机器学习从语义嵌入中构建人格特质失败,不如大五模型有效。

Machine learning methods fail to provide cohesive atheoretical construction of personality traits from semantic embeddings

  • 用机器学习从词汇表自底向上建模人格
  • 生成的集群无法区分人格,且未恢复外向性特质
  • 适合心理学理论验证,不适合替代已有理论

词汇假说认为人格特质编码于语言,是大五模型的基础。我们基于经典形容词列表,利用机器学习构建了自下而上的个性模型,并通过分析一百万条Reddit评论,将其描述能力与大五模型进行对比。结果显示,大五模型尤其是宜人性、尽责性和神经质维度,对这些在线社区提供了更强大且可解释的描述。相比之下,机器学习聚类未能产生有意义的区分,未能恢复外向性特质,且缺乏大五模型的心理测量一致性。结果证实了大五模型的稳健性,表明人格的语义结构具有情境依赖性。研究说明,尽管机器学习可用于检验既有心理理论的生态效度,但可能无法取代它们。

原文摘要 · Abstract (English)

The lexical hypothesis posits that personality traits are encoded in language and is foundational to models like the Big Five. We created a bottom-up personality model from a classic adjective list using machine learning and compared its descriptive utility against the Big Five by analyzing one million Reddit comments. The Big Five, particularly Agreeableness, Conscientiousness, and Neuroticism, provided a far more powerful and interpretable description of these online communities. In contrast, our machine-learning clusters provided no meaningful distinctions, failed to recover the Extraversion trait, and lacked the psychometric coherence of the Big Five. These results affirm the robustness of the Big Five and suggest personality's semantic structure is context-dependent. Our findings show that while machine learning can help check the ecological validity of established psychological theories, it may not be able to replace them.

人格建模机器学习大五模型语义嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。