arXiv:2509.04461cs.CLcs.SI2025-09被引 5

用大模型分析社交媒体文本,精准预测人格类型

From Post To Personality: Harnessing LLMs for MBTI Prediction in Social Media

  • 结合检索增强生成与上下文学习,减少大模型幻觉
  • 通过合成数据扩增平衡各人格类型样本,提升小众类型识别率
  • 在真实社交数据上超越10种主流方法,性能领先

从社交媒体文本中进行人格预测是心理学与社会学中的关键任务。迈尔斯-布里格斯人格类型指标(MBTI)传统上由机器学习与深度学习方法预测。近期大语言模型(LLMs)在理解与推断社交内容人格特征方面展现出巨大潜力。然而,直接使用LLMs进行MBTI预测面临两大挑战:模型固有的幻觉问题,以及人群中共有类型的自然不平衡分布。本文提出PostToPersonality(PtoP)框架,一种基于大模型的社交媒体文本人格预测新方法。PtoP采用检索增强生成结合上下文学习,有效缓解大模型幻觉;同时,通过合成少数类样本对预训练模型进行微调,以平衡类别分布。在真实社交数据集上的实验表明,相比10种主流机器学习与深度学习基线方法,PtoP达到当前最佳性能。

原文摘要 · Abstract (English)

Personality prediction from social media posts is a critical task that implies diverse applications in psychology and sociology. The Myers Briggs Type Indicator (MBTI), a popular personality inventory, has been traditionally predicted by machine learning (ML) and deep learning (DL) techniques. Recently, the success of Large Language Models (LLMs) has revealed their huge potential in understanding and inferring personality traits from social media content. However, directly exploiting LLMs for MBTI prediction faces two key challenges: the hallucination problem inherent in LLMs and the naturally imbalanced distribution of MBTI types in the population. In this paper, we propose PostToPersonality (PtoP), a novel LLM based framework for MBTI prediction from social media posts of individuals. Specifically, PtoP leverages Retrieval Augmented Generation with in context learning to mitigate hallucination in LLMs. Furthermore, we fine tune a pretrained LLM to improve model specification in MBTI understanding with synthetic minority oversampling, which balances the class imbalance by generating synthetic samples. Experiments conducted on a real world social media dataset demonstrate that PtoP achieves state of the art performance compared with 10 ML and DL baselines.

人格预测大模型社交媒体分类平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。