用社交媒体文本做性格评估,既准又可解释。
A Computational Framework for Interpretable Text-Based Personality Assessment from Social Media
- 通过匹配用户语句与标准问卷项,实现可解释的性格分析
- 基于1700万条评论的PANDORA数据集支持多模型验证
- 框架通用性强,适合复杂标签体系的智能分析
人格指个体在行为、思维和情感上的差异。随着数字足迹(尤其是社交媒体)的日益丰富,自动化人格评估方法变得愈发重要。自然语言处理(NLP)使对非结构化文本的分析成为可能,从而识别人格特征。但两大挑战仍存:大规模带人格标签的数据集稀缺,以及人格心理学与NLP之间的脱节,制约了模型的有效性与可解释性。为此,本文构建了两个数据集——MBTI9k与PANDORA,均来自以匿名性和多样性著称的Reddit平台。其中,PANDORA包含超过10,000名用户的1700万条评论,整合了MBTI与大五人格模型及人口统计信息,解决了数据量、质量与标签覆盖的局限。实验表明,人口统计变量影响模型有效性。为此提出SIMPA(Statement-to-Item Matching Personality Assessment)框架——一种可解释的性格评估计算框架,通过机器学习与语义相似性匹配用户生成语句与标准化问卷项。实验证明,SIMPA的评估结果与人工评估相当,同时保持高可解释性与效率。尽管聚焦于人格评估,其模型无关设计、分层线索检测与可扩展性使其适用于涉及复杂标签体系和可变线索关联的多种研究与实际场景。
原文摘要 · Abstract (English)
Personality refers to individual differences in behavior, thinking, and feeling. With the growing availability of digital footprints, especially from social media, automated methods for personality assessment have become increasingly important. Natural language processing (NLP) enables the analysis of unstructured text data to identify personality indicators. However, two main challenges remain central to this thesis: the scarcity of large, personality-labeled datasets and the disconnect between personality psychology and NLP, which restricts model validity and interpretability. To address these challenges, this thesis presents two datasets -- MBTI9k and PANDORA -- collected from Reddit, a platform known for user anonymity and diverse discussions. The PANDORA dataset contains 17 million comments from over 10,000 users and integrates the MBTI and Big Five personality models with demographic information, overcoming limitations in data size, quality, and label coverage. Experiments on these datasets show that demographic variables influence model validity. In response, the SIMPA (Statement-to-Item Matching Personality Assessment) framework was developed - a computational framework for interpretable personality assessment that matches user-generated statements with validated questionnaire items. By using machine learning and semantic similarity, SIMPA delivers personality assessments comparable to human evaluations while maintaining high interpretability and efficiency. Although focused on personality assessment, SIMPA's versatility extends beyond this domain. Its model-agnostic design, layered cue detection, and scalability make it suitable for various research and practical applications involving complex label taxonomies and variable cue associations with target concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。