用可解释AI分析文本心理特征,提升精神健康预测的可信度。
Natural Language Processing Psychometrics

- 通过控制人格的LLM生成文本,构建心理量表评分与语言特征的关联。
- 模型对生活满意度解释率达70.8%,抑郁、焦虑、压力分别达68.5%~76.0%。
- 无需重新训练即可区分高/低分人格日记,临床与对照组分类准确率68%。
自然语言处理(NLP)模型预测心理健康常不明确其测量维度:是上下文知识、情感内容还是句法结构?本文提出NLP Psychometrics,将文本心理预测视为心理测量问题,建立评分与可解释语言证据的联系,并超越训练文本格式进行验证。九个大语言模型(LLMs)在受控人格(认知数字影子)设定下完成心理量表并提供逐项文本解释。通过文本思维网络提取情感特征与句法-语义结构,结合人格与社会人口学变量,使用剪枝随机森林(RF)回归器建模,并用SHAP分析各特征贡献方向。完整模型对生活满意度(SWLS)解释方差达70.8%,抑郁(PHQ-9)55.7%,DASS-21中抑郁68.5%、焦虑76.0%、压力72.4%。仅社会人口学变量无法解释抑郁、焦虑或压力,但可解释生活满意度,其中情绪特征与收入为强预测因子;神经质和网络拓扑则主导抑郁与焦虑,且方向相反。未重新训练的模型可区分低/高分人格日记(相关系数r达0.91),仅用网络与情绪特征即在真实语篇中实现临床与对照者分类,最高准确率68%。结果表明合成数据有潜力暴露模型偏见、复现临床反刍模式,支持无匹配问卷的人类文本心理预测,但不可替代人工验证。NLP Psychometrics通过可解释人工智能与网络/情绪特征,使这些差异变得明确、可测、可验。
原文摘要 · Abstract (English)
Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure. NLP Psychometrics treats psychological prediction from text as a psychometric problem, linking scores to interpretable linguistic evidence and testing beyond the training text format. Nine LLMs, conditioned on controlled personas (cognitive digital shadows), completed psychometric questionnaires with textual explanations per item. We extracted emotional profiles and syntactic-semantic structure via textual forma mentis networks, combined with personality and sociodemographic variables in ablated random forest (RF) regressors, using SHAP to identify which features drove performance and in which direction. Full RF models explained up to 70.8% of variance in life satisfaction (SWLS), 55.7% in depression (PHQ-9), and, for DASS-21, 68.5% depression, 76.0% anxiety, 72.4% stress. Sociodemographics alone explained no meaningful variance in depression, anxiety, or stress, but did so for life satisfaction, where emotion features and income were the strongest predictors; neuroticism and network topology instead dominated depression and anxiety, reversing direction between them. Without retraining, RF models separated diaries from low- and high-score personas ($r$ up to 0.91) and, using only network/emotion features, classified clinical from control participants in real transcripts with up to 68% accuracy. These results show the promise and limits of synthetic data: LLM personas can expose model biases, recover patterns consistent with clinical rumination, and support psychometric prediction from human text without a matched questionnaire, but cannot substitute for human validation. NLP Psychometrics makes these distinctions explicit, measurable, and testable through interpretable AI and network/emotional features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。