arXiv:2511.07011cs.CLcs.LG2025-11

分析语音中的词汇特征,可客观预测抑郁严重程度。

Multilingual Lexical Feature Analysis of Spoken Language for Predicting Major Depression Symptom Severity

  • 用线性混合模型识别与抑郁相关的五类可解释词汇特征。
  • 词汇特征和向量嵌入均提升对PHQ-8评分的预测准确率。
  • 适合关注心理健康监测与自然语言分析交叉研究者。

背景:远程采集的语音数据可提供抑郁症状严重程度的客观、定期指标。但现有研究多基于非临床、横断面书面语言,且采用复杂机器学习模型,可解释性差。方法:我们使用线性混合效应模型,分析来自英国、荷兰和西班牙的RADAR-MDD研究数据,该数据包含467名参与者、5,846段智能手机录音及对应的PHQ-8评分。系统评估了可解释的词汇特征与高维向量嵌入在预测PHQ-8评分上的表现,对比其在社会人口学及混杂因素基础上的提升效果。结果:抑郁严重程度与五类词汇特征相关,包括词数减少、第一人称复数代词使用降低及积极词汇频率下降。这些关联在各国间基本稳定,仅积极词汇频率例外。词汇特征与向量嵌入均显著提升预测精度。局限:样本年龄中位数为53岁(IQR 35–62),女性占多数(n=357),可能影响普适性;非英语语言的NLP工具匮乏,限制特征选择。结论:需进一步研究以实现语音词汇标记在临床研究与实践中的价值,包括扩大样本规模、优化采集协议及发展考虑个体内外变异的分析方法。

原文摘要 · Abstract (English)

Background: Remotely captured spoken language could provide objective, regular indicators of depression symptom severity. However, research to date has largely used non-clinical, cross-sectional written language and complex machine learning (ML) approaches with limited interpretability. Methods: We used linear mixed-effect models to identify interpretable lexical features associated with symptom severity in data from the RADAR-MDD study that comprised 5,846 smartphone recordings and Patient Health Questionnaire (PHQ-8) scores from 467 participants in the UK, Netherlands and Spain. We then developed ML models and systematically assessed via nested cross-validation whether interpretable lexical features or high-dimensional vector embeddings improved the accuracy of PHQ-8 prediction over sociodemographic and confounding features. Results: Depression symptom severity was associated with five lexical features, including reductions in word count measures, use of first-person plural pronouns and positive word frequency. Associations were stable across countries, except for positive word frequency. Lexical features and vector embeddings did improve prediction accuracy beyond baseline models. Limitations: Our cohort was skewed in age (median = 53, IQR 35 to 62) and majority female (n=357), potentially affecting the generalizability of our results. A lack of natural language processing tools for non-English languages restricted our feature choices. Conclusion: Further research is required to realise the value of spoken lexical markers in clinical research and practice including larger and more diverse samples, elicitation protocol development and analytical methods that account for within- and between-individual variations.

抑郁症语音分析词汇特征量化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。