用语音特征辅助精神健康评估,结果可解释且稳定
Exploration of Perceptual Speech Features for Clinical Decision-Support in Mental Health Care

- 基于语音的声学与语言特征分析框架,涵盖语调、语义连贯性等
- 发现语音不规则度与抑郁、焦虑、多动症症状严重程度相关
- 适合临床医生和研究者用于可解释的精神健康辅助诊断
语音与语言技术为精神健康评估提供了客观可解释的线索。本文提出一个系统性的基于特征的分析框架,利用感知基础的声学与语言特征,包括语调、声音质量、语义连贯性、句法结构和反讽表达。通过统计分析与可解释机器学习(XGBoost结合SHAP和LIME),研究语音特征与抑郁症、焦虑症及注意力缺陷多动障碍(ADHD)标准化症状量表之间的关联。在多个受控基准数据集(StressID、DAIC-WOZ、Androids、EATD)和真实临床数据集上验证,框架揭示了症状严重程度与声音不规则性(如颤音、抖动)、词汇-句法模式及情感基调之间稳定且一致的关系。跨所有数据集的消融研究进一步识别出最具信息量的特征组。本工作探索了一种透明且临床可解释的语音辅助精神健康分析方法。
原文摘要 · Abstract (English)
Speech and language technologies offer valuable opportunities for supporting mental health assessment through objective and interpretable cues. We present a systematic feature-based analysis framework leveraging perceptually grounded acoustic and linguistic characteristics, including prosody, vocal quality, semantic coherence, syntactic structure, and sarcasm. Using statistical analysis and interpretable machine learning (XGBoost with SHAP and LIME), we examine associations between speech features and validated symptom measures of depression, anxiety, and ADHD. Evaluated on both controlled benchmark datasets (StressID, DAIC-WOZ, Androids, EATD) and a real-world clinical dataset, the framework reveals stable and consistent relationships between symptom severity and vocal irregularities (e.g., shimmer, jitter), lexical-syntactic patterns, and affective tone. An ablation study conducted across all datasets further identifies the most informative feature groups. This work explores a transparent and clinically interpretable approach to speech-based mental health analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。