用语音声学特征和机器学习评估自杀风险,发现关键语音差异并提升识别准确率。
Acoustic and Machine Learning Methods for Speech-Based Suicide Risk Assessment: A Systematic Review
- 分析33篇论文的语音声学特征,聚焦抖动、基频、梅尔系数等指标差异
- 分类器AUC最高达0.985,准确率可达99.85%,多模态融合效果更优
- 适合心理健康研究者、临床筛查系统开发者参考,尤其关注语音分析应用
自杀仍是重大公共卫生挑战,亟需改进检测方法以实现及时干预。本系统综述依据PRISMA指南,从PubMed、Cochrane、Scopus和Web of Science数据库中筛选出33篇文献,最后一次检索时间为2025年2月。采用PROBAST工具评估偏倚风险。纳入研究比较有自杀风险(RS)与无风险(NRS)人群在语音声学特征上的差异,排除缺乏语音数据、未聚焦自杀或方法描述不足的研究。样本量差异较大,按参与者或语音片段报告。结果基于声学特征和分类器性能进行叙述性综合。结果显示,RS与NRS群体在抖动(jitter)、基频(F0)、梅尔频率倒谱系数(MFCC)和功率谱密度(PSD)等特征上存在显著差异。分类器表现受算法、模态及语音诱导方式影响,融合声学、语言和元数据的多模态方法表现更优。在29项基于分类器的研究中,报告的AUC值范围为0.62至0.985,准确率在60%至99.85%之间。多数数据集偏向非风险组,且很少分组报告性能指标,难以明确效应方向。
原文摘要 · Abstract (English)
Suicide remains a public health challenge, necessitating improved detection methods to facilitate timely intervention and treatment. This systematic review evaluates the role of Artificial Intelligence (AI) and Machine Learning (ML) in assessing suicide risk through acoustic analysis of speech. Following PRISMA guidelines, we analyzed 33 articles selected from PubMed, Cochrane, Scopus, and Web of Science databases. The last search was conducted in February 2025. Risk of bias was assessed using the PROBAST tool. Studies analyzing acoustic features between individuals at risk of suicide (RS) and those not at risk (NRS) were included, while studies lacking acoustic data, a suicide-related focus, or sufficient methodological details were excluded. Sample sizes varied widely and were reported in terms of participants or speech segments, depending on the study. Results were synthesized narratively based on acoustic features and classifier performance. Findings consistently showed significant acoustic feature variations between RS and NRS populations, particularly involving jitter, fundamental frequency (F0), Mel-frequency cepstral coefficients (MFCC), and power spectral density (PSD). Classifier performance varied based on algorithms, modalities, and speech elicitation methods, with multimodal approaches integrating acoustic, linguistic, and metadata features demonstrating superior performance. Among the 29 classifier-based studies, reported AUC values ranged from 0.62 to 0.985 and accuracies from 60% to 99.85%. Most datasets were imbalanced in favor of NRS, and performance metrics were rarely reported separately by group, limiting clear identification of direction of effect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。