探究语音大模型如何更好识别心理健康状态
Probing mental health information in speech foundation models
- 筛选最适合心理健康检测的预训练任务
- 发现特定模型层对心理状态特征编码效果最佳
- 适用于精神健康研究与临床辅助诊断
非侵入式精神健康诊断方法,如语音分析,为现代医学提供了巨大潜力。近年来,机器学习尤其是语音基础模型在捕捉多种特征方面表现出显著能力,可用于识别精神健康状态。本研究探究了哪些预训练任务最有利于向精神健康检测迁移,并分析不同模型层对相关特征的编码方式。同时,我们考察了音频片段的最佳长度和最优池化策略以提升检测准确率。基于Callyope-GP和Androids数据集,评估了模型在多语言及不同语音任务下的表现,旨在增强基于语音的精神健康诊断泛化能力。该方法在Androids数据集上的抑郁检测任务中达到当前最优(SOTA)性能。
原文摘要 · Abstract (English)
Non-invasive methods for diagnosing mental health conditions, such as speech analysis, offer promising potential in modern medicine. Recent advancements in machine learning, particularly speech foundation models, have shown significant promise in detecting mental health states by capturing diverse features. This study investigates which pretext tasks in these models best transfer to mental health detection and examines how different model layers encode features relevant to mental health conditions. We also probed the optimal length of audio segments and the best pooling strategies to improve detection accuracy. Using the Callyope-GP and Androids datasets, we evaluated the models' effectiveness across different languages and speech tasks, aiming to enhance the generalizability of speech-based mental health diagnostics. Our approach achieved SOTA scores in depression detection on the Androids dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。