研究语音输入时长对抑郁分类的影响,优化人机筛查流程
Optimizing Speech-Input Length for Speaker-Independent Depression Classification
- 分析1400小时语音数据,发现自然时长和响应顺序影响分类效果
- 表现更好的系统在更长响应下仍有效,且存在最优饱和阈值
- 建议在响应饱和时换问题,提升抑郁检测准确率
基于语音的抑郁分类机器学习模型在医疗应用中前景广阔。尽管相关研究不断增多,但语音输入时长如何影响模型性能仍不明确。本研究基于超过1400小时的人机健康筛查语音数据,分析两种不同性能的NLP系统在说话人无关抑郁分类任务中的表现。结果表明,模型性能受自然时长、已用时长及响应在会话中的顺序影响。两类系统均存在最小输入长度阈值,但表现更优的系统具有更高的响应饱和阈值;当达到饱和时,提出新问题比继续当前回应更有利于分类。这些发现为设计更高效的语音输入采集与处理策略提供了依据。
原文摘要 · Abstract (English)
Machine learning models for speech-based depression classification offer promise for health care applications. Despite growing work on depression classification, little is understood about how the length of speech-input impacts model performance. We analyze results for speaker-independent depression classification using a corpus of over 1400 hours of speech from a human-machine health screening application. We examine performance as a function of response input length for two NLP systems that differ in overall performance. Results for both systems show that performance depends on natural length, elapsed length, and ordering of the response within a session. Systems share a minimum length threshold, but differ in a response saturation threshold, with the latter higher for the better system. At saturation it is better to pose a new question to the speaker, than to continue the current response. These and additional reported results suggest how applications can be better designed to both elicit and process optimal input lengths for depression classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。