用原始音频时域信号高效识别抑郁程度,避免频域转换损失。
Efficient Long Speech Sequence Modelling for Time-Domain Depression Level Estimation
- 采用状态空间模型与双路径结构处理长语音序列
- 在AVEC2013/2014上超越当前最优方法
- 适合需要真实对话场景建模的临床情绪分析
抑郁症显著影响情绪、思维和日常活动。研究表明语音信号包含抑郁程度的关键线索,推动了基于音频的深度学习方法发展。然而,现有方法多依赖时频表示,因傅里叶变换和梅尔尺度转换导致信息丢失;且将真实语音分割为短段落会破坏连续性,难以反映抑郁者常有的停顿与语速变慢等特征。为此,本文提出一种基于时域长语音信号的高效抑郁程度估计方法。该方法结合状态空间模型、基于双路径结构的长序列建模模块与时间外部注意力模块,重建并增强原始波形中隐藏的抑郁线索。在AVEC2013和AVEC2014数据集上的实验表明,该方法能有效捕捉关键长序列特征,在性能上优于现有最先进方法。
原文摘要 · Abstract (English)
Depression significantly affects emotions, thoughts, and daily activities. Recent research indicates that speech signals contain vital cues about depression, sparking interest in audio-based deep-learning methods for estimating its severity. However, most methods rely on time-frequency representations of speech which have recently been criticized for their limitations due to the loss of information when performing time-frequency projections, e.g. Fourier transform, and Mel-scale transformation. Furthermore, segmenting real-world speech into brief intervals risks losing critical interconnections between recordings. Additionally, such an approach may not adequately reflect real-world scenarios, as individuals with depression often pause and slow down in their conversations and interactions. Building on these observations, we present an efficient method for depression level estimation using long speech signals in the time domain. The proposed method leverages a state space model coupled with the dual-path structure-based long sequence modelling module and temporal external attention module to reconstruct and enhance the detection of depression-related cues hidden in the raw audio waveforms. Experimental results on the AVEC2013 and AVEC2014 datasets show promising results in capturing consequential long-sequence depression cues and demonstrate outstanding performance over the state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。