用记忆增强提升语音抑郁程度估计,效果优于现有方法。
MA-DLE: Speech-based Automatic Depression Level Estimation via Memory Augmentation

- 引入记忆模块,选择性融合历史特征与动态情绪变化信息。
- 在DAIC-WOZ和E-DAIC数据集上达到当前最优准确率。
- 适合关注心理健康检测与语音分析的开发者与研究者。
基于语音的抑郁症水平自动评估对实现早期发现和及时干预至关重要,尤其在资源有限的心理健康环境中。近年来,深度学习在情感计算与心理健康评估等领域表现出色。现有方法多依赖于基于RNN的架构(如LSTM和GRU)建模时序信息进行抑郁程度估计,但提取的特征常仅关注少数相邻语音片段,难以捕捉长程依赖。为克服此局限,我们提出一种基于记忆的特征增强方法,以提升GRU提取特征的表征能力。记忆库并非盲目整合历史数据,而是有选择地融合两类成分:(1) 与当前GRU输出高度相似的历史时序特征,提供互补上下文;(2) 基于特征变异性的动态记忆特征,捕捉反映抑郁症状的行为与情绪波动。为有效融合记忆增强特征与GRU输出,我们进一步设计了分层注意力融合(HAF)模块。该方法在广泛使用的DAIC-WOZ和E-DAIC数据集上进行了评估,表现达到当前最优水平。
原文摘要 · Abstract (English)
Speech-based automatic estimation of depression levels is essential for enabling early detection and timely intervention, particularly in resource-constrained mental health settings. In recent years, deep learning has demonstrated impressive success across various domains, including affective computing and mental health assessment. Most existing approaches rely on RNN-based architectures (such as LSTM and GRU) to model temporal information for depression estimation. However, the extracted features often emphasize only a few adjacent speech segments, limiting their ability to capture long-range dependencies. To overcome this limitation, we introduce a memory-based feature augmentation method that enhances the representational capacity of GRU-extracted features. Rather than indiscriminately incorporating historical data, our memory bank is designed to selectively integrate two types of components in order to reduce redundancy and irrelevance: (1) historical temporal features that closely resemble the current GRU output, offering complementary contextual information; and (2) dynamic memory features identified based on feature variability, which capture behavioral and emotional fluctuations indicative of depressive symptoms. To effectively fuse the memory-augmented features with GRU outputs, we further design a Hierarchical Attention Fusion (HAF) module. Our method is evaluated on the widely used DAIC-WOZ and E-DAIC datasets, achieving state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。