通过频段感知增强网络,提升语音评估抑郁和多动症的准确性
A Frequency-aware Augmentation Network for Mental Disorders Assessment from Audio
- 基于频谱图输入,用多尺度卷积聚焦与心理疾病相关的频段
- 动态卷积捕捉时变特征,实现9.23的抑郁程度估计RMSE和89.8%的多动症检测准确率
- 适合从事情感计算、精神健康智能评估的研究者与开发者
抑郁症和注意力缺陷多动障碍(ADHD)是当前常见的心理健康问题。在情感计算中,语音信号是精神障碍评估的有效生物标志物。现有研究依赖人工设计特征或简单的时频表示,常因未考虑不同频段和时间波动的差异性影响而忽略关键细节。为此,我们提出一种频段感知增强网络,结合动态卷积用于抑郁和ADHD评估。该方法以频谱图为输入,采用多尺度卷积使网络关注与精神疾病相关的判别性频段;设计动态卷积,基于非输入依赖的注意力机制动态聚合多个卷积核以捕捉动态信息;最后引入特征增强模块,强化特征表示能力并充分利用提取信息。在AVEC 2014和自采集的ADHD数据集上的实验表明,该方法具有鲁棒性,抑郁严重程度估计的RMSE达到9.23,ADHD检测准确率达89.8%。
原文摘要 · Abstract (English)
Depression and Attention Deficit Hyperactivity Disorder (ADHD) stand out as the common mental health challenges today. In affective computing, speech signals serve as effective biomarkers for mental disorder assessment. Current research, relying on labor-intensive hand-crafted features or simplistic time-frequency representations, often overlooks critical details by not accounting for the differential impacts of various frequency bands and temporal fluctuations. Therefore, we propose a frequency-aware augmentation network with dynamic convolution for depression and ADHD assessment. In the proposed method, the spectrogram is used as the input feature and adopts a multi-scale convolution to help the network focus on discriminative frequency bands related to mental disorders. A dynamic convolution is also designed to aggregate multiple convolution kernels dynamically based upon their attentions which are input-independent to capture dynamic information. Finally, a feature augmentation block is proposed to enhance the feature representation ability and make full use of the captured information. Experimental results on AVEC 2014 and self-recorded ADHD dataset prove the robustness of our method, an RMSE of 9.23 was attained for estimating depression severity, along with an accuracy of 89.8\% in detecting ADHD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。