基于脑启发机制的音频特征提取器,提升抑郁症诊断精度与抗噪能力。
RBA-FE: A Robust Brain-Inspired Audio Feature Extractor for Depression Diagnosis
- 融合六种声学特征与改进脉冲神经元模型,模拟大脑注意力机制。
- 在MODMA数据集上达到0.8974准确率,各项指标均超现有模型。
- 适合临床辅助诊断、语音分析及脑启发计算研究者使用。
本文提出一种鲁棒的脑启发音频特征提取器(RBA-FE),用于抑郁症诊断,采用改进的分层网络架构。多数深度学习模型在图像诊断任务中表现优异,却忽视了音频特征的应用。为应对噪声挑战,RBA-FE从原始音频中提取六种声学特征,捕捉空间特性与时间依赖性,缓解传统模型如深度残差压缩网络在音频特征提取中的精度瓶颈。模型引入改进的脉冲神经元模型——自适应率平滑漏电积分-发放(ARSLIF),模拟大脑注意系统中“细胞信号选择性重调”机制,增强对环境噪声的鲁棒性。实验表明,RBA-FE在MODMA数据集上分别取得0.8750、0.8974、0.8750和0.8750的精确率、准确率、召回率与F1分数,达到当前最优水平。在AVEC2014和DAIC-WOZ数据集上的大量实验也验证了其抗噪性能优势。对比分析进一步显示,ARSLIF模型揭示了抑郁音频数据中特征提取的异常放电模式,提供脑启发可解释性。
原文摘要 · Abstract (English)
This article proposes a robust brain-inspired audio feature extractor (RBA-FE) model for depression diagnosis, using an improved hierarchical network architecture. Most deep learning models achieve state-of-the-art performance for image-based diagnostic tasks, ignoring the counterpart audio features. In order to tailor the noise challenge, RBA-FE leverages six acoustic features extracted from the raw audio, capturing both spatial characteristics and temporal dependencies. This hybrid attribute helps alleviate the precision limitation in audio feature extraction within other learning models like deep residual shrinkage networks. To deal with the noise issues, our model incorporates an improved spiking neuron model, called adaptive rate smooth leaky integrate-and-fire (ARSLIF). The ARSLIF model emulates the mechanism of ``retuning of cellular signal selectivity" in the brain attention systems, which enhances the model robustness against environmental noises in audio data. Experimental results demonstrate that RBA-FE achieves state-of-the-art accuracy on the MODMA dataset, respectively with 0.8750, 0.8974, 0.8750 and 0.8750 in precision, accuracy, recall and F1 score. Extensive experiments on the AVEC2014 and DAIC-WOZ datasets both show enhancements in noise robustness. It is further indicated by comparison that the ARSLIF neuron model suggest the abnormal firing pattern within the feature extraction on depressive audio data, offering brain-inspired interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。