改进的ECAPA-TDNN模型提升婴儿哭声情绪识别准确率
Infant Cry Emotion Recognition Using Improved ECAPA-TDNN with Multiscale Feature Fusion and Attention Enhancement
- 融合多尺度特征与注意力机制,增强时频关系建模
- 在公开数据集上达到82.20%准确率,参数仅1.43MB
- 适合智能育儿与新生儿医疗场景应用
婴儿哭声情绪识别对育儿和医疗应用至关重要,但面临情绪细微差异、噪声干扰和数据有限等挑战。现有方法难以有效融合多尺度特征与时频关系。本文提出一种改进的强调通道注意力、传播与聚合的时间延迟神经网络(ECAPA-TDNN),结合多尺度特征融合与注意力增强。在公开数据集上的实验表明,该方法实现82.20%的准确率,参数量为1.43 MB,计算量为0.32 Giga FLOPs。相比基线方法,本方法在准确率上具有明显优势。代码已开源:https://github.com/kkpretend/IETMA。
原文摘要 · Abstract (English)
Infant cry emotion recognition is crucial for parenting and medical applications. It faces many challenges, such as subtle emotional variations, noise interference, and limited data. The existing methods lack the ability to effectively integrate multi-scale features and temporal-frequency relationships. In this study, we propose a method for infant cry emotion recognition using an improved Emphasized Channel Attention, Propagation and Aggregation in Time Delay Neural Network (ECAPA-TDNN) with both multi-scale feature fusion and attention enhancement. Experiments on a public dataset show that the proposed method achieves accuracy of 82.20%, number of parameters of 1.43 MB and FLOPs of 0.32 Giga. Moreover, our method has advantage over the baseline methods in terms of accuracy. The code is at https://github.com/kkpretend/IETMA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。