用动态谱特征与卡尔曼平滑提升语音情绪识别准确率
Enhancing Speech Emotion Recognition using Dynamic Spectral Features and Kalman Smoothing
- 引入动态谱特征和卡尔曼平滑降噪
- 在RAVDESS上达87%准确率,优于现有方法
- 适合语音情绪识别中抗噪声与稳定性需求
语音情绪识别系统通常使用静态特征如梅尔频率倒谱系数(MFCCs)、零交叉率(ZCR)和均方根能量(RMSE)。由于这些特征对声学噪声敏感,容易导致情绪误判。为此,本文引入动态谱特征(一阶差分与二阶差分)并结合卡尔曼平滑算法,有效降低噪声影响,提升分类稳定性。由于情绪随时间变化,卡尔曼平滑使输出更连续。在RAVDESS数据集上的实验表明,该方法达到87%的准确率,显著减少具有相似声学特征的情绪之间的误判。
原文摘要 · Abstract (English)
Speech Emotion Recognition systems often use static features like Mel-Frequency Cepstral Coefficients (MFCCs), Zero Crossing Rate (ZCR), and Root Mean Square Energy (RMSE). Because of this, they can misclassify emotions when there is acoustic noise in vocal signals. To address this, we added dynamic features using Dynamic Spectral features (Deltas and Delta-Deltas) along with the Kalman Smoothing algorithm. This approach reduces noise and improves emotion classification. Since emotion changes over time, the Kalman Smoothing filter also helped make the classifier outputs more stable. Tests on the RAVDESS dataset showed that this method achieved a state-of-the-art accuracy of 87\% and reduced misclassification between emotions with similar acoustic features
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。