arXiv:2412.10469cs.SDcs.LG2024-12被引 3

用深度学习比较两种语音特征在情绪识别中的表现

Comparative Analysis of Mel-Frequency Cepstral Coefficients and Wavelet Based Audio Signal Processing for Emotion Detection and Mental Health Assessment in Spoken Speech

  • 对比小波特征与梅尔倒谱系数在语音情绪识别中的效果
  • CNN模型准确率达61%,优于LSTM的56%
  • 适合关注语音情感分析与心理健康评估的研究者

技术与心理健康交叉催生了基于音频数据计算分析的新方法。本研究利用卷积神经网络(CNN)和长短期记忆网络(LSTM)对小波提取特征和梅尔倒谱系数(MFCCs)进行情绪识别。通过数据增强、特征提取、归一化与模型训练,评估模型在分类情绪状态中的表现。结果表明,CNN模型准确率为61%,高于LSTM的56%。两者在识别惊讶与愤怒等特定情绪时表现更优,依赖于音高与语速变化等声学特征。建议进一步探索先进数据增强、融合特征提取及语言分析与语音特征结合的方法,以提升心理健康诊断精度。呼吁建立标准化数据集以推动情感计算与心理健康干预发展。

原文摘要 · Abstract (English)

The intersection of technology and mental health has spurred innovative approaches to assessing emotional well-being, particularly through computational techniques applied to audio data analysis. This study explores the application of Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) models on wavelet extracted features and Mel-frequency Cepstral Coefficients (MFCCs) for emotion detection from spoken speech. Data augmentation techniques, feature extraction, normalization, and model training were conducted to evaluate the models' performance in classifying emotional states. Results indicate that the CNN model achieved a higher accuracy of 61% compared to the LSTM model's accuracy of 56%. Both models demonstrated better performance in predicting specific emotions such as surprise and anger, leveraging distinct audio features like pitch and speed variations. Recommendations include further exploration of advanced data augmentation techniques, combined feature extraction methods, and the integration of linguistic analysis with speech characteristics for improved accuracy in mental health diagnostics. Collaboration for standardized dataset collection and sharing is recommended to foster advancements in affective computing and mental health care interventions.

情绪识别语音分析深度学习心理健康

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。