arXiv:2503.19677cs.SDcs.AI2025-03被引 7

用卷积神经网络分析语音梅尔频谱,提升情绪识别准确率。

Deep Learning for Speech Emotion Recognition: A CNN Approach Utilizing Mel Spectrograms

  • 将语音转为梅尔频谱图,用CNN自动提取情绪特征。
  • 模型在公开数据集上达到85%以上分类准确率。
  • 提供可视化界面,适合教育场景实时应用。

本文研究了卷积神经网络(CNN)在语音情绪识别中的应用,通过音频文件的梅尔频谱表示进行情绪分类。传统方法如高斯混合模型和隐马尔可夫模型在实际部署中表现不足,促使转向深度学习技术。将音频数据转换为视觉形式后,CNN模型能够自主学习复杂模式,提升分类准确性。所开发的模型集成于用户友好的图形界面中,支持实时预测,具备在教育环境中的潜在应用价值。本研究旨在推进深度学习在语音情绪识别中的理解,评估模型可行性,并促进技术在学习场景中的融合。

原文摘要 · Abstract (English)

This paper explores the application of Convolutional Neural Networks CNNs for classifying emotions in speech through Mel Spectrogram representations of audio files. Traditional methods such as Gaussian Mixture Models and Hidden Markov Models have proven insufficient for practical deployment, prompting a shift towards deep learning techniques. By transforming audio data into a visual format, the CNN model autonomously learns to identify intricate patterns, enhancing classification accuracy. The developed model is integrated into a user-friendly graphical interface, facilitating realtime predictions and potential applications in educational environments. The study aims to advance the understanding of deep learning in speech emotion recognition, assess the models feasibility, and contribute to the integration of technology in learning contexts

语音情绪识别卷积神经网络梅尔频谱实时应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。