arXiv:2411.13217eess.SPcs.SD2024-11被引 10

用脑电波区分音乐、语音和听感偏好,准确率最高达98.66%。

Energy-based features and bi-LSTM neural network for EEG-based music and voice classification

  • 基于脑电能量关系构建特征矩阵,结合双向LSTM分类
  • 语音与音乐二分类准确率达98.66%,四种音乐流派多分类达61.59%
  • 适用于脑机接口中的听觉刺激分析,适合神经科学与人机交互研究者

人类大脑通过多种方式接收外界刺激,其中音频是沟通、娱乐、警示等的重要信息来源。本文旨在提升对不同音乐类型及语音类声音的脑电响应分类能力。设计两项实验,采集受试者聆听不同音乐流派歌曲及多语言句子时的脑电图(EEG)信号。提出一种新方法:构建基于各脑电通道能量关系的特征矩阵,并使用双向长短期记忆网络(bi-LSTM)进行分类。评估了语音与音乐的二分类、四种音乐流派的多分类,以及听众是否喜欢所听歌曲的二分类任务。结果表明,该方案表现良好:语音与音乐二分类准确率达98.66%;四种音乐流派多分类准确率为61.59%;音乐偏好二分类准确率达96.96%。

原文摘要 · Abstract (English)

The human brain receives stimuli in multiple ways; among them, audio constitutes an important source of relevant stimuli for the brain regarding communication, amusement, warning, etc. In this context, the aim of this manuscript is to advance in the classification of brain responses to music of diverse genres and to sounds of different nature: speech and music. For this purpose, two different experiments have been designed to acquiere EEG signals from subjects listening to songs of different musical genres and sentences in various languages. With this, a novel scheme is proposed to characterize brain signals for their classification; this scheme is based on the construction of a feature matrix built on relations between energy measured at the different EEG channels and the usage of a bi-LSTM neural network. With the data obtained, evaluations regarding EEG-based classification between speech and music, different musical genres, and whether the subject likes the song listened to or not are carried out. The experiments unveil satisfactory performance to the proposed scheme. The results obtained for binary audio type classification attain 98.66% of success. In multi-class classification between 4 musical genres, the accuracy attained is 61.59%, and results for binary classification of musical taste rise to 96.96%.

脑电分析音乐识别语音识别深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。