融合脑电与音频空间谱,实现14方向听觉注意力精准解码。
Multi-class Decoding of Attended Speaker Direction Using Electroencephalogram and Audio Spatial Spectrum
- 用脑电+音频空间谱双模态数据解码听觉注意力方向。
- 在1秒决策窗口下,14类方向解码准确率达57.19%。
- 适用于听力障碍者脑机接口,尤其适合多方向场景。
从听者脑电图(EEG)信号中解码关注说话人的方向,对提升听力障碍人士生活质量的脑机接口至关重要。以往研究主要集中在左右二分类解码,但精确识别具体方向对有效语音处理更为关键。此外,音频空间信息未被充分利用,导致解码效果不佳。本文在新提出的14类方向关注数据集上发现,仅依赖EEG输入的模型在留一被试和留一试验场景下准确率显著偏低。通过将音频空间谱与EEG特征融合,可有效提升解码性能。采用CNN、LSM-CNN及Deformer模型从EEG信号与音频空间谱中解码方向注意力。所提出的Sp-EEG-Deformer模型在留一被试与留一试验场景下分别达到55.35%和57.19%的14类解码准确率,决策窗口为1秒。实验表明,当候选方向数减少时,解码准确率提升。结果验证了双模态方向注意力解码策略的有效性。
原文摘要 · Abstract (English)
Decoding the directional focus of an attended speaker from listeners' electroencephalogram (EEG) signals is essential for developing brain-computer interfaces to improve the quality of life for individuals with hearing impairment. Previous works have concentrated on binary directional focus decoding, i.e., determining whether the attended speaker is on the left or right side of the listener. However, a more precise decoding of the exact direction of the attended speaker is necessary for effective speech processing. Additionally, audio spatial information has not been effectively leveraged, resulting in suboptimal decoding results. In this paper, it is found that on the recently presented dataset with 14-class directional focus, models relying exclusively on EEG inputs exhibit significantly lower accuracy when decoding the directional focus in both leave-one-subject-out and leave-one-trial-out scenarios. By integrating audio spatial spectra with EEG features, the decoding accuracy can be effectively improved. The CNN, LSM-CNN, and Deformer models are employed to decode the directional focus from listeners' EEG signals and audio spatial spectra. The proposed Sp-EEG-Deformer model achieves notable 14-class decoding accuracies of 55.35% and 57.19% in leave-one-subject-out and leave-one-trial-out scenarios with a decision window of 1 second, respectively. Experiment results indicate increased decoding accuracy as the number of alternative directions reduces. These findings suggest the efficacy of our proposed dual modal directional focus decoding strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。