arXiv:2601.16442eess.SPcs.HC2026-01

无需空间线索,用脑电解码听觉注意力,准确率提升超22%。

Auditory Attention Decoding without Spatial Information: A Diotic EEG Study

  • 通过共享潜空间映射脑电信号与语音特征,消除左右耳差异影响
  • 在双耳相同语音混合场景下达到72.70%准确率,比现有方法高22.58%
  • 适用于真实混杂环境,适合智能助听器与客观听力检测研究

听觉注意力解码(AAD)通过解码脑电信号(如EEG)识别多说话人环境中关注的语音流,对解决鸡尾酒会问题的智能助听器和客观听力评估系统至关重要。现有研究主要依赖双耳分听环境,即不同语音分别输入左右耳,使模型基于方向判断注意力,但此方式受限于空间线索,难以应用于真实场景中说话人重叠或动态移动的情况。为此,本文提出一种适用于双耳同源(diotic)环境的AAD框架,即在双耳同时呈现相同的语音混合信号,消除空间提示。方法采用独立编码器将EEG与语音映射至共享潜空间:语音特征使用wav2vec 2.0提取,并通过两层一维卷积神经网络(CNN)编码;脑电信号则采用BrainNetwork架构进行编码。模型通过计算脑电与语音表示间的余弦相似度来判定关注的语音。在公开的双耳同源EEG数据集上验证,取得72.70%的准确率,较当前最优的方向型方法高出22.58%。

原文摘要 · Abstract (English)

Auditory attention decoding (AAD) identifies the attended speech stream in multi-speaker environments by decoding brain signals such as electroencephalography (EEG). This technology is essential for realizing smart hearing aids that address the cocktail party problem and for facilitating objective audiometry systems. Existing AAD research mainly utilizes dichotic environments where different speech signals are presented to the left and right ears, enabling models to classify directional attention rather than speech content. However, this spatial reliance limits applicability to real-world scenarios, such as the "cocktail party" situation, where speakers overlap or move dynamically. To address this challenge, we propose an AAD framework for diotic environments where identical speech mixtures are presented to both ears, eliminating spatial cues. Our approach maps EEG and speech signals into a shared latent space using independent encoders. We extract speech features using wav2vec 2.0 and encode them with a 2-layer 1D convolutional neural network (CNN), while employing the BrainNetwork architecture for EEG encoding. The model identifies the attended speech by calculating the cosine similarity between EEG and speech representations. We evaluate our method on a diotic EEG dataset and achieve 72.70% accuracy, which is 22.58% higher than the state-of-the-art direction-based AAD method.

听觉注意脑电解码双耳同源智能助听

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。