用自监督学习对齐脑电与语音特征,提升听觉注意力检测精度
A contrastive-learning approach for auditory attention detection
- 通过对比学习对齐脑电信号与目标语音的潜在表示
- 在有限数据下实现验证集上最先进的分类性能
- 适合脑机接口与听力障碍辅助技术研究者
在多声源环境中进行对话是极具挑战性的任务,因为声音在时间和频率上重叠,难以分离出单一声源。一种可能的解决方案是通过解码脑电图(EEG)并利用统计或机器学习方法识别出被关注的语音源。然而,与其他机器学习问题相比,该任务的数据量有限,且不同EEG记录间存在分布差异,因此亟需一种适用于小样本的自监督学习方法以获得更鲁棒的解法。本文提出一种基于自监督学习的方法,旨在最小化被关注语音信号与其对应EEG信号之间的潜在表示差异。该网络随后在听觉注意力分类任务上进行微调。实验结果表明,本方法在验证集上优于已有方法,达到当前最优性能。
原文摘要 · Abstract (English)
Carrying conversations in multi-sound environments is one of the more challenging tasks, since the sounds overlap across time and frequency making it difficult to understand a single sound source. One proposed approach to help isolate an attended speech source is through decoding the electroencephalogram (EEG) and identifying the attended audio source using statistical or machine learning techniques. However, the limited amount of data in comparison to other machine learning problems and the distributional shift between different EEG recordings emphasizes the need for a self supervised approach that works with limited data to achieve a more robust solution. In this paper, we propose a method based on self supervised learning to minimize the difference between the latent representations of an attended speech signal and the corresponding EEG signal. This network is further finetuned for the auditory attention classification task. We compare our results with previously published methods and achieve state-of-the-art performance on the validation set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。