用对比损失提升脑电语音解码准确率
Decoding Speech Envelopes from Electroencephalogram with a Contrastive Pearson Correlation Coefficient Loss
- 设计对比皮尔逊相关损失,强化关注与未关注语音包络的差异
- 在三个公开数据集上提升语音分离度和听觉注意力解码精度
- 揭示了不同数据集与模型架构下的失败模式,指导后续优化
近年来,从脑电图(EEG)信号中重构语音包络的技术推动了多说话人环境中连续听觉注意力解码(AAD)的发展。大多数基于深度神经网络(DNN)的包络重构模型通过最大化关注语音包络与重建包络之间的皮尔逊相关系数(PCC)来训练(即关注PCC)。然而,关注PCC与未关注PCC之间的差异在听觉注意力解码中起关键作用,现有方法往往仅聚焦于最大化关注PCC。为此,本文提出一种对比性PCC损失,用于表征关注与未关注PCC之间的差异。该方法在三个公开的EEG-AAD数据集上,使用四种DNN架构进行评估。实验表明,所提目标函数在多种设置下均提升了包络可分性和AAD准确率,同时揭示了数据集与模型架构依赖的失败案例。
原文摘要 · Abstract (English)
Recent advances in reconstructing speech envelopes from Electroencephalogram (EEG) signals have enabled continuous auditory attention decoding (AAD) in multi-speaker environments. Most Deep Neural Network (DNN)-based envelope reconstruction models are trained to maximize the Pearson correlation coefficients (PCC) between the attended envelope and the reconstructed envelope (attended PCC). While the difference between the attended PCC and the unattended PCC plays an essential role in auditory attention decoding, existing methods often focus on maximizing the attended PCC. We therefore propose a contrastive PCC loss which represents the difference between the attended PCC and the unattended PCC. The proposed approach is evaluated on three public EEG AAD datasets using four DNN architectures. Across many settings, the proposed objective improves envelope separability and AAD accuracy, while also revealing dataset- and architecture-dependent failure cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。