用因果卷积网络直接解码听觉注意力,提升多说话人场景下的识别准确率。
CA-TCN: A Causal-Anticausal Temporal Convolutional Network for Direct Auditory Attention Decoding
- 采用因果与反因果卷积并行处理,对齐声音刺激与神经响应的时间方向。
- 在多个数据集上比现有模型准确率提升0.5%~3.2%,显著优于AADNet。
- 结构稳定、适合在线处理,适用于真实场景的听觉注意力系统。
在复杂听觉环境中,通过神经记录直接解码听觉注意力(AAD)是一种有前景的方法,旨在从多重说话人场景中识别出关注的语音流。传统的基于同步性的AAD方法通常依赖干净语音源和脑电图(EEG)信号,利用神经反应与目标刺激间的低频相关性。本文提出CA-TCN,一种因果-反因果时间卷积网络,可直接分类关注的说话人。该架构融合序列建模中卷积网络的最佳实践,通过分别使用因果与反因果卷积,以不同感受野在相反时间方向上对齐听觉刺激与神经反应。实验表明,相较于三种基线模型,CA-TCN在多个数据集与决策窗口中持续提升解码准确率:在跨被试模型中提升0.5%~3.2%,在个体特定模型中提升0.8%~2.9%,均优于表现最佳的AADNet。在六个评估设置中,四组比较显示最小期望切换持续时间分布具有统计显著性。此外,模型在空间上表现出稳健性,不同数据集间脑电空间滤波模式保持稳定。总体而言,该工作提出了一种高效统一的AAD模型,在准确率与实用性上均超越现有方法,推动了其在真实系统中的应用。
原文摘要 · Abstract (English)
A promising approach for steering auditory attention in complex listening environments relies on Auditory Attention Decoding (AAD), which aim to identify the attended speech stream in a multiple speaker scenario from neural recordings. Entrainment-based AAD approaches, typically assume access to clean speech sources and electroencephalography (EEG) signals to exploit low-frequency correlations between the neural response and the attended stimulus. In this study, we propose CA-TCN, a Causal-Anticausal Temporal Convolutional Network that directly classifies the attended speaker. The proposed architecture integrates several best practices from convolutional neural networks in sequence processing tasks. Importantly, it explicitly aligns auditory stimuli and neural responses by employing separate causal and anticausal convolutions respectively, with distinct receptive fields operating in opposite temporal directions. Experimental results, obtained through comparisons with three baseline AAD models, demonstrated that CA-TCN consistently improved decoding accuracy across datasets and decision windows, with gains ranging from 0.5% to 3.2% for subject-independent models and from 0.8% to 2.9% for subject-specific models compared with the next best-performing model, AADNet. Moreover, these improvements were statistically significant in four of the six evaluated settings when comparing Minimum Expected Switch Duration distributions. Beyond accuracy, the model demonstrated spatial robustness across different conditions, as the EEG spatial filters exhibited stable patterns across datasets. Overall, this work introduces an accurate and unified AAD model that outperforms existing methods while considering practical benefits for online processing scenarios. These findings contribute to advancing the state of AAD and its applicability in real-world systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。