用动态融合模型提升脑电语音包络重建精度
DECAF: Dynamic Envelope Context-Aware Fusion for Speech-Envelope Reconstruction from EEG
- 引入状态空间模型,动态融合脑电信号与语音上下文
- 在ICASSP 2023数据集上显著优于传统静态方法
- 适合神经解码与听觉注意力研究者参考
从头皮脑电(EEG)重构语音音频包络是解码听者注意焦点的核心任务,适用于神经控制助听器等场景。现有方法多将此视为静态回归问题,孤立处理每个EEG窗口,忽视连续语音中的丰富时序结构。本文提出一种动态框架,利用语音上下文作为预测性时序先验。设计了一种状态空间融合模型,通过可学习门控机制自适应平衡来自EEG的直接神经估计与近期语音上下文的预测信号。在ICASSP 2023语音重构基准测试中验证了该方法,显著优于仅依赖EEG的静态基线。分析揭示神经信号与时序信息之间存在强大协同效应。本工作将包络重构重新定义为动态状态估计问题,为构建更精确、连贯的神经解码系统开辟新方向。
原文摘要 · Abstract (English)
Reconstructing the speech audio envelope from scalp neural recordings (EEG) is a central task for decoding a listener's attentional focus in applications like neuro-steered hearing aids. Current methods for this reconstruction, however, face challenges with fidelity and noise. Prevailing approaches treat it as a static regression problem, processing each EEG window in isolation and ignoring the rich temporal structure inherent in continuous speech. This study introduces a new, dynamic framework for envelope reconstruction that leverages this structure as a predictive temporal prior. We propose a state-space fusion model that combines direct neural estimates from EEG with predictions from recent speech context, using a learned gating mechanism to adaptively balance these cues. To validate this approach, we evaluate our model on the ICASSP 2023 Stimulus Reconstruction benchmark demonstrating significant improvements over static, EEG-only baselines. Our analyses reveal a powerful synergy between the neural and temporal information streams. Ultimately, this work reframes envelope reconstruction not as a simple mapping, but as a dynamic state-estimation problem, opening a new direction for developing more accurate and coherent neural decoding systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。