用脑电引导的动态门控,实现听觉注意力切换时的稳定语音分离。
SAGE: Switch-Aware EEG-Guided Soft Gating for Target Speaker Extraction with In-Trial Switching

- 将注意力切换视为动态选择,通过脑电信号生成平滑融合权重。
- 在真实切换场景中实现8.67 dB SI-SDR与2.04秒平均延迟。
- 适合需要实时脑机协同语音增强的智能助听或人机交互系统。
在听觉注意力发生试验内切换时,脑电引导的目标语音提取面临神经噪声和固有延迟问题,导致注意力追踪延迟或不稳定。传统方法难以应对动态切换,常在切换点产生不连续。为此,我们提出SAGE——一种面向切换的脑电引导软门控框架,将试验内切换视为动态选择。SAGE通过稳健分离器生成两条候选语音流,并利用脑电引导的切换感知门控模块生成平滑融合权重,抑制过渡伪影。进一步引入延迟补偿对齐与不确定性驱动的保守策略,以应对延迟差异和脑电信号可靠性波动。SAGE超越基线方法,在目标语音分离任务中达到8.67 dB SI-SDR与88.24% STOI,同时将平均切换延迟降至2.04秒。该方法通过耦合神经解码与语音分离,实现了动态场景下的鲁棒目标提取。
原文摘要 · Abstract (English)
EEG-guided target speaker extraction is challenging under in-trial auditory attention switching, where neural noise and intrinsic latency can delay or destabilize attention tracking. Conventional methods struggle with dynamic switches and often cause discontinuities at switching points. Therefore, we propose SAGE, a switch-aware EEG-guided soft gating framework that treats in-trial switching as dynamic selection. SAGE generates two candidate speech streams with a robust separator and uses an EEG-guided switch-aware gating module to produce smooth fusion weights and suppress transition artifacts. We further integrate latency-compensated alignment and an uncertainty-driven conservative strategy to handle latency discrepancies and fluctuating EEG reliability. SAGE outperforms baselines, achieving 8.67 dB SI-SDR and 88.24% STOI while reducing average switching latency to 2.04 s. By coupling neural decoding with speech separation, it enables robust target extraction in dynamic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。