arXiv:2606.24164eess.AScs.AI2026-06中稿 · Interspeech 2026

解决脑电引导语音提取中的试次捷径问题,提升跨试次泛化能力。

Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training

论文配图:Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training
图 1 · 摘自论文原文
  • 分两阶段训练,用对比学习抑制试次身份线索
  • 在KUL和DTU数据集上跨试次性能显著优于基线
  • 适合开发可靠神经接口的语音重建系统

近期端到端脑电引导目标语音提取模型取得了显著成果,展现出神经调控听觉技术的巨大潜力。然而,我们的分析发现,高试内性能可能源于试次特异的脑电信号结构,这些结构充当了目标选择的捷径,导致在未见试次上泛化能力差。为克服这一缺陷,我们提出TRUST-TSE,一种两阶段框架以缓解捷径学习。通过引入带注意力说话人负采样的对比预训练,促使脑电编码器捕捉精细的脑电-语音对齐关系,同时抑制试次身份线索。此外,还采用基于脑电-源信号相似性的置信度加权提取目标,指导利用学习到的表示进行语音重建。在KUL和DTU数据集上的实验表明,TRUST-TSE在严格的跨试次协议下优于端到端基线模型,解决了现有方法的关键可靠性瓶颈。

原文摘要 · Abstract (English)

Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis reveals that high within-trial performance can be driven by trial-specific EEG structure that acts as shortcuts for target selection, leading to poor generalization on unseen trials. To overcome this gap, we propose TRUST-TSE, a two-stage framework to mitigate shortcut learning. By introducing contrastive pretraining with attended-speaker negative sampling, we encourage the EEG encoder to capture fine-grained EEG--speech alignment while suppressing trial-identity cues. We also employ a confidence-weighted extraction objective based on EEG--source similarity to guide extraction using the learned representations. Experiments on KUL and DTU datasets show that TRUST-TSE outperforms end-to-end baselines under strict cross-trial protocols, addressing a key reliability bottleneck of existing approaches.

脑电解码语音提取捷径学习跨试次泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。