用跨会话重复刺激构建对比学习,提升脑电转语音的准确性。
Contrastive Learning with Variational Regularization for Multi-Session EEG-to-Speech Decoding
- 以跨会话相同刺激为正样本,设计对比学习框架。
- 在日语数据集上,字符错误率降低,语音重建质量保持稳定。
- 适合脑机接口、神经解码方向研究者参考。
从非侵入性脑电图(EEG)重构听到的语音具有挑战性,主要因为信噪比低和跨会话差异大。虽然试次平均可提高信噪比,但难以应用于连续语音。本文利用同一刺激在不同会话中的重复脑电响应作为正样本,构建对比学习机制,并引入变分正则化,与对比目标结合,使编码器表示空间保持广泛性。在日语EEG数据集上的实验表明,结合会话不变策略与变分正则化,能有效降低字符错误率(CER),同时保持梅尔频谱图重建质量。会话探查分析证实,编码器表示实现了会话不变性。
原文摘要 · Abstract (English)
Reconstructing heard speech from non-invasive electroencephalography (EEG) is challenging due to a low signal-to-noise ratio (SNR) and inter-session variability. While trial averaging improves the SNR, it is difficult to apply to continuous speech. We instead use repeated EEG responses to the same stimulus across different sessions as positive pairs for contrastive learning, and introduce variational regularization that, combined with this contrastive objective, keeps the encoder representation space broad. Experiments on a Japanese EEG dataset show that combining the session-invariant strategy with variational regularization improves the character error rate (CER) while maintaining mel-spectrogram reconstruction fidelity. Session probing confirms that the encoder representations achieve session-invariance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。