用廉价脑电数据解码语音,提升脑机接口说话能力
A Penny for Your Thoughts: Decoding Speech from Inexpensive Brain Signals
- 用对比学习对齐脑电与语音嵌入,构建跨模态映射
- 引入个性化结构后字错误率降低1.87%,最高提升0.45%
- 适合脑机接口、神经工程领域研究者参考
我们探索神经网络能否通过将脑电图(EEG)记录映射到音频表示来解码语音。基于受试者聆听自然语言时采集的EEG数据,训练模型采用对比性CLIP损失,使脑电衍生嵌入与预训练的基于Transformer的语音模型嵌入对齐。在Meta现有最先进EEG解码器基础上,提出三项架构改进:(i) 个体化注意力层(字错误率降低0.15%),(ii) 个性化空间注意力(降低0.45%),(iii) 带注意力的双路径RNN(降低1.87%)。其中两项改进提升性能,表明个性化架构在脑-语音解码中的潜力及其在脑机接口中的应用前景。
原文摘要 · Abstract (English)
We explore whether neural networks can decode brain activity into speech by mapping EEG recordings to audio representations. Using EEG data recorded as subjects listened to natural speech, we train a model with a contrastive CLIP loss to align EEG-derived embeddings with embeddings from a pre-trained transformer-based speech model. Building on the state-of-the-art EEG decoder from Meta, we introduce three architectural modifications: (i) subject-specific attention layers (+0.15% WER improvement), (ii) personalized spatial attention (+0.45%), and (iii) a dual-path RNN with attention (-1.87%). Two of the three modifications improved performance, highlighting the promise of personalized architectures for brain-to-speech decoding and applications in brain-computer interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。