用视觉变压器和对比学习直接从脑电图解码语音,支持无线植入系统。
Neural Decoding of Overt Speech from ECoG Using Vision Transformers and Contrastive Representation Learning
- 结合视觉变换器与对比学习,端到端回归语音波形。
- 在临床电极与可植入无线系统上均实现有效语音重建。
- 首次在可植入无线系统上完成语音解码,适合长期脑机接口应用。
言语脑机接口(BCI)为严重瘫痪患者提供了沟通希望。近期研究已能通过表面皮层脑电图(ECoG)或皮层内记录,预测音素或词序列,并借助下游语言模型生成有意义句子。当前挑战在于实现流式语音重建,即直接将皮层信号回归为声学语音。尽管已有研究使用皮层内数据实现了该目标,但针对表面ECoG记录仍需进一步工作,尤其需要优化神经解码器。本文提出一种基于编码器-解码器架构的离线语音解码流程,融合视觉变换器与对比学习,以增强从ECoG信号直接回归语音的能力。该方法在两个数据集上评估:一个来自癫痫患者使用临床硬膜下电极采集的数据,另一个来自运动BCI试验中使用完全植入式WIMAGINE硬膜外系统的受试者数据。据我们所知,这是首次在完全植入且无线的硬膜外记录系统上实现语音解码,为长期应用提供了前景。
原文摘要 · Abstract (English)
Speech Brain Computer Interfaces (BCIs) offer promising solutions to people with severe paralysis unable to communicate. A number of recent studies have demonstrated convincing reconstruction of intelligible speech from surface electrocorticographic (ECoG) or intracortical recordings by predicting a series of phonemes or words and using downstream language models to obtain meaningful sentences. A current challenge is to reconstruct speech in a streaming mode by directly regressing cortical signals into acoustic speech. While this has been achieved recently using intracortical data, further work is needed to obtain comparable results with surface ECoG recordings. In particular, optimizing neural decoders becomes critical in this case. Here we present an offline speech decoding pipeline based on an encoder-decoder deep neural architecture, integrating Vision Transformers and contrastive learning to enhance the direct regression of speech from ECoG signals. The approach is evaluated on two datasets, one obtained with clinical subdural electrodes in an epileptic patient, and another obtained with the fully implantable WIMAGINE epidural system in a participant of a motor BCI trial. To our knowledge this presents a first attempt to decode speech from a fully implantable and wireless epidural recording system offering perspectives for long-term use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。