arXiv:2512.22146eess.SPcs.LG2025-12被引 1

无需对齐时间,直接从脑电波还原说话语音和想象中的语音。

EEG-to-Voice Decoding of Spoken and Imagined speech Using Non-Invasive EEG

  • 用个性化生成器直接从脑电数据生成语音频谱图。
  • 4秒内语音还原准确率高,词错误率降低30%以上。
  • 适合脑机接口、语言神经机制研究者阅读。

恢复由神经信号产生的言语通信是脑机接口研究的核心目标,但受空间分辨率有限、噪声敏感及想象言语缺乏时间对齐声学目标等因素影响,基于脑电图(EEG)的言语重建仍具挑战性。本研究提出一种非侵入式EEG-to-Voice范式,无需动态时间规整(DTW)或显式时间对齐,直接从EEG信号生成语音。该流程采用个体专属生成器在开环模式下生成梅尔频谱图,再通过预训练声码器与自动语音识别(ASR)模块合成语音波形并解码文本。分别针对真实言语与想象言语训练独立生成器,并通过迁移学习实现从真实言语到想象言语的领域自适应。引入基于最小语言模型的纠错模块,可有效修正部分ASR错误,同时保持语义连贯性。在2秒和4秒言语条件下,采用声学级指标(PCC、RMSE、MCD)与语言级指标(CER、WER)进行评估。结果显示,真实与想象言语均实现稳定声学重构,语言准确性相当;尽管长句声学相似度下降,但文本解码性能基本保持,词位置分析显示句子后半部分错误略有增加。语言模型纠错模块持续降低CER与WER,未引入语义失真。结果表明,无需显式时间对齐,即可实现真实与想象言语的直接、开环式EEG-to-Voice重建。

原文摘要 · Abstract (English)

Restoring speech communication from neural signals is a central goal of brain-computer interface research, yet EEG-based speech reconstruction remains challenging due to limited spatial resolution, susceptibility to noise, and the absence of temporally aligned acoustic targets in imagined speech. In this study, we propose an EEG-to-Voice paradigm that directly reconstructs speech from non-invasive EEG signals without dynamic time warping (DTW) or explicit temporal alignment. The proposed pipeline generates mel-spectrograms from EEG in an open-loop manner using a subject-specific generator, followed by pretrained vocoder and automatic speech recognition (ASR) modules to synthesize speech waveforms and decode text. Separate generators were trained for spoken speech and imagined speech, and transfer learning-based domain adaptation was applied by pretraining on spoken speech and adapting to imagined speech. A minimal language model-based correction module was optionally applied to correct limited ASR errors while preserving semantic structure. The framework was evaluated under 2 s and 4 s speech conditions using acoustic-level metrics (PCC, RMSE, MCD) and linguistic-level metrics (CER, WER). Stable acoustic reconstruction and comparable linguistic accuracy were observed for both spoken speech and imagined speech. While acoustic similarity decreased for longer utterances, text-level decoding performance was largely preserved, and word-position analysis revealed a mild increase in decoding errors toward later parts of sentences. The language model-based correction consistently reduced CER and WER without introducing semantic distortion. These results demonstrate the feasibility of direct, open-loop EEG-to-Voice reconstruction for spoken speech and imagined speech without explicit temporal alignment.

脑机接口语音生成脑电信号想象言语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。