用预训练模型从脑电数据还原清晰语音,提升可懂度。
MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors

- 分两路重建:语义内容与语音细节,再融合生成
- 在EEG和MEG上显著优于现有方法,可懂度大幅提升
- 适合脑机接口与听觉神经科学研究者
从非侵入式神经信号重构连续语音是探究人类听觉感知与构建安全、可扩展的语音脑机接口的核心挑战。尽管近期取得进展,但因非侵入式记录固有的噪声大、空间模糊及对感知语音信息保留不全,可懂语音重构仍难以实现。现有方法直接将神经活动映射为纠缠的语音表示,再通过神经声码器合成波形,结果仅在频谱上相似却难以理解。为此,我们提出MindVoice,一种利用预训练模型弥补神经记录中语义与声学信息不完整性的神经到语音重构框架。MindVoice将重建分解为两条互补路径:一路恢复高层语义内容,另一路估计细粒度声学特征。这些推断出的表征随后与强大的语音生成模型及上下文语音克隆结合,合成自然且可懂的语音。在EEG和MEG上的大量实验表明,MindVoice在多种指标上显著优于现有方法。结果表明,预训练先验为弥合噪声神经记录与自然语音之间的差距提供了原则性路径,为听觉神经科学研究与非侵入式语音脑机接口展示出光明前景。
原文摘要 · Abstract (English)
Reconstructing continuous speech from non-invasive neural recordings is a fundamental problem for probing human auditory perception and building safe, scalable speech brain-computer interfaces. Despite recent progress, intelligible reconstruction remains elusive, as non-invasive recordings are inherently noisy, spatially blurred, and only partially preserve information about perceived speech. Existing methods directly map neural activity to entangled speech representations before synthesizing waveforms with neural vocoders, resulting in spectral-similar but unintelligible results. To overcome these limitations, we introduce MindVoice, a neuro-to-speech reconstruction framework that uses pretrained models to compensate for the incomplete semantic and acoustic information in neural recordings. MindVoice disentangles reconstruction into two complementary pathways: one recovers high-level semantic content, while the other estimates fine-grained acoustic attributes. These inferred representations are then fused with powerful speech generation models and in-context voice cloning to synthesize natural and intelligible utterances. Extensive experiments on EEG and MEG demonstrate that MindVoice substantially outperforms existing methods on various metrics. These results show that pretrained priors provide a principled way to bridge the gap between noisy neural recordings and natural speech, highlighting a promising attempt for auditory neuroscience research and non-invasive speech brain-computer interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。