直接从脑电数据生成可懂语音,无中间文本步骤,实时且准确。
Brain2Speech-Net: Intelligible, Real-Time Brain-to-Speech Synthesis Without Text Decoding

- 单阶段框架,用可微分音素瓶颈保留语言结构。
- 在有限数据下实现高可懂度,语音生成速度超实时。
- 适合神经康复与脑机接口研究者,无需文本中间环节。
失语症患者因瘫痪丧失语言能力。通过直接从神经活动合成语音极具挑战:皮层数据稀缺且缺乏对齐标签,多数系统依赖神经信号→文本→语音的级联流程,导致延迟并传播错误。我们提出 Brain2Speech-Net,是首个在数据有限情况下仍保持可懂度的单阶段脑到语音合成框架,无需中间文本解码。其采用可微分音素瓶颈保留语言结构;轻量级深度隐马尔可夫模型对齐器将该瓶颈映射至语音合成潜空间中的上下文音素表示。该对齐器在无帧级监督下学习神经记录与音素段之间的单调对齐,继承强声学先验,实现数据高效训练。在皮层数据集上,Brain2Speech-Net 在客观测试和听觉测试中均表现出高可懂度,且运行速度超过实时。相比级联系统延迟高、直接语音单元模型可懂度差,本方法同时实现可懂与实时语音合成。
原文摘要 · Abstract (English)
The loss of speech limits communication for individuals with paralysis. Restoring speech by synthesizing it directly from neural activity is challenging: intracortical data are scarce and lack aligned targets, so most systems rely on cascaded neural-to-text-to-speech pipelines that add latency and propagate errors. We present Brain2Speech-Net, among the first single-stage frameworks to remain intelligible under limited data while removing intermediate text decoding. A differentiable phoneme bottleneck preserves linguistic structure without explicit text decoding. A lightweight deep-HMM aligner then maps this bottleneck to contextual phoneme representations in a TTS latent space. It learns monotonic alignment between neural recordings and phoneme segments without frame-level supervision, inheriting strong acoustic priors for data-efficient training. On an intracortical dataset, Brain2Speech-Net achieves strong intelligibility in objective and listening tests while running faster than real time. Unlike cascaded systems that incur high latency and direct speech-unit models that lack intelligibility, it delivers both intelligible and real-time speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。