用脑电波生成自然语音,提升发音清晰度与语调准确率。
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
- 结合小波分析与Transformer,从脑电信号中提取多维度特征并预测语调。
- 重建语音与原始语音的梅尔谱相关性达0.91,优于现有方法。
- 适合神经假肢、脑机接口领域研究者,尤其关注语音恢复应用。
从颅内脑电(iEEG)信号中合成语音为严重言语障碍患者恢复交流能力提供了新途径。然而,由于特征表示不足、语调建模困难和相位重建不准确,实现清晰自然的语音仍具挑战。本文提出MiSTR,一种深度学习框架,包含:1)基于小波的特征提取,以捕捉iEEG信号的精细时间、频谱及神经生理特征;2)基于Transformer的解码器,实现考虑语调的声谱图预测;3)神经相位声码器,通过自适应频谱校正保证谐波一致性。在公开iEEG数据集上评估,MiSTR达到当前最优语音可懂度,重建声谱与原始声谱的平均皮尔逊相关系数为0.91,显著优于现有神经语音合成基线。
原文摘要 · Abstract (English)
Speech synthesis from intracranial EEG (iEEG) signals offers a promising avenue for restoring communication in individuals with severe speech impairments. However, achieving intelligible and natural speech remains challenging due to limitations in feature representation, prosody modeling, and phase reconstruction. We introduce MiSTR, a deep-learning framework that integrates: 1) Wavelet-based feature extraction to capture fine-grained temporal, spectral, and neurophysiological representations of iEEG signals, 2) A Transformer-based decoder for prosody-aware spectrogram prediction, and 3) A neural phase vocoder enforcing harmonic consistency via adaptive spectral correction. Evaluated on a public iEEG dataset, MiSTR achieves state-of-the-art speech intelligibility, with a mean Pearson correlation of 0.91 between reconstructed and original Mel spectrograms, improving over existing neural speech synthesis baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。