用实时核磁记录嘴部动作,实现高精度语音合成。
MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI
- 用自监督模型从核磁数据预测文本,结合时长预测对齐。
- 在未见过的说话人上达到15.18%词错误率,显著优于现有方法。
- 适用于语言康复、语音控制等需非声学输入的场景。
以往基于实时核磁(rtMRI)的语音合成模型严重依赖含噪声的真实语音作为监督信号。直接在真实语音的梅尔频谱上施加损失,会将语音内容与MRI噪声混淆,导致语音可懂度差。本文提出一种新方法:利用多模态自监督的AV-HuBERT模型从rtMRI中预测文本,并引入基于流的时长预测器实现说话人特异性对齐。预测出的文本和时长由语音解码器生成与输入对齐的语音,可在任意新声音中合成。我们在两个数据集上进行了充分实验,验证了方法对未见说话人的泛化能力。通过遮蔽rtMRI视频不同部位,评估各发音器官对文本预测的影响。该方法在USC-TIMIT MRI语料库上达到15.18%的词错误率,远超当前最先进水平。语音样例见https://mri2speech.github.io/MRI2Speech/
原文摘要 · Abstract (English)
Previous real-time MRI (rtMRI)-based speech synthesis models depend heavily on noisy ground-truth speech. Applying loss directly over ground truth mel-spectrograms entangles speech content with MRI noise, resulting in poor intelligibility. We introduce a novel approach that adapts the multi-modal self-supervised AV-HuBERT model for text prediction from rtMRI and incorporates a new flow-based duration predictor for speaker-specific alignment. The predicted text and durations are then used by a speech decoder to synthesize aligned speech in any novel voice. We conduct thorough experiments on two datasets and demonstrate our method's generalization ability to unseen speakers. We assess our framework's performance by masking parts of the rtMRI video to evaluate the impact of different articulators on text prediction. Our method achieves a $15.18\%$ Word Error Rate (WER) on the USC-TIMIT MRI corpus, marking a huge improvement over the current state-of-the-art. Speech samples are available at https://mri2speech.github.io/MRI2Speech/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。