arXiv:2503.06588cs.SDcs.CV2025-03

用动态MRI生成语音,解决数据丢失问题。

Speech Audio Generation from dynamic MRI via a Knowledge Enhanced Conditional Variational Autoencoder

  • 用未标注MRI数据增强知识,再通过变分推断生成语音
  • 在真实语音波形生成上优于传统深度学习方法
  • 适合语音康复、脑机接口等临床与科研场景

动态磁共振成像(MRI)在语音运动研究中日益普及。然而,由于成像环境不可预测,常出现数据丢失、噪声污染和音频文件损坏。在此背景下,从图像中重建语音对临床与研究应用至关重要。现有方法如去噪与多阶段合成在语音保真度和泛化能力上存在局限。为此,本文提出知识增强的条件变分自编码器(KE-CVAE),采用两步框架:先利用未标注的MRI数据进行知识增强,再通过变分推断提升生成建模能力。该方法是首个直接从动态MRI视频序列合成语音的尝试。模型在公开的语音动态声门区MRI数据集上训练与评估,实验结果表明其能有效生成自然语音波形,克服特定于MRI的声学挑战,显著优于传统深度学习合成方法。

原文摘要 · Abstract (English)

Dynamic Magnetic Resonance Imaging (MRI) of the vocal tract has become an increasingly adopted imaging modality for speech motor studies. Beyond image signals, systematic data loss, noise pollution, and audio file corruption can occur due to the unpredictability of the MRI acquisition environment. In such cases, generating audio from images is critical for data recovery in both clinical and research applications. However, this remains challenging due to hardware constraints, acoustic interference, and data corruption. Existing solutions, such as denoising and multi-stage synthesis methods, face limitations in audio fidelity and generalizability. To address these challenges, we propose a Knowledge Enhanced Conditional Variational Autoencoder (KE-CVAE), a novel two-step "knowledge enhancement + variational inference" framework for generating speech audio signals from cine dynamic MRI sequences. This approach introduces two key innovations: (1) integration of unlabeled MRI data for knowledge enhancement, and (2) a variational inference architecture to improve generative modeling capacity. To the best of our knowledge, this is one of the first attempts at synthesizing speech audio directly from dynamic MRI video sequences. The proposed method was trained and evaluated on an open-source dynamic vocal tract MRI dataset recorded during speech. Experimental results demonstrate its effectiveness in generating natural speech waveforms while addressing MRI-specific acoustic challenges, outperforming conventional deep learning-based synthesis approaches.

语音生成MRI变分自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。