arXiv:2603.11845eess.AS2026-03被引 1

用普通录音替代带噪MRI录音,实现高精度发音重建。

Acoustic-to-Articulatory Inversion of Clean Speech Using an MRI-Trained Model

  • 用干净语音替代去噪MRI语音训练模型
  • 重建误差仅1.56毫米,接近MRI原生数据表现
  • 适合无MRI设备但需发音建模的研究者

发音声学反演旨在从语音中重建声道形状。实时磁共振成像(rt-MRI)可同步获取语音信号与发音信息,但其录制的音频受扫描仪噪声严重污染,需去噪才能使用。为实际应用,需在无MRI噪声环境下实现反演。本研究探讨以干净声学环境录制的语音作为去噪MRI语音的替代方案。对比同一说话人、相同语句的两组信号,通过音素分割对齐。评估在去噪MRI语音上训练的模型在去噪MRI语音与干净语音上的表现,并测试仅在干净语音上训练和测试的模型。结果表明,干净语音可有效支持发音反演,达到1.56毫米的均方根误差(RMSE),接近基于MRI数据的表现。

原文摘要 · Abstract (English)

Articulatory acoustic inversion reconstructs vocal tract shapes from speech. Real-time magnetic resonance imaging (rt-MRI) allows simultaneous acquisition of both the acoustic speech signal and articulatory information. Besides the complexity of rt-MRI acquisition, the recorded audio is heavily corrupted by scanner noise and requires denoising to be usable. For practical use, it must be possible to invert speech recorded without MRI noise. In this study, we investigate the use of speech recorded in a clean acoustic environment as an alternative to denoised MRI speech. To this end we compare two signals from the same speaker with identical sentences which are aligned using phonetic segmentation. A model trained on denoised MRI speech is evaluated on both denoised MRI and clean speech. We also assess a model trained and tested only on clean speech. Results show that clean speech supports articulatory inversion effectively, achieving an RMSE of 1.56 mm, close to MRI-based performance.

发音重建语音分析声学反演干净语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。