arXiv:2603.28723eess.AS2026-03

用MRI轮廓图实现全声道发音逆向建模,精度超像素级。

Acoustic-to-articulatory Inversion of the Complete Vocal Tract from RT-MRI with Various Audio Embeddings and Dataset Sizes

  • 从MRI图像自动提取声道轮廓,避开冗余像素信息
  • 平均误差1.48毫米,优于1.62毫米像素精度
  • 适用于语音合成与发音研究,尤其适合高精度建模

发音到声学的逆向建模高度依赖数据类型。以往研究多基于EMA(电磁发音仪),受限于传感器数量且仅覆盖部分发音器官。本文提出一种完整声道逆向建模方法,从声门至双唇。利用单名说话人约3.5小时的实时磁共振(RT-MRI)数据,创新性地采用从MRI图像自动提取的发音器轮廓作为输入,而非原始图像。该方法聚焦声道几何动态,忽略冗余像素信息。结合降噪音频,使用双向长短期记忆网络(Bi-LSTM)处理。实验分为两部分:(1) 比较三种音频嵌入(MFCC、LCC、HuBERT)对模型性能的影响;(2) 分析数据量从10分钟到3.5小时的变化影响。评估指标包括均方根误差(RMSE)、中位误差及声道变量,并新增喉部高度测量。结果表明,平均RMSE为1.48毫米,优于1.62毫米像素尺寸,验证了基于RT-MRI实现完整声道逆向建模的可行性。

原文摘要 · Abstract (English)

Articulatory-to-acoustic inversion strongly depends on the type of data used. While most previous studies rely on EMA, which is limited by the number of sensors and restricted to accessible articulators, we propose an approach aiming at a complete inversion of the vocal tract, from the glottis to the lips. To this end, we used approximately 3.5 hours of RT-MRI data from a single speaker. The innovation of our approach lies in the use of articulator contours automatically extracted from MRI images, rather than relying on the raw images themselves. By focusing on these contours, the model prioritizes the essential geometric dynamics of the vocal tract while discarding redundant pixel-level information. These contours, alongside denoised audio, were then processed using a Bi-LSTM architecture. Two experiments were conducted: (1) the analysis of the impact of the audio embedding, for which three types of embeddings were evaluated as input to the model (MFCCs, LCCs, and HuBERT), and (2) the study of the influence of the dataset size, which we varied from 10 minutes to 3.5 hours. Evaluation was performed on the test data using RMSE, median error, as well as Tract Variables, to which we added an additional measurement: the larynx height. The average RMSE obtained is 1.48\,mm, compared with the pixel size (1.62\,mm). These results confirm the feasibility of a complete vocal-tract inversion using RT-MRI data.

语音建模医学影像深度学习声道逆向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。