arXiv:2603.11847eess.AS2026-03

用语音和发音信息重建声腔结构,手动校正后效果接近基准模型。

Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data

  • 分三阶段引入发音信息:自动转录、对齐分割、专家手动修正
  • 手动修正后重建精度逼近基于MFCC的基线模型
  • 适合语音合成与声学建模研究者参考

发音声学反演旨在从语音信号中重建完整的声腔几何结构。本文对比了多种发音分割精度等级,并与先前工作提出的基于梅尔频率倒谱系数(MFCC)的基线方法进行比较。所有方法均基于去噪语音信号,通过三个逐步细化的发音信息层次来探究其影响:未经校正的自动转录、时间对齐的发音分割,以及对齐后的专家手动修正。模型训练目标是预测从声腔MRI图像中提取的发音轮廓,采用自动轮廓追踪方法。结果表明,在依赖发音表示的模型中,对齐后的人工修正表现最佳,性能接近基线模型。

原文摘要 · Abstract (English)

Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a comparison to the baseline introduced in our previous work, which is based on Mel-Frequency Cepstral Coefficients (MFCCs). All the approaches considered are based on a denoised speech signal and aim to investigate the impact of incorporating phonetic information through three successive levels: an uncorrected automatic transcription, a temporally aligned phonetic segmentation, and an expert manual correction following alignment. The models are trained to predict articulatory contours extracted from vocal tract MRI images using an automatic contour tracking method. The results show that, among the models relying on phonetic representations, manual correction after alignment yields the best performance, approaching that of the baseline.

语音生成声腔重建发音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。