用实时MRI数据实现舌形全程重建,精度达2.21毫米。
Complete reconstruction of the tongue contour through acoustic to articulatory inversion using real-time MRI data
- 基于语音信号与舌部轮廓的反演网络,融合双方向LSTM和自编码器。
- 在1帧MFCC特征下,舌形重建中位误差为2.21毫米。
- 适合语音合成、语言康复等需精确发音控制的研究者。
声学到发音器官的反演是语音处理中的关键挑战,应用涵盖语音合成及语言学习与康复反馈系统。近年来,深度学习方法仅能反演少数几个可接触发音部位的几何位置,无法还原从舌根到舌尖的完整舌形。本文利用高质量实时MRI数据追踪舌轮廓,以未结构化的语音信号和舌轮廓作为输入,探索了多种基于双向门控循环单元(Bi-MSTM)的架构,包括是否使用自编码器降低潜在空间维度、是否结合音素分割。结果表明,在仅依赖1帧MFCC特征(含静态、一阶差分和二阶差分倒谱特征)的情况下,舌轮廓可实现2.21毫米(即1.37像素)的中位重建精度。
原文摘要 · Abstract (English)
Acoustic articulatory inversion is a major processing challenge, with a wide range of applications from speech synthesis to feedback systems for language learning and rehabilitation. In recent years, deep learning methods have been applied to the inversion of less than a dozen geometrical positions corresponding to sensors glued to easily accessible articulators. It is therefore impossible to know the shape of the whole tongue from root to tip. In this work, we use high-quality real-time MRI data to track the contour of the tongue. The data used to drive the inversion are therefore the unstructured speech signal and the tongue contours. Several architectures relying on a Bi-MSTM including or not an autoencoder to reduce the dimensionality of the latent space, using or not the phonetic segmentation have been explored. The results show that the tongue contour can be recovered with a median accuracy of 2.21 mm (or 1.37 pixel) taking a context of 1 MFCC frame (static, delta and double-delta cepstral features).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。