首次实现从声音还原完整声道运动,精度达1.65毫米。
Reconstruction of the Complete Vocal Tract Contour Through Acoustic to Articulatory Inversion Using Real-Time MRI Data
- 用实时MRI数据训练双向LSTM模型,同步反演声道各部位
- 测试集平均误差1.65毫米,接近像素分辨率1.62毫米
- 适合语音合成、发音研究及临床语音分析领域
声学到发音的逆问题通常局限于声道局部,因传统电磁动觉测量(EMA)需在可触及的发音器官上粘贴传感器。本文提出的声学到发音逆模型,实现了从声门到整个舌头、软腭直至双唇的完整声道重构。该模型基于超过3小时的实时动态MRI语音数据库,输入为降噪后的语音信号与自动分割的发音器轮廓。采用多种基于双向LSTM的方法,分别对各发音器单独建模或联合建模。据我们所知,这是首个完整的声道逆向重建工作。测试集平均均方根误差(RMSE)为1.65毫米,与像素尺寸1.62毫米相当。
原文摘要 · Abstract (English)
Acoustic to articulatory inversion has often been limited to a small part of the vocal tract because the data are generally EMA (ElectroMagnetic Articulography) data requiring sensors to be glued to easily accessible articulators. The presented acoustic to articulation model focuses on the inversion of the entire vocal tract from the glottis, the complete tongue, the velum, to the lips. It relies on a realtime dynamic MRI database of more than 3 hours of speech. The data are the denoised speech signal and the automatically segmented articulator contours. Several bidirectional LSTM-based approaches have been used, either inverting each articulator individually or inverting all articulators simultaneously. To our knowledge, this is the first complete inversion of the vocal tract. The average RMSE precision on the test set is 1.65 mm to be compared with the pixel size which is 1.62 mm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。