跨语言语音反演模型可从语音中还原发音器官与声源特征。
Towards Language-Agnostic Speech Inversion

- 基于美式英语数据训练,同时估计发音部位和声源参数。
- 在法语和俄语上分别达到0.83和0.74的相关性。
- 首次验证了语音反演模型的跨语言泛化能力。
语音信号中包含声道构型与声源激励的特征时间模式。以往研究证明,语音反演(SI)系统能从语音中恢复这些时间模式,包括舌位、唇部收缩等声道变量,以及基频、周期性与非周期性能量等声源信息。本文构建了一个同时估计声道变量与三种声源参数的SI系统,使用共录制的美式英语语音音频与发音运动数据进行训练,并通过评估其在未见语言上的表现,检验跨语言泛化能力。在未训练的法语和俄语上,估计值与真实测量值的相关系数分别达到0.83和0.74,涵盖声道变量与声源信息。
原文摘要 · Abstract (English)
Characteristic timing patterns are reflected in the acoustic speech signal, encompassing both vocal tract configuration and acoustic excitation. Previous studies have demonstrated that speech inversion (SI) systems can recover these timing patterns from speech, including oral tract variables (tongue and lip constrictions) and source information such as periodic and aperiodic energies and fundamental frequency. In this study, we develop an SI system that simultaneously estimates oral tract variables and three source information parameters trained on co-recorded American English speech audio and articulatory kinematics and investigate cross-linguistic generalizability by evaluating performance on previously unseen languages. Pearson product-moment correlation scores of 0.83 and 0.74 were achieved on untrained French and Russian respectively, across oral tract variables and source information when comparing estimated data with ground-truth measurements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。