通过引入鼻音度提升语音逆向建模,更准确还原口鼻腔运动。
Enhancing Acoustic-to-Articulatory Speech Inversion by Incorporating Nasality
- 融合鼻音度与声源特征,协同建模口腔与软腭运动。
- 相比基线模型,口腔参数估计提升5%,鼻音度预测提升9%。
- 适合语音生理建模、发音障碍研究等场景使用。
语音由唇、舌、软腭和声门等声道约束器官协同产生。以往研究构建了语音逆向(SI)系统,用于恢复唇舌等口腔区域的运动参数(称为口部变量,TVs),并通过引入周期性与非周期性能量、基频(F0)等声源信息作为声门控制的代理。对比高速鼻咽镜数据发现,鼻音度(nasalance)可作为真实标签,以该标签训练的SI系统能可靠还原美国英语说话者的软腭运动模式。本文比较两种训练方法:基础模型独立估计口部变量和鼻音度,以及将口部变量、声源特征与鼻音度联合建模的协同模型。结果表明,协同模型在口部变量估计上相对提升5%,在鼻音度估计上相对提升9%。
原文摘要 · Abstract (English)
Speech is produced through the coordination of vocal tract constricting organs: lips, tongue, velum, and glottis. Previous works developed Speech Inversion (SI) systems to recover acoustic-to-articulatory mappings for lip and tongue constrictions, called oral tract variables (TVs), which were later enhanced by including source information (periodic and aperiodic energies, and F0 frequency) as proxies for glottal control. Comparison of the nasometric measures with high-speed nasopharyngoscopy showed that nasalance can serve as ground truth, and that an SI system trained with it reliably recovers velum movement patterns for American English speakers. Here, two SI training approaches are compared: baseline models that estimate oral TVs and nasalance independently, and a synergistic model that combines oral TVs and source features with nasalance. The synergistic model shows relative improvements of 5% in oral TVs estimation and 9% in nasalance estimation compared to the baseline models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。