arXiv:2504.13308cs.SDcs.CL2025-04被引 1

用机器学习从声音反推发音器官运动,助力语言治疗与语音研究

Acoustic to Articulatory Inversion of Speech; Data Driven Approaches, Challenges, Applications, and Future Scope

  • 基于十年内机器学习方法,从语音信号反推舌头等发音器官位置
  • 使用EMA、rtMRI等多模态数据,通过相关系数与均方误差评估精度
  • 适合语音病理治疗、发音训练系统开发人员参考

本文综述了2011至2021年间声学到发音器官反演(AAI)的数据驱动方法及其应用。涵盖说话人相关与无关的AAI研究,目标包括发音近似、发音特征空间选择、自动语音识别(ASR)、声学与发音特征相关性探索,以及计算机辅助语言训练框架。研究采用同步录制的语音(wav)与医学成像数据,如电磁性发音仪(EMA)、电腭图(EPG)、喉镜、电子喉图(EGG)、X射线动态成像、超声及实时磁共振成像(rtMRI)。所有方法均基于机器学习,性能评估主要采用相关系数(CC)、均方根误差(RMSE),也考虑均方误差(MSE)与平均格式误差(MFE)。AAI模型可生成可视化发音轨迹反馈,尤其适用于舌位变化的直观呈现,为语音障碍患者提供语音治疗与发音训练支持。

原文摘要 · Abstract (English)

This review is focused on the data-driven approaches applied in different applications of Acoustic-to-Articulatory Inversion (AAI) of speech. This review paper considered the relevant works published in the last ten years (2011-2021). The selection criteria includes (a) type of AAI - Speaker Dependent and Speaker Independent AAI, (b) objectives of the work - Articulatory approximation, Articulatory Feature space selection and Automatic Speech Recognition (ASR), explore the correlation between acoustic and articulatory features, and framework for Computer-assisted language training, (c) Corpus - Simultaneously recorded speech (wav) and medical imaging models such as ElectroMagnetic Articulography (EMA), Electropalatography (EPG), Laryngography, Electroglottography (EGG), X-ray Cineradiography, Ultrasound, and real-time Magnetic Resonance Imaging (rtMRI), (d) Methods or models - recent works are considered, and therefore all the works are based on machine learning, (e) Evaluation - as AAI is a non-linear regression problem, the performance evaluation is mostly done by Correlation Coefficient (CC), Root Mean Square Error (RMSE), and also considered Mean Square Error (MSE), and Mean Format Error (MFE). The practical application of the AAI model can provide a better and user-friendly interpretable image feedback system of articulatory positions, especially tongue movement. Such trajectory feedback system can be used to provide phonetic, language, and speech therapy for pathological subjects.

语音合成发音反演语言治疗机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。