arXiv:2509.15689eess.IVcs.SD2025-09被引 2

用实时核磁分析发音时口腔动作,提升语音识别准确率与可解释性。

Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition

  • 提取舌、唇等部位的运动区域特征,结合原始视频和光流数据
  • 多特征融合模型错误率最低达0.34,显著优于单一特征
  • 揭示舌部与唇部运动对识别贡献最大,适合语音研究与临床应用

实时磁共振成像(rtMRI)可可视化发音过程中的声道动态,但信号维度高且噪声大,难以解析。本文研究从矢状面声道rtMRI视频中提取紧凑的时空发音动态表示,用于音素识别。对比三种特征类型:(1) 原始视频,(2) 光流,(3) 六个语言相关的感兴趣区域(ROIs)的发音器运动。分别训练基于每种表示的模型,并评估多特征组合效果。结果表明,多特征模型始终优于单特征基线,其中结合ROI与原始视频的模型取得最低音素错误率(PER)0.34。时序保真度实验显示模型依赖精细发音动态,ROI消融实验表明舌与唇的贡献显著。研究验证了rtMRI特征在提升识别精度与可解释性方面的价值,并为语音处理中利用发音数据提供有效策略。

原文摘要 · Abstract (English)

Real-time Magnetic Resonance Imaging (rtMRI) visualizes vocal tract action, offering a comprehensive window into speech articulation. However, its signals are high dimensional and noisy, hindering interpretation. We investigate compact representations of spatiotemporal articulatory dynamics for phoneme recognition from midsagittal vocal tract rtMRI videos. We compare three feature types: (1) raw video, (2) optical flow, and (3) six linguistically-relevant regions of interest (ROIs) for articulator movements. We evaluate models trained independently on each representation, as well as multi-feature combinations. Results show that multi-feature models consistently outperform single-feature baselines, with the lowest phoneme error rate (PER) of 0.34 obtained by combining ROI and raw video. Temporal fidelity experiments demonstrate a reliance on fine-grained articulatory dynamics, while ROI ablation studies reveal strong contributions from tongue and lips. Our findings highlight how rtMRI-derived features provide accuracy and interpretability, and establish strategies for leveraging articulatory data in speech processing.

语音识别核磁成像可解释性发音建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。