用单说话人MRI数据分离语音与发音的贡献,揭示两者在音素识别中的差异。
Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
- 基于单说话人长时序MRI构建音频、视频及多模态模型,避免跨说话人干扰。
- 音频与多模态模型在发音方式上表现相似,但在发音部位上出现性能差异。
- 潜空间分析显示音素结构编码一致,但注意力权重揭示发音与语音时间上的不同。
尽管先前研究已利用实时MRI数据探索言语过程中声道运动的视听动态,但受限于多说话人语料库,这些研究难以学习声学与发音之间的精细关联。本研究采用长时序单说话人MRI语料,构建单模态音频与视频模型以及多模态模型用于音素识别,旨在解耦并解释各模态的贡献。音频与多模态模型在不同发音方式类别上表现相近,但在发音部位上出现差异。对模型潜空间的分析显示,音频与多模态模型均编码了相似的音素空间,而注意力权重则突显了特定音素在声学与发音时间上的差异。
原文摘要 · Abstract (English)
Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on multi-speaker corpora. This prevents such models from learning a detailed relationship between acoustics and articulation due to considerable cross-speaker variability. In this study, we develop unimodal audio and video models as well as multimodal models for phoneme recognition using a long-form single-speaker MRI corpus, with the goal of disentangling and interpreting the contributions of each modality. Audio and multimodal models show similar performance on different phonetic manner classes but diverge on places of articulation. Interpretation of the models' latent space shows similar encoding of the phonetic space across audio and multimodal models, while the models' attention weights highlight differences in acoustic and articulatory timing for certain phonemes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。