arXiv:2608.09767cs.CLcs.SD2026-08

用语音模型提取的音位特征,提升发音肌动影像的分类效果。

Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

论文配图:Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification
图 1 · 摘自论文原文
  • 从语音模型中提取音位结构特征,用于辅助肌动影像分析。
  • 在未见说话人和未见语音场景下,音位分类准确率显著提升。
  • 可解释的发音模式揭示了如/t/颤音化等语言现象,适合语音建模研究者。

实时磁共振成像使我们能够观察说话时声道的运动,但将这些运动模式映射到音素和音位类别仍具挑战性。本文研究了基于语音训练的PhonoQ模型能否为音动建模提供有用信息。具体地,从PhonoQ的Conformer模块中提取表示,其训练受制于对发音方式、发音部位、声带振动和元音特征的监督。结合同步的音频特征与肌动轮廓,对比了WavLM-large和HuBERT-large基线模型,以及融合PhonoQ特征的模型。在未见说话人和未见语音设置下,新特征提升了对发音方式、部位、声带振动、元音高度和元音后部等音位目标的宏平均F1分数,也提高了细粒度39个音素的分类性能。在仅使用轮廓的推理设置中,来自音频的教师监督带来稳定但有限的增益,表明音位信息可部分从同步音频传递至肌动模型。后验分析显示,模型表现出与/ t /颤音化、/t/-/r/后缩或塞擦音化、鼻音部位同化等一致的可解释表面敏感模式。

原文摘要 · Abstract (English)

Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based model trained to recognize structured phonological features, provides useful information for audio--articulatory modeling. Specifically, we extract representations from PhonoQ's Conformer module, whose training is shaped by supervision for manner, place, voicing, and vowel features. Using articulatory contours with synchronized audio-derived features, we compare WavLM-large and HuBERT-large baselines with models that incorporate PhonoQ-derived representations. Across unseen-speech and unseen-subject settings, these features improve macro-F1 for phonological targets including manner, place, voicing, vowel height, and vowel backness, and also improve fine-grained 39-phoneme classification. In a contour-only inference setting, audio-derived teacher supervision yields modest but consistent gains over contour-only training, indicating that phonological information from synchronized audio can be partially transferred to articulatory models. Finally, posterior analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation.

语音建模肌动影像音位特征多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。