探索大模型能否通过舌位图像理解元音发音机制
Tonguescape: Exploring Language Models Understanding of Vowel Articulation
- 用实时核磁数据构建图文数据集,测试模型对舌位与元音关系的理解
- 有参考示例时模型能关联舌位与元音,无参考则表现不佳
- 为语言学习和语音认知研究提供可解释的AI工具
元音主要由舌位决定。人类通过自身经验与如MRI等客观观测手段掌握了元音发音的舌位特征,这些知识有助于语言学习者掌握发音。由于语言模型(LMs)在包含语言学和医学领域的海量数据上训练,初步研究表明其具备解释元音发音机制的能力。然而,多模态语言模型(如视觉语言模型)是否能将文本信息与视觉信息对齐尚不明确。本文基于现有实时核磁数据,构建了视频与图像数据集,探究视觉信息下语言模型对元音发音中舌位的理解能力。结果表明:在提供参考示例的情况下,语言模型表现出理解元音与舌位关联的潜力;但缺乏参考时则难以建立对应关系。相关代码已开源于GitHub。
原文摘要 · Abstract (English)
Vowels are primarily characterized by tongue position. Humans have discovered these features of vowel articulation through their own experience and explicit objective observation such as using MRI. With this knowledge and our experience, we can explain and understand the relationship between tongue positions and vowels, and this knowledge is helpful for language learners to learn pronunciation. Since language models (LMs) are trained on a large amount of data that includes linguistic and medical fields, our preliminary studies indicate that an LM is able to explain the pronunciation mechanisms of vowels. However, it is unclear whether multi-modal LMs, such as vision LMs, align textual information with visual information. One question arises: do LMs associate real tongue positions with vowel articulation? In this study, we created video and image datasets from the existing real-time MRI dataset and investigated whether LMs can understand vowel articulation based on tongue positions using vision-based information. Our findings suggest that LMs exhibit potential for understanding vowels and tongue positions when reference examples are provided while they have difficulties without them. Our code for dataset building is available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。