arXiv:2508.02527cs.CLcs.LG2025-08被引 1

LLaMA 3.2 能无监督学会人类发音规律,靠的是内部的音素表征。

I Have No Mouth, and I Must Rhyme: Uncovering Internal Phonetic Representations in LLaMA 3.2

  • 通过分析模型潜空间,发现其对音素有高阶组织结构。
  • 识别出一个促进押韵的‘音素移动头’,输出接近标准国际音标元音图。
  • 无需语音监督,模型自学到类人发音知识,适合研究语言表征者阅读。

大型语言模型在无需显式语音或听觉训练的情况下,也能完成押韵等语音任务。本文研究了 Llama-3.2-1B-Instruct 模型如何表征词级语音信息。结果表明,该模型内部存在丰富的音素表示,并在潜在空间中呈现出音素的高级组织结构。我们还发现了一个‘音素移动头’,它在押韵任务中增强语音信息。可视化该头的输出空间后发现,尽管存在差异,但模型学习到的元音分布与人类的标准国际音标(IPA)元音图高度相似,且未接受任何直接语音监督。

原文摘要 · Abstract (English)

Large language models demonstrate proficiency on phonetic tasks, such as rhyming, without explicit phonetic or auditory grounding. In this work, we investigate how \verb|Llama-3.2-1B-Instruct| represents token-level phonetic information. Our results suggest that Llama uses a rich internal model of phonemes to complete phonetic tasks. We provide evidence for high-level organization of phoneme representations in its latent space. In doing so, we also identify a ``phoneme mover head" which promotes phonetic information during rhyming tasks. We visualize the output space of this head and find that, while notable differences exist, Llama learns a model of vowels similar to the standard IPA vowel chart for humans, despite receiving no direct supervision to do so.

音素表征语言模型潜在空间押韵生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。