语言模型如何用发音信息选词形?
How Do Language Models Represent and Use Phonological Information for Allomorph Selection?
- 在英文不定冠词a/an选择中,发音条件由触发词嵌入向量中的单一线性方向编码
- 模型在生成时会预判后续触发词的发音特征并据此选择冠词,验证了因果作用
- 该机制不仅适用于英语,还扩展到其他语言的词形选择和发音判断任务
语言模型在分词文本上训练,隐藏了词语的发音结构,但仍能可靠生成受发音条件制约的词形。目前尚不清楚其依赖的是特定词条记忆还是规则式泛化,若为后者,这种泛化如何实现仍不明确。本文探究语言模型是否表征发音条件及其在词形选择中的因果作用。针对英语不定冠词a/an,研究发现发音条件在触发词嵌入中沿单一线性方向编码,该方向在词元级wug测试中直接驱动冠词选择;在预测冠词的位置,模型会提前预判触发词发音特征,并据此选择冠词。进一步考察该规则泛化是否超越英语冠词选择,延伸至其他语言的词形选择及显式发音判断任务。结果提供了一个关于语言模型中发音条件词形选择的机制解释,并将生成阶段的能力与显式元语言判断区分开来。
原文摘要 · Abstract (English)
Language models are trained on tokenized text that obscures the sound structure of words, yet they reliably produce morphemes whose form is phonologically conditioned. It remains unclear whether they rely on item-specific memorization or rule-like generalization and, if the latter, how that generalization is implemented. We therefore ask whether this phonological condition is represented within language models and how it is causally used for allomorph selection. For the English indefinite article a/an, we show that the phonological condition is encoded along a single linear direction in trigger-token embeddings, that this direction causally drives article selection in token-level wug tests, and that, at the article-prediction position, the model forecasts the upcoming trigger token and uses the forecasted trigger's phonological feature to choose the article. We then ask whether this rule-like generalization extends beyond English article selection, both to allomorph selection in other languages and to explicit phonological judgment. Together, these results provide a mechanistic account of phonologically conditioned allomorph selection in language models, and dissociate this generation-time ability from explicit metalinguistic judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。