arXiv:2502.12835cs.CL2025-02ACL被引 7

子词模型学词能力差,但意外度可揭示隐藏的词汇学习能力。

Subword models struggle with word learning, but surprisal hides it

  • 用词义判断任务测试模型词汇学习能力
  • 字符模型准确率超90%,子词模型需上下文才达标
  • 适合研究语言习得底层机制的研究者参考

我们通过心理语言学中的词汇判断任务,研究了子词与字符级语言模型在词汇学习上的表现。尽管子词语言模型在区分单词与非单词时准确率较低,字符语言模型则能轻松且一致地完成该任务。只有在提供额外上下文信息后,子词模型的表现才接近字符模型。此外,在分析词级和句法学习轨迹时发现,字符模型中词汇学习先于句法学习,而子词模型中两者同时发生。这质疑了子词模型在建模语言习得过程中的适用性,并将字符模型定位为研究句法以下层级语言机制的可行替代方案。

原文摘要 · Abstract (English)

We study word learning in subword and character language models with the psycholinguistic lexical decision task. While subword LMs struggle to discern words and non-words with high accuracy, character LMs solve this task easily and consistently. Only when supplied with further contexts do subword LMs perform similarly to character models. Additionally, when looking at word-level and syntactic learning trajectories, we find that both processes are separable in character LMs. Word learning happens before syntactic learning, whereas both occur simultaneously in subword LMs. This raises questions about the adequacy of subword LMs for modeling language acquisition and positions character LMs as a viable alternative to study processes below the syntactic level.

语言模型词汇学习认知建模字符级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。