用字符级词汇表训练的小模型也能实现强语言能力。
Small Language Models Also Work With Small Vocabularies: Probing the Linguistic Abilities of Grapheme- and Phoneme-Based Baby Llamas
- 采用字符级词表,避免子词分词限制。
- 小模型在句法和新词任务上表现优异。
- 拼音模型接近字形模型,适合语言习得研究。
近期研究探讨语言模型是否能从发展性合理的数据量中学习人类般的语言泛化与表征。然而,这些模型处理的基本语言单位由子词分词决定,限制了其在词级及以下层次学习建模的有效性。本文探索无分词、基于音素和字形的语言模型潜力。我们证明,使用字符级词汇表训练的微型Llama模型可在标准句法及新颖词汇/语音基准测试中取得强劲表现。此外,音素模型在标准任务和新评估中几乎达到字形模型水平。结果表明,构建更符合语言习得过程的计算模型具有前景,尤其适用于语言获取与处理的计算研究。
原文摘要 · Abstract (English)
Recent work investigates whether LMs learn human-like linguistic generalizations and representations from developmentally plausible amounts of data. Yet, the basic linguistic units processed in these LMs are determined by subword-based tokenization, which limits their validity as models of learning at and below the word level. In this paper, we explore the potential of tokenization-free, phoneme- and grapheme-based language models. We demonstrate that small models based on the Llama architecture can achieve strong linguistic performance on standard syntactic and novel lexical/phonetic benchmarks when trained with character-level vocabularies. We further show that phoneme-based models almost match grapheme-based models in standard tasks and novel evaluations. Our findings suggest a promising direction for creating more linguistically plausible language models that are better suited for computational studies of language acquisition and processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。