为拉丁语设计基于形态的分词方法,提升低资源语言模型性能
Contextual morphologically-guided tokenization for Latin encoder models
- 基于形态信息优化拉丁语分词,兼顾语言学合理性
- 在4个下游任务上均提升性能,跨域文本效果更显著
- 适合研究低资源复杂形态语言的学者参考
分词是语言模型预训练的关键环节,但传统方法多关注信息论目标(如高压缩率、低碎片化),忽视语言学目标如形态对齐。这在形态丰富的语言中表现不佳,直接影响下游任务表现。本文针对拉丁语——一种形态丰富、预训练数据中等但词汇资源丰富的语言——提出形态引导的分词方法。实验表明,该方法在4个下游任务中均提升整体性能,尤其在跨域文本上表现更优,凸显模型泛化能力增强。研究证明,语言学资源对复杂形态语言建模具有实际价值。对于缺乏大规模预训练数据的低资源语言,构建和利用语言学资源是提升语言模型性能的有效替代路径。
原文摘要 · Abstract (English)
Tokenization is a critical component of language model pretraining, yet standard tokenization methods often prioritize information-theoretical goals like high compression and low fertility rather than linguistic goals like morphological alignment. In fact, they have been shown to be suboptimal for morphologically rich languages, where tokenization quality directly impacts downstream performance. In this work, we investigate morphologically-aware tokenization for Latin, a morphologically rich language that is medium-resource in terms of pretraining data, but high-resource in terms of curated lexical resources -- a distinction that is often overlooked but critical in discussions of low-resource language modeling. We find that morphologically-guided tokenization improves overall performance on four downstream tasks. Performance gains are most pronounced for out of domain texts, highlighting our models' improved generalization ability. Our findings demonstrate the utility of linguistic resources to improve language modeling for morphologically complex languages. For low-resource languages that lack large-scale pretraining data, the development and incorporation of linguistic resources can serve as a feasible alternative to improve LM performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。