研究变压器模型如何处理土耳其语和希伯来语复杂动词形态,揭示分词策略的关键作用。
Modelling the Morphology of Verbal Paradigms: A Case Study in the Tokenization of Turkish and Hebrew
- 对比原子分词与子词分词对动词形态的捕捉效果
- 希伯来语中单语模型在词素感知分词下表现更优
- 合成数据上所有模型性能均提升
我们研究了变压器模型在土耳其语和现代希伯来语中对复杂动词形态的表征能力,重点关注分词策略的影响。基于自然数据的Blackbird Language Matrices任务显示:对于具有透明形态标记的土耳其语,单语和多语模型在原子分词或子词分词下均表现良好;而希伯来语中,使用字符级分词的多语模型无法捕捉非连接性形态,但采用词素感知分词的单语模型表现优异。在更合成的数据集上,所有模型性能均有提升。
原文摘要 · Abstract (English)
We investigate how transformer models represent complex verb paradigms in Turkish and Modern Hebrew, concentrating on how tokenization strategies shape this ability. Using the Blackbird Language Matrices task on natural data, we show that for Turkish -- with its transparent morphological markers -- both monolingual and multilingual models succeed, either when tokenization is atomic or when it breaks words into small subword units. For Hebrew, instead, monolingual and multilingual models diverge. A multilingual model using character-level tokenization fails to capture the language non-concatenative morphology, but a monolingual model with morpheme-aware segmentation performs well. Performance improves on more synthetic datasets, in all models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。