检验阿拉伯语分词与大模型如何处理根形构词,发现分词对生成能力无决定作用。
Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs
- 对比七种分词器与大模型的形态忠实度,用标准分割作为基准
- 大模型在根形生成任务中表现不一,但分词对齐非必要条件
- 揭示分词设计未必提升生成性能,适合研究多语言模型机制者阅读
本研究探究大型语言模型(LLMs)及其分词方案对阿拉伯语根形构词结构的表征与生成能力,检验其是否捕捉真实形态结构,或仅依赖表面记忆。阿拉伯语复杂的非拼接式形态系统为分析模型处理复杂形态提供了丰富测试场景。研究首先评估阿拉伯语及多语言分词器在标准分割上的形态忠实度,随后使用新构建的测试集分析七种阿拉伯语中心及多语言大模型在根形生成任务中的表现。结果表明,分词器的形态对齐既非生成能力的必要条件,也非充分条件,挑战了形态分词在下游性能中的关键作用。
原文摘要 · Abstract (English)
This work investigates how effectively large language models (LLMs) and their tokenization schemes represent and generate Arabic root-pattern morphology, probing whether they capture genuine morphological structure or rely on surface memorization. Arabic morphological system provides a rich testbed for analyzing how LLMs handle complex, non-concatenative forms and how tokenization choices influence this process. Our study begins with an evaluation of morphological fidelity across Arabic and multilingual tokenizers against gold-standard segmentation, followed by an analysis of LLM performance in productive root-pattern generation using a newly developed test set. Our findings across seven Arabic-centric and multilingual LLMs and their respective tokenizers reveal that tokenizer morphological alignment is not necessary nor sufficient for morphological generation, which questions the role of morphological tokenization in downstream performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。