arXiv:2508.13058cs.CL2025-08被引 6

为土耳其语等复杂语言建立精准的分词评估标准

Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi

  • 提出新评估框架,用5项指标衡量分词效果
  • 发现特定语言分词比例比纯度更影响模型表现
  • 强调定制化分词对低资源语言至关重要

分词是自然语言处理中的基础预处理步骤,显著影响大语言模型捕捉语言与语义细节的能力。本研究针对形态丰富且资源匮乏的语言(如土耳其语)提出新型评估框架。基于土耳其MMLU(TR-MMLU)数据集——包含6200道来自土耳其教育系统的多选题,评估了不同分词器在词汇量、分词数、处理时间、语言特定分词比例(%TR)和分词纯度(%Pure)等方面的性能。这些新提出的指标用于衡量分词器保持语言结构的有效性。分析表明,语言特定分词比例与下游任务表现(如MMLU得分)的相关性更强于分词纯度。此外,单纯增加模型参数并不必然提升语言性能,凸显了针对特定语言定制分词方法的重要性。所提框架为形态复杂的语言建立了稳健且实用的分词标准。

原文摘要 · Abstract (English)

Tokenization is a fundamental preprocessing step in Natural Language Processing (NLP), significantly impacting the capability of large language models (LLMs) to capture linguistic and semantic nuances. This study introduces a novel evaluation framework addressing tokenization challenges specific to morphologically-rich and low-resource languages such as Turkish. Utilizing the Turkish MMLU (TR-MMLU) dataset, comprising 6,200 multiple-choice questions from the Turkish education system, we assessed tokenizers based on vocabulary size, token count, processing time, language-specific token percentages (\%TR), and token purity (\%Pure). These newly proposed metrics measure how effectively tokenizers preserve linguistic structures. Our analysis reveals that language-specific token percentages exhibit a stronger correlation with downstream performance (e.g., MMLU scores) than token purity. Furthermore, increasing model parameters alone does not necessarily enhance linguistic performance, underscoring the importance of tailored, language-specific tokenization methods. The proposed framework establishes robust and practical tokenization standards for morphologically complex languages.

分词评估土耳其语大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。