arXiv:2601.04469cs.CLcs.IR2026-01中稿 · the 10th Internati…被引 2

自参考方法生成乌拉尔语系形态词典,优化子词分词器的词汇量选择。

SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers

  • 基于最小描述长度思想,用内部结构线索筛选复合词形态
  • 发现芬兰语等语言在8k-256k词汇量下存在性能拐点
  • 为低资源语种提供可复现的分词器优化工具和数据

子词分词质量对大语言模型至关重要,但评估乌拉尔语系等形态丰富的语言时缺乏纯净的形态词典。我们提出SampoNLP,一个无需语料库的形态词典生成工具,采用受信息论启发的自参考原子性评分,通过内部结构线索过滤复合形式,适用于低资源场景。利用SampoNLP为芬兰语、匈牙利语和爱沙尼亚语生成高纯度词典,系统评估了不同词汇量(8k–256k)下BPE分词器的表现。提出统一指标集成性能得分(IPS),平衡形态覆盖与过度切分问题。通过分析IPS曲线,识别出收益递减的“拐点”,首次提供这些语言最优词汇量(k)的实证建议。研究不仅提供实用指导,还定量揭示标准BPE在高度黏着语言中的局限性。SampoNLP库及所有生成资源已公开:https://github.com/AragonerUA/SampoNLP

原文摘要 · Abstract (English)

The quality of subword tokenization is critical for Large Language Models, yet evaluating tokenizers for morphologically rich Uralic languages is hampered by the lack of clean morpheme lexicons. We introduce SampoNLP, a corpus-free toolkit for morphological lexicon creation using MDL-inspired Self-Referential Atomicity Scoring, which filters composite forms through internal structural cues - suited for low-resource settings. Using the high-purity lexicons generated by SampoNLP for Finnish, Hungarian, and Estonian, we conduct a systematic evaluation of BPE tokenizers across a range of vocabulary sizes (8k-256k). We propose a unified metric, the Integrated Performance Score (IPS), to navigate the trade-off between morpheme coverage and over-splitting. By analyzing the IPS curves, we identify the "elbow points" of diminishing returns and provide the first empirically grounded recommendations for optimal vocabulary sizes (k) in these languages. Our study not only offers practical guidance but also quantitatively demonstrates the limitations of standard BPE for highly agglutinative languages. The SampoNLP library and all generated resources are made publicly available: https://github.com/AragonerUA/SampoNLP

分词器形态分析低资源语言建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。