arXiv:2608.07727cs.CL2026-08

针对德拉维达语系,训练单语与多语模型对比,发现单语模型更优。

Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages

  • 分别训练四语单语模型和一语多语模型,使用不同词表大小的分词器
  • 单语模型在情感分类和命名实体识别上优于mGPT,分词效率更高
  • 适合低资源语言研究者参考,尤其关注南印度语言建模

德拉维达语系(主要为泰米尔语、泰卢固语、卡纳达语和马拉雅拉姆语)仅占多语言语言模型训练数据的一小部分,因此这些模型对各语言的真实能力尚不明确。本文从零训练了五个基于GPT-2架构的模型:四个分别为泰米尔语、泰卢固语、卡纳达语和马拉雅拉姆语的单语模型(各使用32K词表子词分词器),以及一个共享64K词表子词分词器的多语模型。所有模型均在清洗后的CC-100、Wikipedia和Samanantar数据集上训练。通过困惑度、每字节比特数、分词器效率及微调结果进行评估,并与mGPT对比。结果显示,单语模型在情感分类和命名实体识别任务中表现优于mGPT,且其分词器在所有测试语言中均比共享多语模型更高效。

原文摘要 · Abstract (English)

Dravidian languages, mainly Tamil, Telugu, Kannada, and Malayalam make up only a small part of the data used to train multilingual language models, so it's not clear how much per-language ability these models actually keep. I have trained five GPT-2 architecture models from scratch to compare four monolingual models (one each for Tamil, Telugu, Kannada, and Malayalam, each with its own 32K-vocabulary subword tokenizer) against one multilingual model sharing a 64K-vocabulary subword tokenizer across all four languages. All the 5 models are trained on cleaned CC-100, Wikipedia, and Samanantar data. I have tested the models on perplexity, bits-per-byte, tokenizer efficiency, and fine-tuning results which are compared against mGPT. The monolingual models outperform mGPT on sentiment classification and named entity recognition, and their tokenizers proved more efficient than the shared multilingual model across all the languages tested.

多语模型单语模型南亚语言分词效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。