arXiv:2411.11171cs.CLcs.AI2024-11ACL被引 15

开源两款自研德语小模型,性能媲美主流同类。

LLäMmlein: Transparent, Compact and Competitive German-Only Language Models from Scratch

  • 从零构建德语专属模型与分词器,全程透明可复现。
  • 120M和1B参数模型在SuperGLEBer上表现媲美甚至超越同规模模型。
  • 揭示模型大小与性能关系,为未来研发提供资源分配参考。

我们从零开始透明训练并发布两款纯德语解码模型:LLäMmlein 120M 和 1B,同时公开训练数据,供德语自然语言处理研究社区使用。训练过程包括大规模数据预处理、自定义德语分词器构建、模型训练及多基准测试评估。通过SuperGLEBer基准持续监控多个检查点的学习动态。在SuperGLEBer上,两款模型均表现竞争力,稳定达到或超过同等参数量的现有先进模型。结果表明模型质量随规模提升符合预期,但部分任务性能早期即达上限,为后续模型开发中的资源分配提供了宝贵洞见。

原文摘要 · Abstract (English)

We create two German-only decoder models, LLäMmlein 120M and 1B, transparently from scratch and publish them, along with the training data, for the German NLP research community to use. The model training involved several key steps, including extensive data preprocessing, the creation of a custom German tokenizer, the training itself, as well as the evaluation of the final models on various benchmarks. Throughout the training process, multiple checkpoints were saved and analyzed using the SuperGLEBer benchmark to monitor the models' learning dynamics. Compared to state-of-the-art models on the SuperGLEBer benchmark, both LLäMmlein models performed competitively, consistently matching or surpassing models with similar parameter sizes. The results show that the models' quality scales with size as expected, but performance improvements on some tasks plateaued early, offering valuable insights into resource allocation for future model development.

德语模型小模型自研模型透明训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。