arXiv:2411.07854cs.CLcs.AI2024-11被引 26

打造2000亿词的葡萄牙语语料库,训练出性能更强的生成模型

Tucano: Advancing Neural Text Generation for Portuguese

  • 用2000亿词构建GigaVerbo语料库,训练Tucano系列生成模型
  • 在多个葡萄牙语评测中表现优于或媲美同类模型
  • 开源所有资源,适合研究低资源语言生成的学者使用

近年来自然语言处理取得显著进展,但当前深度学习语言建模需大量数据与计算资源。这种高数据依赖性导致语言间差距扩大:主流发展集中在高资源语言,而低资源语言难以达到同等性能。本研究旨在推动葡萄牙语神经文本生成发展,提出GigaVerbo语料库——由去重后的葡萄牙语文本拼接而成,总量达2000亿词。基于此语料,我们训练了一系列解码器-变换器模型(Tucano)。这些模型在多个葡萄牙语基准测试中表现等同或优于其他同规模葡萄牙语及多语言模型。评估还发现,现有葡萄牙语NLP社区常用基准测试的结果与训练时词元数量增长之间相关性极弱,揭示了当前评估方式在评估生成型语言模型时的局限性。本研究所有衍生成果均开源发布于GitHub和Hugging Face。

原文摘要 · Abstract (English)

Significant advances have been made in natural language processing in recent years. However, our current deep learning approach to language modeling requires substantial resources in terms of data and computation. One of the side effects of this data-hungry paradigm is the current schism between languages, separating those considered high-resource, where most of the development happens and resources are available, and the low-resource ones, which struggle to attain the same level of performance and autonomy. This study aims to introduce a new set of resources to stimulate the future development of neural text generation in Portuguese. In this work, we document the development of GigaVerbo, a concatenation of deduplicated Portuguese text corpora amounting to 200 billion tokens. Via this corpus, we trained a series of decoder-transformers named Tucano. Our models perform equal or superior to other Portuguese and multilingual language models of similar size in several Portuguese benchmarks. The evaluation of our models also reveals that model performance on many currently available benchmarks used by the Portuguese NLP community has little to no correlation with the scaling of token ingestion during training, highlighting the limitations of such evaluations when it comes to the assessment of Portuguese generative language models. All derivatives of our study are openly released on GitHub and Hugging Face. See https://nkluge-correa.github.io/Tucano/

文本生成葡萄牙语大模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。