NorBERTo是用3310亿词元训练的葡萄牙语BERT模型,性能领先。
NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
- 基于ModernBERT架构,支持长文本和高效注意力机制。
- 在PLUE数据集上达0.9191 F1(MRPC)和0.7689准确率(RTE)。
- 适合实际部署,可作为葡萄牙语NLP任务的通用骨干模型。
高质量语料对推进葡萄牙语自然语言处理至关重要。我们基于此前的编码器模型如BERTimbau和Albertina PT-BR,提出NorBERTo,一个基于ModernBERT架构的现代编码器,具备长上下文支持和高效注意力机制。该模型在Aurora-PT语料库上训练,该语料由来自多样化网络来源和现有多语言数据集的3310亿个GPT-2词元组成。我们在标准化数据集(如ASSIN 2和PLUE)上系统评估NorBERTo,涵盖语义相似性、文本蕴含和分类任务。在PLUE上,NorBERTo-large在所有评估的编码器中表现最佳,尤其在MRPC任务上达到0.9191 F1,在RTE任务上达到0.7689准确率。在ASSIN 2上,NorBERTo-large实现约0.904的蕴含F1,虽略低于Albertina-900M和BERTimbau-large,但表现优异。据我们所知,Aurora-PT是目前最大且公开可用的单语葡萄牙语语料库。NorBERTo是一个现代化、中等规模的编码器,设计用于现实部署:易于微调,服务效率高,适用于检索增强生成及其他下游葡萄牙语NLP系统。
原文摘要 · Abstract (English)
High-quality corpora are essential for advancing Natural Language Processing (NLP) in Portuguese. Building on previous encoder-only models such as BERTimbau and Albertina PT-BR, we introduce NorBERTo, a modern encoder based on the ModernBERT architecture, featuring long-context support and efficient attention mechanisms. NorBERTo is trained on Aurora-PT, a newly curated Brazilian Portuguese corpus comprising 331 billion GPT-2 tokens collected from diverse web sources and existing multilingual datasets. We systematically benchmark NorBERTo against Strong baselines on semantic similarity, textual entailment and classification tasks using standardized datasets such as ASSIN 2 and PLUE. On PLUE, NorBERTo-large achieves the best results among the encoder models we evaluated, notably reaching 0.9191 F1 on MRPC and 0.7689 accuracy on RTE. On ASSIN 2, NorBERTo-large attains the highest entailment F1 (~0.904) among all encoders considered, although Albertina-900M and BERTimbau-large still hold an advantage. To the best of our knowledge, Aurora-PT is currently the largest openly available monolingual Portuguese corpus, surpassing previous resources. NorBERTo provides a modern, mid-sized encoder designed for realistic deployment scenarios: it is straight-forward to fine-tune, efficient to serve, and well suited as a backbone for retrieval-augmented generation and other downstream Portuguese NLP systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。