arXiv:2603.15005cs.CL2026-03被引 1

为拉脱维亚语训练专用编码器,性能超越现有模型。

Pretraining and Benchmarking Modern Encoders for Latvian

  • 基于RoBERTa、DeBERTaV3等架构,定制拉脱维亚语编码器。
  • 最佳模型lv-deberta-base(111M参数)表现领先于大模型和旧模型。
  • 适合拉脱维亚语NLP研究与实际应用,资源已公开。

编码器型Transformer在实际自然语言处理任务中仍至关重要。尽管多语言模型的进步提升了跨语言能力,但低资源语言如拉脱维亚语在预训练语料中仍严重缺失,且现存的单语拉脱维亚语编码器极少。本文通过基于RoBERTa、DeBERTaV3和ModernBERT架构的拉脱维亚语专用编码器预训练,包括长上下文变体,并在多样化的拉脱维亚语诊断与语言学基准上进行评估。所提模型在性能上可与现有单语及多语编码器媲美,同时受益于最新的架构与效率改进。最佳模型lv-deberta-base(111M参数)整体表现最强,优于更大规模的多语基线及先前的拉脱维亚语专用编码器。所有预训练模型与评估资源均已发布,以支持拉脱维亚语NLP的进一步研究与实际应用。

原文摘要 · Abstract (English)

Encoder-only transformers remain essential for practical NLP tasks. While recent advances in multilingual models have improved cross-lingual capabilities, low-resource languages such as Latvian remain underrepresented in pretraining corpora, and few monolingual Latvian encoders currently exist. We address this gap by pretraining a suite of Latvian-specific encoders based on RoBERTa, DeBERTaV3, and ModernBERT architectures, including long-context variants, and evaluating them across a diverse set of Latvian diagnostic and linguistic benchmarks. Our models are competitive with existing monolingual and multilingual encoders while benefiting from recent architectural and efficiency advances. Our best model, lv-deberta-base (111M parameters), achieves the strongest overall performance, outperforming larger multilingual baselines and prior Latvian-specific encoders. We release all pretrained models and evaluation resources to support further research and practical applications in Latvian NLP.

拉脱维亚语编码器预训练NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。