arXiv:2510.21364cs.CL2025-10被引 4

首个面向土耳其语的大型RoBERTa模型,助力语言研究与技术落地。

SindBERT, the Sailor: Charting the Seas of Turkish NLP

  • 从312GB土耳其语数据训练,推出base与large双版本编码器。
  • 在4项任务中表现优异,大模型在2项任务上夺冠但无持续增益。
  • 强调语料质量比数量更重要,适合土耳其语研究者使用。

Transformer模型已革新自然语言处理,但许多形态丰富的语言仍缺乏大规模预训练支持。本文提出SindBERT,首个基于RoBERTa的土耳其语大规模编码器,从零训练于312GB土耳其语文本(mC4、OSCAR23、Wikipedia),发布base与large两个版本,是首个公开可用的土耳其语纯编码器模型。我们在词性标注、命名实体识别、攻击性语言检测及TurBLiMP可接受性评估中测试其性能。结果表明,SindBERT在多数任务上表现媲美现有土耳其语和多语言模型,大版本在两项任务中达到最佳,但整体未见持续规模收益。这一平坦缩放趋势与XLM-R、EuroBERT一致,暗示当前土耳其语基准可能已达饱和。同时,与更小但更精炼的BERTurk对比显示,语料质量与多样性可超越数据量。SindBERT不仅作为开源资源推动土耳其语NLP发展,也为形态丰富语言中规模与语料设计的权衡提供实证参考。模型以MIT许可开放,支持fairseq与Huggingface格式。

原文摘要 · Abstract (English)

Transformer models have revolutionized NLP, yet many morphologically rich languages remain underrepresented in large-scale pre-training efforts. With SindBERT, we set out to chart the seas of Turkish NLP, providing the first large-scale RoBERTa-based encoder for Turkish. Trained from scratch on 312~GB of Turkish text (mC4, OSCAR23, Wikipedia), SindBERT is released in both base and large configurations, representing the first large-scale encoder-only language model available for Turkish. We evaluate SindBERT on part-of-speech tagging, named entity recognition, offensive language detection, and the TurBLiMP linguistic acceptability benchmark. Our results show that SindBERT performs competitively with existing Turkish and multilingual models, with the large variant achieving the best scores in two of four tasks but showing no consistent scaling advantage overall. This flat scaling trend, also observed for XLM-R and EuroBERT, suggests that current Turkish benchmarks may already be saturated. At the same time, comparisons with smaller but more curated models such as BERTurk highlight that corpus quality and diversity can outweigh sheer data volume. Taken together, SindBERT contributes both as an openly released resource for Turkish NLP and as an empirical case study on the limits of scaling and the central role of corpus composition in morphologically rich languages. The SindBERT models are released under the MIT license and made available in both fairseq and Huggingface formats.

土耳其语预训练模型RoBERTa语料质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。