arXiv:2510.20386cs.CL2025-10被引 1

新BERT模型专攻希伯来语,性能全面超越现有方法。

NeoDictaBERT: Pushing the Frontier of BERT models for Hebrew

  • 基于NeoBERT架构改造,专为希伯来语优化的BERT模型。
  • 在几乎所有希伯来语基准上表现更优,多任务性能领先。
  • 适合希伯来语NLP研究者与需要多语言检索的应用场景。

自发布以来,BERT模型在多种任务中表现出色,尽管参数量相对较小(BERT-base约1亿)。然而,其架构已落后于Llama3、Qwen3等新型Transformer模型。近期,ModernBERT和NeoBERT在英文基准上取得显著进步,并大幅扩展了上下文窗口。受此启发,我们提出NeoDictaBERT与NeoDictaBERT-bilingual:基于NeoBERT架构的BERT类模型,专注于希伯来语文本。这些模型在几乎所有希伯来语基准上均优于现有方法,为下游任务提供强大基础。尤其值得注意的是,NeoDictaBERT-bilingual在检索任务中表现优异,超越同规模多语言模型。本文详细描述训练过程并报告多基准结果。我们已将模型开源,旨在推动希伯来语自然语言处理的研究与发展。

原文摘要 · Abstract (English)

Since their initial release, BERT models have demonstrated exceptional performance on a variety of tasks, despite their relatively small size (BERT-base has ~100M parameters). Nevertheless, the architectural choices used in these models are outdated compared to newer transformer-based models such as Llama3 and Qwen3. In recent months, several architectures have been proposed to close this gap. ModernBERT and NeoBERT both show strong improvements on English benchmarks and significantly extend the supported context window. Following their successes, we introduce NeoDictaBERT and NeoDictaBERT-bilingual: BERT-style models trained using the same architecture as NeoBERT, with a dedicated focus on Hebrew texts. These models outperform existing ones on almost all Hebrew benchmarks and provide a strong foundation for downstream tasks. Notably, the NeoDictaBERT-bilingual model shows strong results on retrieval tasks, outperforming other multilingual models of similar size. In this paper, we describe the training process and report results across various benchmarks. We release the models to the community as part of our goal to advance research and development in Hebrew NLP.

BERT希伯来语多语言NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。