arXiv:2510.21372cs.CL2025-10被引 2

为希伯来语构建了首个大规模鲁棒编码器,提升多项语言任务表现。

HalleluBERT: Let Every Token That Has Meaning Bear Its Weight

  • 基于罗伯塔架构从头训练,使用专有字节级BPE分词器
  • 在3个希伯来语基准上均超越现有模型,平均得分最高
  • 开源权重与分词器,支持可复现的希伯来语NLP研究

基于Transformer的模型已推动自然语言处理发展,但希伯来语仍缺乏大规模训练且发布基线与大型变体的RoBERTa编码器。本文提出HalleluBERT,一个基于RoBERTa的编码器家族,从零开始在49.1~GB去重的希伯来语网络文本与维基百科数据上训练,并采用专有的字节级BPE词汇表。在希伯来语命名实体识别(BMC、NEMO)和情感分类(SMCD)的原生基准测试中,HalleluBERT优于单语及多语基线模型,在三个基准上的未加权平均得分最高。我们以MIT许可证发布模型权重与分词器,以支持可复现的希伯来语自然语言处理研究。

原文摘要 · Abstract (English)

Transformer-based models have advanced NLP, yet Hebrew still lacks a RoBERTa encoder that is trained at scale and released in both base and large variants. We present HalleluBERT, a RoBERTa-based encoder family trained from scratch on 49.1~GB of deduplicated Hebrew web text and Wikipedia using a Hebrew-specific byte-level BPE vocabulary. On native Hebrew benchmarks for named entity recognition (BMC, NEMO) and sentiment classification (SMCD), HalleluBERT outperforms monolingual and multilingual baselines, and yields the highest unweighted mean score across the three benchmarks. We release model weights and tokenizer under the MIT license to support reproducible Hebrew NLP research.

希伯来语RoBERTaNLP编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。