arXiv:2605.11255cs.CL2026-05被引 1

首个支持长文本的开源希伯来语专家混合模型,推理效率高且性能强。

HEBATRON: A Hebrew-Specialized Open-Weight Mixture-of-Experts Language Model

论文配图:HEBATRON: A Hebrew-Specialized Open-Weight Mixture-of-Experts Language Model
图 1 · 摘自论文原文
  • 基于稀疏专家混合架构,分三阶段渐进训练并防遗忘。
  • 希伯来语推理平均得分73.8%,激活仅30亿参数却达9倍吞吐量。
  • 适合希伯来语与闪米特语种NLP研究者使用,权重完全开源。

我们提出Hebatron,一个基于NVIDIA Nemotron-3稀疏专家混合架构的希伯来语专用开源大模型。训练采用由易到难的三阶段课程学习,并结合持续防遗忘锚定机制,随后在200万对双语希伯来-英语样本上进行监督微调。仅课程顺序优化就比反向配置提升3个百分点的综合基准表现。Hebatron在希伯来语推理任务中平均得分为73.8%,优于DictaLM-3.0-24B-Thinking(68.9%),在GSM8K-HE和Israeli Trivia上与Gemma-3-27B-IT相当,同时每次前向传播仅激活30亿参数(总参数300亿),在最长65,536标记的原生上下文长度下实现约9倍更高的推理吞吐量。据我们所知,这是首个针对特定语言适配Nemotron-3架构的实例,也是首个具备原生长上下文支持的开源希伯来语专家混合模型。模型权重已公开,以促进希伯来语及闪米特语种自然语言处理研究。

原文摘要 · Abstract (English)

We present Hebatron, a Hebrew-specialized open-weight large language model built on the NVIDIA Nemotron-3 sparse Mixture-of-Experts architecture. Training employs a three-phase easy-to-hard curriculum with continuous anti-forgetting anchoring, followed by supervised fine-tuning on 2 million bilingual Hebrew--English samples. The curriculum ordering alone yields a 3-point aggregate benchmark gain over the reversed configuration. Hebatron achieves a Hebrew reasoning average of 73.8\%, outperforming DictaLM-3.0-24B-Thinking (68.9\%) and remaining competitive with Gemma-3-27B-IT on GSM8K-HE and Israeli Trivia, while activating only 3B parameters per forward pass across a 30B-parameter model, delivering approximately 9 times higher inference throughput at native context lengths up to 65,536 tokens. To our knowledge, this is the first language-specific adaptation of the Nemotron-3 architecture for any target language, and the first open-weight Hebrew-specialized MoE model with native long-context support. Model weights are released openly to support further research in Hebrew and Semitic-language NLP.

希伯来语专家混合长文本开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。