开源3种规模的希伯来语大模型,支持长文本与工具调用。
Dicta-LM 3.0: Advancing The Frontier of Hebrew Sovereign LLMs
- 基于Mistral、NVIDIA和Qwen基模型,微调适配希伯来语与英语。
- 推出24B/12B/1.7B三版本,上下文长度达65k tokens。
- 专为低资源语言设计,适合多语言NLP研究与应用开发。
前沿实验室已发布开放权重的大语言模型,但非英语主权大模型仍供不应求。训练低资源语言如希伯来语的大模型面临独特挑战。本文介绍Dicta-LM 3.0:一个基于大规模希伯来语与英语语料库训练的开放权重大模型系列。模型分为三种尺寸:24B(基于Mistral-Small-3.1)、12B(基于NVIDIA Nemotron Nano V2)和1.7B(基于Qwen3-1.7B)。每个模型提供多个变体,原生上下文长度为65k token,包含基础模型与支持工具调用的对话模型。为严格评估,我们引入新的希伯来语对话模型基准测试套件,涵盖翻译、摘要、Winograd、以色列趣味知识题和音符标注(nikud)等任务。本工作不仅解决低资源语言建模难题,还提出可复用于其他非英语语言的适配框架,推动多语言自然语言处理发展。
原文摘要 · Abstract (English)
Open-weight LLMs have been released by frontier labs; however, sovereign Large Language Models (for languages other than English) remain low in supply yet high in demand. Training large language models (LLMs) for low-resource languages such as Hebrew poses unique challenges. In this paper, we introduce Dicta-LM 3.0: an open-weight collection of LLMs trained on substantially-sized corpora of Hebrew and English texts. The model is released in three sizes: 24B - adapted from the Mistral-Small-3.1 base model, 12B - adapted from the NVIDIA Nemotron Nano V2 model, and 1.7B - adapted from the Qwen3-1.7B base model. We are releasing multiple variants of each model, each with a native context length of 65k tokens; base model and chat model with tool-calling support. To rigorously evaluate our models, we introduce a new benchmark suite for evaluation of Hebrew chat-LLMs, covering a diverse set of tasks including Translation, Summarization, Winograd, Israeli Trivia, and Diacritization (nikud). Our work not only addresses the intricacies of training LLMs in low-resource languages but also proposes a framework that can be leveraged for adapting other LLMs to various non-English languages, contributing to the broader field of multilingual NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。