为挪威语打造的开源大模型,提升北欧语言AI能力
NorwAI's Large Language Models: Technical Report
- 基于Transformer架构,用25B-88.45B tokens训练,专优化挪威语
- 指令微调模型在对话任务中表现优秀,适合实际应用
- 面向北欧机构开放,支持研究与实验,可自由使用
挪威语(约500万使用者)在自然语言处理领域仍被严重忽视。为弥补这一差距,诺瓦AI的NorLLM团队开发了一套专为挪威语及北欧语言设计的大模型系列,基于GPT、Mistral、Llama2、Mixtral和Magistral等多种Transformer架构。这些模型通过250亿至884.5亿个标记从头预训练或持续预训练,采用扩展的挪威语分词器与先进后训练策略,以优化性能、增强鲁棒性并提升跨任务适应能力。其中指令微调版本(如Mistral-7B-Instruct和Mixtral-8x7B-Instruct)展现出强大的助手式交互能力,具备在实际场景中部署的潜力。所有模型均对北欧组织、企业及学生开放,可用于研究与实验。本报告详细记录了模型架构、训练数据、分词器设计、微调策略、部署方案与评估结果。
原文摘要 · Abstract (English)
Norwegian, spoken by approximately five million people, remains underrepresented in many of the most significant breakthroughs in Natural Language Processing (NLP). To address this gap, the NorLLM team at NorwAI has developed a family of models specifically tailored to Norwegian and other Scandinavian languages, building on diverse Transformer-based architectures such as GPT, Mistral, Llama2, Mixtral and Magistral. These models are either pretrained from scratch or continually pretrained on 25B - 88.45B tokens, using a Norwegian-extended tokenizer and advanced post-training strategies to optimize performance, enhance robustness, and improve adaptability across various real-world tasks. Notably, instruction-tuned variants (e.g., Mistral-7B-Instruct and Mixtral-8x7B-Instruct) showcase strong assistant-style capabilities, underscoring their potential for practical deployment in interactive and domain-specific applications. The NorwAI large language models are openly available to Nordic organizations, companies and students for both research and experimental use. This report provides detailed documentation of the model architectures, training data, tokenizer design, fine-tuning strategies, deployment, and evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。