基于专家混合架构的高效俄语大模型,助力俄语NLP研究与应用
GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture
- 采用专家混合架构降低计算开销,实现高效俄语语言建模
- 推出多尺寸基座与指令微调模型,支持多种应用场景
- 开源三款模型,推动俄语NLP研究与产业落地
生成式大语言模型(LLMs)在现代自然语言处理研究与应用中日益重要。然而,针对俄语的基础模型发展受限,主要由于所需大量计算资源。本文提出GigaChat家族俄语大模型,涵盖多种规模,包括基座模型和指令微调版本。详细报告了模型架构、预训练过程及实验设计选择。评估其在俄语和英语基准上的表现,并与多语言模型对比。展示顶级模型的系统应用:可通过API、Telegram机器人和网页界面访问。此外,已在HuggingFace上开源三款GigaChat模型(https://huggingface.co/ai-sage),旨在拓展俄语自然语言处理研究机会,支持俄语领域工业解决方案的发展。
原文摘要 · Abstract (English)
Generative large language models (LLMs) have become crucial for modern NLP research and applications across various languages. However, the development of foundational models specifically tailored to the Russian language has been limited, primarily due to the significant computational resources required. This paper introduces the GigaChat family of Russian LLMs, available in various sizes, including base models and instruction-tuned versions. We provide a detailed report on the model architecture, pre-training process, and experiments to guide design choices. In addition, we evaluate their performance on Russian and English benchmarks and compare GigaChat with multilingual analogs. The paper presents a system demonstration of the top-performing models accessible via an API, a Telegram bot, and a Web interface. Furthermore, we have released three open GigaChat models in open-source (https://huggingface.co/ai-sage), aiming to expand NLP research opportunities and support the development of industrial solutions for the Russian language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。