80亿参数的哈萨克语大模型,让本地语言也能用上顶尖AI。
Sherkala-Chat: Building a State-of-the-Art LLM for Kazakh in a Moderately Resourced Setting
- 基于LLaMA-3.1微调,用453亿词训练多语言数据
- 在哈萨克语上性能超越同类模型,英语表现也达标
- 开源模型+本土化指令数据,适合研究与实用
Sherkala-Chat (8B) 是一个专为哈萨克语设计的先进指令微调开源生成式大语言模型。该模型基于 LLaMA-3.1-8B 构建,使用包含哈萨克语、英语、俄语和土耳其语的 453 亿个标记进行训练。拥有 80 亿参数,其在哈萨克语中展现出强大的知识与推理能力,显著优于现有同规模的开放哈萨克语及多语言模型,并在英语任务上达到竞争性水平。为确保有效且负责任的对齐,研究者采用了翻译的指令数据集、自动构建并人工验证的哈萨克斯坦本地指令数据集以及针对哈萨克语的安全数据。模型已作为开源权重发布,并附有详细的训练、对齐与评估说明,以支持哈萨克语研究与实际应用。
原文摘要 · Abstract (English)
Llama-3.1-Sherkala-8B-Chat, or Sherkala-Chat (8B) for short, is a state-of-the-art instruction-tuned open generative large language model (LLM) designed for Kazakh. Sherkala-Chat (8B) aims to enhance the inclusivity of LLM advancements for Kazakh speakers. Adapted from the LLaMA-3.1-8B model, Sherkala-Chat (8B) is trained on 45.3B tokens across Kazakh, English, Russian, and Turkish. With 8 billion parameters, it demonstrates strong knowledge and reasoning abilities in Kazakh, significantly outper-forming existing open Kazakh and multilingual models of similar scale while achieving competitive performance in English. To ensure effective and responsible alignment, we leverage translated instruction datasets, a Kazakhstan-specific instruction dataset that is automatically constructed and manually verified, and Kazakh-specific safety data. We release Sherkala-Chat (8B) as an open-weight model, along with a detailed description of its training, alignment, and evaluation, to support research and real-world applications for Kazakh speakers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。