Kanana用高效训练法打造低成本双语模型,韩语表现突出。
Kanana: Compute-efficient Bilingual Language Models
- 通过分阶段预训练与剪枝压缩降低算力消耗
- 2.1B~32.5B参数模型在韩语上超越主流,英语具竞争力
- 适合关注韩语AI、资源受限部署的研究者
我们提出Kanana系列双语语言模型,在韩语上表现优异,英语表现具备竞争力,计算成本显著低于同规模先进模型。报告详述了预训练中采用的高效技术,包括高质量数据过滤、分阶段预训练、深度扩展以及剪枝与蒸馏。后训练阶段则结合监督微调与偏好优化,提升模型与用户交互能力。此外,还介绍了针对特定场景的适配方法,如嵌入、检索增强生成和函数调用。该系列模型参数量覆盖2.1B至32.5B,其中2.1B基础版、指令版和嵌入版已公开,旨在推动韩语语言模型研究。
原文摘要 · Abstract (English)
We introduce Kanana, a series of bilingual language models that demonstrate exceeding performance in Korean and competitive performance in English. The computational cost of Kanana is significantly lower than that of state-of-the-art models of similar size. The report details the techniques employed during pre-training to achieve compute-efficient yet competitive models, including high quality data filtering, staged pre-training, depth up-scaling, and pruning and distillation. Furthermore, the report outlines the methodologies utilized during the post-training of the Kanana models, encompassing supervised fine-tuning and preference optimization, aimed at enhancing their capability for seamless interaction with users. Lastly, the report elaborates on plausible approaches used for language model adaptation to specific scenarios, such as embedding, retrieval augmented generation, and function calling. The Kanana model series spans from 2.1B to 32.5B parameters with 2.1B models (base, instruct, embedding) publicly released to promote research on Korean language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。