K-EXAONE 2.0是750B参数的多语言MoE模型,支持超长上下文与更强安全能力。
K-EXAONE 2.0 Technical Report

- 基于旧模型升级为750B参数的MoE架构,每令牌激活约37B参数
- 支持256K上下文长度,在长文本理解与智能编码任务中表现突出
- 聚焦韩语文化安全,适合需要多语言与长文本处理的研究与应用
本技术报告介绍由LG AI Research开发的开源权重多语言基础模型K-EXAONE 2.0,是迈向全球前沿规模模型的重要一步。该模型未从零训练,而是通过复用并扩展K-EXAONE架构,构建出总参数量达750B、每令牌激活约37B参数的Mixture-of-Experts(MoE)模型,容量超过前代三倍以上。其支持最长256K tokens的上下文长度,并将多语言覆盖从六种扩展至十种。训练流程融合持续预训练、难度导向的中期训练及后训练,强化推理、代理式编程、多语言能力与基于韩国社会文化背景的安全性。在九个反映实际应用场景的评估类别中,K-EXAONE 2.0优于K-EXAONE,且与开源模型相比保持竞争力,尤其在代理编程和长上下文理解方面提升显著,长期上下文检索与安全性表现尤为突出。模型以Apache 2.0许可证发布,旨在推动整个AI生态的评估、部署、适配与创新,标志着全球前沿探索的起点而非终点。
原文摘要 · Abstract (English)
This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward global frontier-scale foundation models. Rather than training from scratch, we upcycle K-EXAONE and expand its architecture, yielding a Mixture-of-Experts (MoE) model with 750B total parameters and approximately 37B activated per token---more than three times the capacity of its predecessor. K-EXAONE 2.0 supports context lengths of up to 256K tokens and expands multilingual coverage from six to ten languages. Its training pipeline combines continual pre-training, difficulty-focused mid-training, and post-training to strengthen reasoning, agentic coding, multilingual capability, and safety grounded in Korean sociocultural contexts. Across nine evaluation categories selected to reflect the conditions of practical use, K-EXAONE 2.0 improves over K-EXAONE and remains competitive with open-weight models, showing its largest gains in agentic coding and long-context understanding and its clearest strengths in long-context retrieval and safety. Released under the Apache 2.0 license, K-EXAONE 2.0 enables the wider AI ecosystem to evaluate, deploy, adapt, and build upon it, while marking the beginning---rather than the endpoint---of our challenge toward the global frontier.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。