arXiv:2601.09200cs.CLcs.AI2026-01

5190亿参数大模型,支持可控推理,韩语表现突出

A.X K1 Technical Report

  • 基于缩放定律设计,固定算力下优化参数与词表规模
  • 预训练10万亿token数据,支持思维模式自由切换
  • 韩语评测领先,适合需要高效推理的部署场景

我们提出A.X K1,一个从头训练的5190亿参数混合专家(MoE)语言模型。设计上利用缩放定律,在固定计算预算下优化训练配置与词汇量。模型在约10万亿标记的语料上预训练,该语料通过多阶段数据处理流程构建。为弥合理论能力与推理效率之间的差距,A.X K1支持显式可控制的推理机制,便于在多样现实场景中规模化部署。我们提出一种简单有效的Think-Fusion训练方法,使单一统一模型内可实现思维与非思维模式的用户可控切换。大量评估表明,A.X K1性能媲美顶尖开源模型,同时在韩语基准测试中展现出显著优势。

原文摘要 · Abstract (English)

We introduce A.X K1, a 519B-parameter Mixture-of-Experts (MoE) language model trained from scratch. Our design leverages scaling laws to optimize training configurations and vocabulary size under fixed computational budgets. A.X K1 is pre-trained on a corpus of approximately 10T tokens, curated by a multi-stage data processing pipeline. Designed to bridge the gap between reasoning capability and inference efficiency, A.X K1 supports explicitly controllable reasoning to facilitate scalable deployment across diverse real-world scenarios. We propose a simple yet effective Think-Fusion training recipe, enabling user-controlled switching between thinking and non-thinking modes within a single unified model. Extensive evaluations demonstrate that A.X K1 achieves performance competitive with leading open-source models, while establishing a distinctive advantage in Korean-language benchmarks.

大模型MoE韩语推理控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。