32B参数韩语大模型,专攻复杂推理与企业级应用
Mi:dm K 2.5 Pro
- 用代码/数学的高质量数据构建训练基础,结合智能评估机制
- 支持128K上下文,通过多阶段训练提升解题能力与工具使用
- 在韩语评测中表现领先,兼顾安全性和对话流畅性
不断演进的大语言模型需突破单一文本生成,转向多步推理、长上下文理解与代理工作流。这一趋势对现有模型提出挑战,尤其在韩语及特定领域场景下,规模扩展仍显不足。我们推出32B参数旗舰模型Mi:dm K 2.5 Pro,专为解决企业级复杂任务而设计,以推理优化为核心。方法上,构建质量导向的数据体系:利用抽象语法树(AST)分析代码,通过补全合成处理数学问题,并采用大模型评估器筛选数据。预训练阶段采用基于层预测的深度扩展(DuS)与渐进策略,实现128K token上下文窗口。后训练引入多阶段流程:推理监督微调(SFT)、模型合并与异步强化学习(RL),发展复杂问题求解能力。随后通过“融合训练”重新平衡,增强对话流畅性、风格一致性与可靠工具调用。评估显示,该模型性能媲美全球与本地领先模型,在韩语专用基准上达到顶尖水平,展现深层语言与文化理解力。负责任AI评估验证其抗攻击安全性,部署时兼具无害性与响应性。
原文摘要 · Abstract (English)
The evolving LLM landscape requires capabilities beyond simple text generation, prioritizing multi-step reasoning, long-context understanding, and agentic workflows. This shift challenges existing models in enterprise environments, especially in Korean-language and domain-specific scenarios where scaling is insufficient. We introduce Mi:dm K 2.5 Pro, a 32B parameter flagship LLM designed to address enterprise-grade complexity through reasoning-focused optimization. Our methodology builds a robust data foundation via a quality-centric curation pipeline utilizing abstract syntax tree (AST) analysis for code, gap-filling synthesis for mathematics, and an LLM-based quality evaluator. Pre-training scales the model via layer-predictor-based Depth Upscaling (DuS) and a progressive strategy supporting a 128K token context window. Post-training introduces a specialized multi-stage pipeline, including Reasoning SFT, model merging, and asynchronous reinforcement learning (RL), to develop complex problem-solving skills. "Fusion Training" then rebalances these capabilities with conversational fluency, consistent response styling, and reliable tool-use. The evaluations show that Mi:dm K 2.5 Pro achieves competitive performance against leading global and domestic models. In addition, it sets state-of-the-art results on Korean-specific benchmarks, showcasing deep linguistic and cultural understanding. Finally, Responsible AI evaluations validate safety against attacks, ensuring a secure profile for deployment with a balance of harmlessness and responsiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。