arXiv:2601.18129cs.CLcs.AI2026-01被引 1

用小规模训练让本地大模型兼具通用能力与本土任务表现

Typhoon-S: Minimal Open Post-Training for Sovereign Large Language Models

  • 仅用监督微调+策略蒸馏+小规模强化微调,降低训练成本
  • 在泰语法律推理和本土知识上显著提升,通用能力不下降
  • 适合资源有限的国家或机构自主打造可控大模型

大语言模型虽发展迅速,但多数先进模型集中于英语、中文等高资源语言,且由少数机构主导,形成技术壁垒。在主权场景下,区域或国家级机构需在资源有限、透明度要求高的条件下,掌控模型权重、数据与部署。为此,我们提出两个核心需求:可采纳性(将基础模型转化为通用助手)与主权能力(完成本地高风险任务,如泰语法律推理与文化知识)。本文探索是否可在不依赖海量指令数据、复杂偏好对齐或大规模强化微调的前提下实现。提出 Typhoon-S:一种最小化、开放的后训练方案,结合监督微调、在线策略蒸馏与小规模强化微调。以泰语为例,验证该方法能有效将主权适配与通用基础模型转化为高性能指令模型;进一步表明,采用 InK-GRPO(在 GRPO 损失中加入下一步词预测损失)的小规模强化微调,显著提升泰语法律推理与本土知识表现,同时保持通用能力。结果表明,精心设计的后训练策略可大幅降低指令数据与算力需求,为学术级资源下构建高质量主权大模型提供可行路径。

原文摘要 · Abstract (English)

Large language models (LLMs) have progressed rapidly; however, most state-of-the-art models are trained and evaluated primarily in high-resource languages such as English and Chinese, and are often developed by a small number of organizations with access to large-scale compute and data. This gatekeeping creates a practical barrier for sovereign settings in which a regional- or national-scale institution or domain owner must retain control and understanding of model weights, training data, and deployment while operating under limited resources and strict transparency constraints. To this end, we identify two core requirements: (1) adoptability, the ability to transform a base model into a general-purpose assistant, and (2) sovereign capability, the ability to perform high-stakes, region-specific tasks (e.g., legal reasoning in local languages and cultural knowledge). We investigate whether these requirements can be achieved without scaling massive instruction corpora or relying on complex preference tuning pipelines and large-scale reinforcement fine-tuning (RFT). We present Typhoon S, a minimal and open post-training recipe that combines supervised fine-tuning, on-policy distillation, and small-scale RFT. Using Thai as a representative case study, we demonstrate that our approach transforms both sovereign-adapted and general-purpose base models into instruction-tuned models with strong general performance. We further show that small-scale RFT with InK-GRPO -- an extension of GRPO that augments the GRPO loss with a next-word prediction loss -- improves Thai legal reasoning and Thai-specific knowledge while preserving general capabilities. Our results suggest that a carefully designed post-training strategy can reduce the required scale of instruction data and computation, providing a practical path toward high-quality sovereign LLMs under academic-scale resources.

主权模型小规模训练泰语强化微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。