arXiv:2510.08620cs.CL2025-10

JAI-1用扩容+系统注入法打造泰国专用大模型,兼顾通用能力与泰语表现。

JAI-1: A Thai-Centric Large Language Model

  • 从英文小模型扩容出发,分阶段注入泰语知识避免原有能力丢失
  • 预训练用1.5万亿词元(含3000亿泰语),后经60万条指令微调
  • 在多个泰语评测中优于Typhoon2-70B,适合泰语应用开发者

本文介绍JAI-1,一个参数量达750亿的泰国中心语言模型。现有泰语模型多基于开源模型进行增量训练,未改架构,易因泰语信息注入而破坏原有知识。JAI-1采用扩容策略:从高性能英文开源LLM出发,扩大参数规模,并利用新增容量系统性融入泰语知识。该方法既保留原模型通用智能,又形成独特架构,便于后续扩展。预训练阶段,模型接触1.5万亿词元数据,其中包含超3000亿泰语词元;随后通过超过60万条指令样本进行监督微调与对齐调优。最终模型在IFEval-TH、MT-Bench-TH和JAI-Hall-Bench等泰语基准测试中表现优于Typhoon2-70B,验证了其扩容与知识融合框架的有效性。

原文摘要 · Abstract (English)

This technical report introduces JAI-1, a Thai-centric language model with 75B parameters. Recent Thai models have primarily relied on existing open-source models, applying additional training without structural modifications to specialize in Thai. However, this approach risks eroding pre-existing knowledge in the model's parameter space during the injection of Thai-specific information, as optimized parameters for general tasks may conflict with new linguistic requirements. In contrast, JAI-1 adopts an upscaling strategy: starting from a smaller, high-performing English open-source LLM, we expanded its parameter space and utilized the newly allocated capacity to systematically integrate Thai-language knowledge. This methodology not only preserves the original model's general intelligence but also establishes a unique architecture distinct from other open-source models, enabling scalable future enhancements. During pre-training, JAI-1 was exposed to 1.5T tokens, including over 300B Thai language tokens. This was followed by post-training stages -- supervised fine-tuning and alignment tuning -- using more than 600K instruction-based examples. The final model demonstrated superior performance compared to Typhoon2-70B on Thai-centric benchmarks (IFEval-TH, MT-Bench-TH, and JAI-Hall-Bench), validating the efficacy of its upscaling and knowledge-integration framework.

大模型泰语语言模型知识注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。