arXiv:2510.06847cs.CLcs.AI2025-10

开源泰国语大模型OpenJAI-v1.0,支持中泰双语任务优化。

OpenJAI-v1.0: An Open Thai Large Language Model

  • 基于Qwen3-14B微调,聚焦指令遵循、长文本理解与工具使用场景。
  • 在多项基准测试中优于现有开源泰语模型,未出现灾难性遗忘。
  • 适合泰国NLP研究者及开发者使用,推动本地AI生态发展。

我们提出OpenJAI-v1.0,一个基于Qwen3-14B的开源多语言大模型,专为泰语和英语设计。通过在指令遵循、长上下文理解和工具使用三个关键应用场景上精心筛选数据,显著提升实际任务表现。评估结果表明,OpenJAI-v1.0在多个基准测试中超越其基础模型,并优于其他主流开源泰语模型,同时避免了灾难性遗忘问题。该模型已公开发布,为泰国人工智能社区提供新的自然语言处理资源。

原文摘要 · Abstract (English)

We introduce OpenJAI-v1.0, an open-source large language model for Thai and English, developed from the Qwen3-14B model. Our work focuses on boosting performance on practical tasks through carefully curated data across three key use cases: instruction following, long-context understanding, and tool use. Evaluation results show that OpenJAI-v1.0 improves on the capabilities of its base model and outperforms other leading open-source Thai models on a diverse suite of benchmarks, while avoiding catastrophic forgetting. OpenJAI-v1.0 is publicly released as another alternative NLP resource for the Thai AI community.

大模型泰语开源NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。