arXiv:2502.09056cs.CLcs.AI2025-02被引 6

用120美元预算,一天内让泰语大模型具备顶级推理能力

Adapting Language-Specific LLMs to a Reasoning Model in One Day via Model Merging -- An Open Recipe

论文配图:Adapting Language-Specific LLMs to a Reasoning Model in One Day via Model Merging -- An Open Recipe
图 1 · 摘自论文原文
  • 通过模型融合与数据筛选,将深求R1推理能力迁移到泰语模型
  • 仅用120美元计算成本,推理性能媲美DeepSeek R1
  • 适合关注低资源语言模型优化的研究者与开发者

本文研究了数据选择与模型融合方法,旨在将DeepSeek R1等先进推理能力引入语言专用大语言模型(LLM),特别聚焦泰语模型。尽管DeepSeek R1在推理方面表现优异,但主要惠及英语和中文等高资源语言。由于训练数据和模型优化以英语为主,低资源语言的性能仍受限,导致代码切换不可靠、任务表现下降。为此,本地化大模型项目致力于提升本地语言表达的准确性。本研究证明,在仅使用公开数据集且计算成本控制在120美元的前提下,可使语言专用模型的推理能力达到DeepSeek R1水平,同时不损害其在目标语言任务上的表现。

原文摘要 · Abstract (English)

This paper investigates data selection and model merging methodologies aimed at incorporating advanced reasoning capabilities such as those of DeepSeek R1 into language-specific large language models (LLMs), with a particular focus on the Thai LLM. Our goal is to enhance the reasoning capabilities of language-specific LLMs while maintaining their target language abilities. DeepSeek R1 excels in reasoning but primarily benefits high-resource languages such as English and Chinese. However, low-resource languages remain underserved due to the dominance of English-centric training data and model optimizations, which limit performance in these languages. This limitation results in unreliable code-switching and diminished effectiveness on tasks in low-resource languages. Meanwhile, local and regional LLM initiatives have attempted to bridge this gap by developing language-specific LLMs that focus on improving local linguistic fidelity. We demonstrate that, with only publicly available datasets and a computational budget of $120, it is possible to enhance the reasoning capabilities of language-specific LLMs to match the level of DeepSeek R1, without compromising their performance on target language tasks.

模型融合低资源语言推理能力成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。