用极少资源让英文大模型学会韩语,效果还更好。
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
- 用小规模韩语数据微调现有英文模型,流程完整可复现。
- 新模型在韩语任务上超越现有顶尖模型,仅需少量算力。
- 开源全流程代码,适合资源有限的团队做多语言适配。
由于顶级大模型在英语和中文之外的语言中表现不佳,提升其在新语言中的能力已成为关键任务。同时,大模型的端到端训练过程因商业机密、技术复杂性和伦理问题,对公众而言仍属未知。本文提出一种低成本方法,将现有英文大模型适配至韩语。完整描述了数据收集、预处理、模型训练、下游评测基准构建与评估的全流程。实验表明,该方法能高效、经济地为大模型添加新语言能力。新推出的双语模型 Thunder-LLM 与 Thunder-LLM-Ins 在韩语任务上表现优于现有最佳模型,且仅使用极少量数据与计算资源。我们公开了全部实践经验与代码。
原文摘要 · Abstract (English)
Since state-of-the-art LLMs often underperform in languages other than English or Chinese, improving the capability of LLMs in new languages has become an essential task. Moreover, LLMs' entire end-to-end training process remains largely unknown to the public due to proprietary reasons, technical complexity, inconsistent documentation, and ethical considerations. The complete picture remains a closely guarded secret within the industry. This paper presents methods to adapt an existing English-based LLM to Korean in a low-budget scenario. We describe the entire end-to-end process: collecting Korean datasets, preprocessing the data, training the model, creating downstream benchmarks, and conducting evaluations. The evaluation results indicate that our method can effectively and cost-efficiently add new language capabilities to existing LLMs. Our new bilingual models, Thunder-LLM and Thunder-LLM-Ins, achieve superior Korean performance compared to state-of-the-art models while utilizing minimal data and computational resources. We share our comprehensive experience and make the code publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。