用本土化数据微调,让双语大模型在印地语和英语上都更准且不增大模型。
Improving Multilingual Capabilities with Cultural and Local Knowledge in Large Language Models While Enhancing Native Performance
- 用48.5万条中英指令数据微调,提升模型跨语言能力。
- 140次训练验证:双语性能提升3%,且不降低母语表现。
- 无需扩容词表或改架构,适合低资源语言研究者使用。
大型语言模型虽表现优异,但主要针对英语等高资源语言开发,许多语言仍被忽视。我们提出名为Mantra-14B的印地语-英语双语大模型,其在两种语言上的基准测试平均得分提升约3%,优于规模翻倍的模型。基于包含48.5万样本的中英指令数据集,我们对Qwen-2.5-14B-Instruct、Phi-4等7种不同参数量的模型进行微调,覆盖超过140次训练实验,验证了在不牺牲母语性能的前提下显著提升多语言能力的可行性。该方法避免了词汇表扩展或架构修改等高成本操作,保持模型体积小。结果表明,仅通过少量融入文化与本地知识的数据微调,即可有效缩小语言性能差距,且计算开销极低。我们已将训练代码、数据集及模型以MIT和Apache许可证开源,支持低资源语言的研究推进。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown remarkable capabilities, but their development has primarily focused on English and other high-resource languages, leaving many languages underserved. We present our latest Hindi-English bi-lingual LLM \textbf{Mantra-14B} with ~3\% average improvement in benchmark scores over both languages, outperforming models twice its size. Using a curated dataset composed of English and Hindi instruction data of 485K samples, we instruction tuned models such as Qwen-2.5-14B-Instruct and Phi-4 to improve performance over both English and Hindi. Our experiments encompassing seven different LLMs of varying parameter sizes and over 140 training attempts with varying English-Hindi training data ratios demonstrated that it is possible to significantly improve multilingual performance without compromising native performance. Further, our approach avoids resource-intensive techniques like vocabulary expansion or architectural modifications, thus keeping the model size small. Our results indicate that modest fine-tuning with culturally and locally informed data can bridge performance gaps without incurring significant computational overhead. We release our training code, datasets, and models under mit and apache licenses to aid further research towards under-represented and low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。