优化土耳其语大模型:新语料筛选与训练方法提升性能
Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training
- 用生成和翻译数据扩充土耳其语语料库
- 少样本与零样本场景下准确率显著提升
- 适合关注低资源语言模型优化的研究者
本研究提出并评估了改进土耳其语大模型性能的新语料选择与训练方法。具体而言,将大语言模型生成的数据及英译土耳其语数据整合进训练流程,显著提升了模型在少样本与零样本学习场景下的准确性。此外,多模型融合进一步增强了性能。人工评估显示,优化后的模型在理解土耳其语和解决逻辑类问题方面表现更优。研究强调,针对土耳其语等低资源语言,优化语料选择策略对提升多语言模型性能至关重要。
原文摘要 · Abstract (English)
In this study, we develop and assess new corpus selection and training methodologies to improve the effectiveness of Turkish language models. Specifically, we adapted Large Language Model generated datasets and translated English datasets into Turkish, integrating these resources into the training process. This approach led to substantial enhancements in model accuracy for both few-shot and zero-shot learning scenarios. Furthermore, the merging of these adapted models was found to markedly improve their performance. Human evaluative metrics, including task-specific performance assessments, further demonstrated that these adapted models possess a greater aptitude for comprehending the Turkish language and addressing logic-based queries. This research underscores the importance of refining corpus selection strategies to optimize the performance of multilingual models, particularly for under-resourced languages like Turkish.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。