用多语言数据增强器提升跨文化指令质量,让大模型更懂不同语言背后的文化。
MIDB: Multilingual Instruction Data Booster for Enhancing Cultural Equality in Multilingual Instruction Synthesis
- 基于16种语言的3.68万条人工修订数据训练,自动修复多语言指令中的内容错误和翻译缺陷。
- 在16种语言中均显著提升指令数据质量,使微调后的模型文化理解能力更强。
- 适合关注多语言公平性、跨文化AI应用的研究者与开发者。
尽管对数据质量存疑,指令合成仍被广泛用于大模型指令微调(IT),作为经济高效的替代方案。近期工作主要提升英文合成指令对齐质量,推动了以英语为中心的大模型发展。然而,多语言合成指令的数据质量问题更为严重,因普遍采用机器翻译(MT)将英文合成数据转译至其他语言,不仅继承原英文数据中的内容错误,还引入了翻译缺陷,并缺乏目标语言的本地化适配,导致训练出的模型存在文化不平等。本文提出MIDB(Multilingual Instruction Data Booster),一个可自动修复多语言合成数据质量的增强框架。MIDB在16种语言上基于约3.68万条由人类语言专家标注的修订样本进行训练,能够有效纠正内容错误、机器翻译缺陷,并改善本地化表达。自动与人工评估均表明,MIDB在16种语言中持续提升了指令数据质量,且使用其增强数据微调的多语言大模型,在指令遵循和文化理解能力上均有显著提升,体现出更优的语言与文化平等性。
原文摘要 · Abstract (English)
Despite doubts on data quality, instruction synthesis has been widely applied into instruction tuning (IT) of LLMs as an economic and rapid alternative. Recent endeavors focus on improving data quality for synthesized instruction pairs in English and have facilitated IT of English-centric LLMs. However, data quality issues in multilingual synthesized instruction pairs are even more severe, since the common synthesizing practice is to translate English synthesized data into other languages using machine translation (MT). Besides the known content errors in these English synthesized data, multilingual synthesized instruction data are further exposed to defects introduced by MT and face insufficient localization of the target languages, leading to cultural inequality in trained LLMs. In this paper, we propose MIDB, a Multilingual Instruction Data Booster to automatically address the quality issues in multilingual synthesized data. MIDB is trained on around 36.8k revision examples across 16 languages by human linguistic experts, thereby can boost the low-quality data by addressing content errors and MT defects, and improving localization in these synthesized data. Both automatic and human evaluation indicate that not only MIDB steadily improved instruction data quality in 16 languages, but also the instruction-following and cultural-understanding abilities of multilingual LLMs fine-tuned on MIDB-boosted data were significantly enhanced, suggesting an improved linguistic and cultural equality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。