用指令感知翻译框架提升非英语指令数据质量
InstaTrans: An Instruction-Aware Translation Framework for Non-English Instruction Datasets
- 设计指令感知翻译框架InstaTrans,确保翻译忠实原意
- 在多个语言上验证,翻译后微调模型性能显著提升
- 适合想低成本拓展大模型多语言能力的研究者
由于长尾现象,生成高质量的非英语指令数据集极具挑战性。为此,我们提出通过翻译高质量英语指令数据集来解决该问题,强调翻译需完整且保持指令语义一致性。我们提出专为指令数据集设计的翻译框架InstaTrans(INSTruction-Aware TRANSlation)。实验表明,InstaTrans在翻译完整性与指令感知性方面优于现有方法,具备低成本扩大大模型多语言覆盖的潜力。进一步验证显示,使用InstaTrans翻译的数据微调大模型,可有效提升其在目标语言上的表现。
原文摘要 · Abstract (English)
It is challenging to generate high-quality instruction datasets for non-English languages due to tail phenomena, which limit performance on less frequently observed data. To mitigate this issue, we propose translating existing high-quality English instruction datasets as a solution, emphasizing the need for complete and instruction-aware translations to maintain the inherent attributes of these datasets. We claim that fine-tuning LLMs with datasets translated in this way can improve their performance in the target language. To this end, we introduces a new translation framework tailored for instruction datasets, named InstaTrans (INSTruction-Aware TRANSlation). Through extensive experiments, we demonstrate the superiority of InstaTrans over other competitors in terms of completeness and instruction-awareness of translation, highlighting its potential to broaden the accessibility of LLMs across diverse languages at a relatively low cost. Furthermore, we have validated that fine-tuning LLMs with datasets translated by InstaTrans can effectively improve their performance in the target language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。