arXiv:2503.22973cs.CLcs.AI2025-03EMNLP

用合成数据提升多语言开放式生成能力,效果显著。

XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation

  • 用XL-Instruct生成高质量跨语言指令数据
  • 8千条数据微调后对齐GPT-4o-Mini胜率升至21.5%
  • 适合做多语言大模型后训练与评估的研究者

跨语言开放式生成——以不同于查询语种回应——是重要但研究不足的问题。本文提出XL-Instruct,一种生成高质量合成数据的新方法,并引入XL-AlpacaEval基准,用于评估大语言模型的跨语言生成能力。实验表明,仅用8000条由XL-Instruct生成的指令进行微调,即可显著提升模型性能,使对GPT-4o-Mini的胜率从7.4%提升至21.5%,并在多个细粒度质量指标上改善。此外,基于LLM在XL-Instruct上微调后,在机器翻译的m-AlpacaEval上展现出强大的零样本问答能力。这些一致的提升凸显了XL-Instruct在多语言大模型后训练中的潜力。最后,我们公开发布XL-Suite,包含训练与评估数据,以推动跨语言开放式生成研究。

原文摘要 · Abstract (English)

Cross-lingual open-ended generation - responding in a language different from that of the query - is an important yet understudied problem. This work proposes XL-Instruct, a novel technique for generating high-quality synthetic data, and introduces XL-AlpacaEval, a new benchmark for evaluating cross-lingual generation capabilities of large language models (LLMs). Our experiments show that fine-tuning with just 8K instructions generated using XL-Instruct significantly improves model performance, increasing the win rate against GPT-4o-Mini from 7.4% to 21.5% and improving on several fine-grained quality metrics. Moreover, base LLMs fine-tuned on XL-Instruct exhibit strong zero-shot improvements to question answering in the same language, as shown on our machine-translated m-AlpacaEval. These consistent gains highlight the promising role of XL-Instruct in the post-training of multilingual LLMs. Finally, we publicly release XL-Suite, a collection of training and evaluation data to facilitate research in cross-lingual open-ended generation.

多语言生成合成数据LLM微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。