arXiv:2503.23714cs.CL2025-03被引 4

用开源大模型生成回复,构建高质量人类指令数据集

Building Instruction-Tuning Datasets from Human-Written Instructions with Open-Weight Large Language Models

  • 以人类编写的指令为输入,用开源大模型生成对应回复
  • 在多个语言上训练的模型均达到当前最优性能
  • 数据集开源可用,适合研究与工业应用

指令微调对使大语言模型解决实际任务至关重要。已有研究证明仅用大模型生成的数据也有效,但本文回答:仍需人类来源信号。我们通过将人类编写的指令与大模型生成的回复配对,构建了当前最先进的指令微调数据集。基于这些数据微调的模型在各项指标上均优于现有方法。该方法可轻松扩展至其他语言,我们在日语上验证并获得最佳表现。分析显示,新语言微调使模型能理解指令,但缺乏该语言的文化知识。所有数据集与微调模型将公开发布,采用宽松许可,支持多样化应用。

原文摘要 · Abstract (English)

Instruction tuning is crucial for enabling Large Language Models (LLMs) to solve real-world tasks. Prior work has shown the effectiveness of instruction-tuning data synthesized solely from LLMs, raising a fundamental question: Do we still need human-originated signals for instruction tuning? This work answers the question affirmatively: we build state-of-the-art instruction-tuning datasets sourced from human-written instructions, by simply pairing them with LLM-generated responses. LLMs fine-tuned on our datasets consistently outperform those fine-tuned on existing ones. Our data construction approach can be easily adapted to other languages; we build datasets for Japanese and confirm that LLMs tuned with our data reach state-of-the-art performance. Analyses suggest that instruction-tuning in a new language allows LLMs to follow instructions, while the tuned models exhibit a notable lack of culture-specific knowledge in that language. The datasets and fine-tuned models will be publicly available. Our datasets, synthesized with open-weight LLMs, are openly distributed under permissive licenses, allowing for diverse use cases.

指令微调开源模型多语言数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。