arXiv:2410.14251cs.AIcs.CL2024-10ACL被引 25

用多智能体模拟生成高质量指令数据,仅2万条就超越1000万条训练的模型。

Synthesizing Post-Training Data for LLMs through Multi-Agent Simulation

  • 设计多智能体系统自动构建真实人类需求场景。
  • 用2万条合成数据训练的模型胜过1000万条真实数据训练的基线。
  • 适合想低成本提升LLM指令遵循能力的研究者和工程师。

后训练对使大语言模型遵循人类指令至关重要,但其效果依赖高质量指令数据,而现实中受限于隐私、数据稀缺和标注成本难以获取。受大语言模型模拟人类社会成功的启发,我们提出MATRIX——一个可自动生成多样化文本场景的多智能体模拟器,以真实且可扩展的方式捕捉广泛的人类需求。基于此输出,我们设计了新型场景驱动的指令生成器MATRIX-Gen,实现可控且高度真实的数据合成。大量实验表明,该框架能有效生成通用及领域特定数据。在AlpacaEval 2和Arena-Hard基准上,仅用2万条由MATRIX-Gen合成的指令-响应对进行后训练的Llama-3-8B-Base模型,性能优于使用超过1000万条数据训练的Meta官方Llama-3-8B-Instruct模型。

原文摘要 · Abstract (English)

Post-training is essential for enabling large language models (LLMs) to follow human instructions. However, its effectiveness depends on high-quality instruction data, which is challenging to obtain in the real world due to privacy concerns, data scarcity, and high annotation costs. To fill this gap, inspired by the recent success of using LLMs to simulate human society, we propose MATRIX, a multi-agent simulator that automatically generates diverse text-based scenarios, capturing a wide range of real-world human needs in a realistic and scalable manner. Leveraging these outputs, we introduce a novel scenario-driven instruction generator MATRIX-Gen for controllable and highly realistic data synthesis. Extensive experiments demonstrate that our framework effectively generates both general and domain-specific data. On AlpacaEval 2 and Arena-Hard benchmarks, Llama-3-8B-Base, post-trained on datasets synthesized by MATRIX-Gen with just 20K instruction-response pairs, outperforms Meta's Llama-3-8B-Instruct model, which was trained on over 10M pairs.

指令微调多智能体数据合成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。