用多智能体模拟教学,自动生成高质量指令微调数据
MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching
- 设计多智能体对话生成教师-学生互动数据
- 在多个基准上提升模型推理与泛化能力
- 适合需要高效数据增强的LLM研究者
指令微调对自然语言处理任务至关重要,可提升预训练模型的指令遵循能力和特定任务表现。然而,由于数据收集困难和生产成本高,获取高质量的大模型微调数据仍具挑战。为此,我们提出MASTER,一种通过具有不同认知水平的多智能体交互来丰富原始数据的新颖数据增强方法。我们模拟了三种基于教育学原理的教学场景,利用多智能体对话生成高质量的师生交互数据。基于MASTER,我们构建了BOOST-QA,一个从Orca-Math-200k、ProcQA和OpenHermes2.5等现有数据集扩展而来的微调数据集。实验表明,使用BOOST-QA微调的模型在多个基准测试中表现优异,展现出强大的多任务泛化能力。尤其值得注意的是,MASTER显著提升了模型在复杂任务中的推理能力,为未来研究提供了重要启示。
原文摘要 · Abstract (English)
Instruction fine-tuning is crucial in NLP tasks, enhancing pretrained models' instruction-following capabilities and task-specific performance. However, obtaining high-quality fine-tuning data for large models is challenging due to data collection difficulties and high production costs. To address this, we propose MASTER, a novel data augmentation method that enriches original data through interactions among multiple agents with varying cognitive levels. We simulate three pedagogically grounded teaching scenarios, leveraging multi-agent conversations to generate high-quality teacher-student interaction data. Utilizing MASTER, we construct BOOST-QA, a fine-tuning dataset augmented from existing datasets like Orca-Math-200k, ProcQA, and OpenHermes2.5. Experiments show that models fine-tuned with BOOST-QA perform excellently across multiple benchmarks, demonstrating strong multitask generalization. Notably, MASTER significantly improves models' reasoning abilities in complex tasks, providing valuable insights for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。