arXiv:2504.12563cs.CLcs.AI2025-04ACL被引 9

用多智能体协作生成高多样性合成数据,仅2500万词元就实现专业领域适配。

MetaSynth: Meta-Prompting-Driven Agentic Scaffolds for Diverse Synthetic Data Generation

  • 通过元提示调度多个专家模型协同生成数据,提升多样性。
  • 仅用2500万词元合成数据,金融和生物医学领域性能分别提升4.08%和13.75%。
  • 适合需要高效领域适配但缺乏真实数据的研究者使用。

当前如Phi-3.5和Phi-4等小型语言模型依赖大模型生成的合成数据。然而合成数据多样性不足,限制了其在模型适配等场景的应用。为此,我们提出MetaSynth,通过元提示驱动多个‘专家’LLM智能体协同生成数据,以增强多样性。仅使用2500万词元的合成数据,我们就成功将Mistral-7B-v0.3适配至金融与生物医学两个专业领域,且未损害其通用能力。在七项自动化指标下,生成数据的多样性接近预训练语料库水平。持续预训练显示,相较于基线模型,金融领域提升4.08%,生物医学领域提升13.75%。而使用模板提示生成的数据即使包含真实样本,性能仍下降。结果表明,仅需数百万词元的高质量合成数据,无需混合真实数据,即可通过MetaSynth实现有效领域适配。

原文摘要 · Abstract (English)

Recent smaller language models such Phi-3.5 and Phi-4 rely on synthetic data generated using larger Language models. Questions remain about leveraging synthetic data for other use cases, such as adapting LLMs to specific domains. A key limitation of synthetic data is low diversity, which negatively impacts its downstream applicability for improving other models. To address this, we propose MetaSynth, a method for generating synthetic data that enhances diversity through meta-prompting, where a language model orchestrates multiple "expert" LLM agents to collaboratively generate data. Using only 25 million tokens of synthetic data generated with MetaSynth, we successfully adapt a well-trained LLM (Mistral-7B-v0.3) to two specialized domains-Finance and Biomedicine-without compromising the capabilities of the resulting model in general tasks. In addition, we evaluate the diversity of our synthetic data using seven automated metrics, and find that it approaches the diversity of LLM pre-training corpora. Continually pre-training Mistral-7B-v0.3 with MetaSynth notably outperforms the base LLM, showing improvements of up to 4.08% in Finance and 13.75% in Biomedicine. The same model shows degraded performance when trained on data generated using a template prompt, even when the template includes prior generations and varying In-Context exemplars of real data. Our findings suggest that a few million tokens of diverse synthetic data without mixing any real data, is sufficient for effective domain adaptation when using MetaSynth.

合成数据领域适配多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。