arXiv:2410.16534cs.LG2024-10被引 3

用数据驱动方法让大模型生成精准合成数据,提升小模型性能。

SoftSRV: Learn to Generate Targeted Synthetic Data

  • 通过损失最小化引导冻结大模型生成目标分布数据
  • 在编码、数学、推理任务上显著提升小模型表现
  • 无需领域定制,通用性强且更贴近真实数据分布

我们提出一种新框架 SoftSRV,用于生成针对特定任务的合成微调数据以提升模型性能。给定目标分布的一个样本,该框架采用数据驱动的损失最小化方法,引导一个冻结的大语言模型生成与目标分布相似的合成序列。相较于依赖人工设计提示模板的常规方法,SoftSRV避免了主观性、耗时且需领域适配的问题。我们在三个不同领域(编程、数学、推理)上对方法进行了评估,未对框架进行任何领域特化,验证其通用性。结果表明,SoftSRV生成的数据能显著提升微调后模型的任务表现,并在 MAUVE 相似度指标上更接近真实目标分布。

原文摘要 · Abstract (English)

We present a novel framework, SoftSRV, that is used to generate targeted synthetic fine-tuning data for improving task-specific model performance. Given a sample from a target distribution, our proposed framework uses a data-driven loss minimization approach to steer a frozen large language model (LLM) to generate synthetic sequences that are similar to those from the target distribution. SoftSRV provides a practical improvement over common prompt engineering approaches that rely on human-engineered prompt-templates, which can be idiosyncratic, labor-intensive to craft, and may need to be specialized per domain. We empirically evaluate our method against standard baselines guiding a large LLM to generate synthetic data to fine-tune a smaller language model on three different domains (coding, math, reasoning). We perform these evaluations without any particular specialization of the framework to each domain, emphasizing the generality of our approach. We find that SoftSRV improves upon typical prompt engineering approaches, generating targeted data that leads to fine-tuned models with significantly better task-specific performance. In addition, SoftSRV-generated data better matches the target distribution according to the MAUVE similarity metric.

合成数据大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。