用基因算法模拟大模型生成更真实多样的文本数据
Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
- 把文本属性当基因,用大模型模拟交叉突变
- 合成数据质量与多样性显著提升,逼近真实分布
- 特别适合数据不平衡场景的模型训练
大型语言模型在生成合成数据方面表现优异,但保证其质量和多样性仍具挑战。本文提出 Genetic Prompt 框架,将语义文本属性视为基因序列,并利用大模型模拟交配和突变操作。该遗传过程通过生成新颖的属性组合,提升数据质量与多样性,使合成数据分布更接近真实数据。为优化父代选择,还引入主动学习机制以扩展后代搜索空间。多任务 NLP 实验表明,Genetic Prompt 显著优于当前最优基线,在不同规模生成模型下均表现稳健。此外,将合成数据与原始训练集融合可显著提升下游模型性能,尤其在类别不平衡场景中效果突出。结果验证了 Genetic Prompt 在多种 NLP 应用中生成高质量合成数据的有效性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel at generating synthetic data, but ensuring its quality and diversity remains challenging. We propose Genetic Prompt, a novel framework that combines genetic algorithms with LLMs to augment synthetic data generation. Our approach treats semantic text attributes as gene sequences and leverages the LLM to simulate crossover and mutation operations. This genetic process enhances data quality and diversity by creating novel attribute combinations, yielding synthetic distributions closer to real-world data. To optimize parent selection, we also integrate an active learning scheme that expands the offspring search space. Our experiments on multiple NLP tasks reveal several key findings: Genetic Prompt not only significantly outperforms state-of-the-art baselines but also shows robust performance across various generator model sizes and scales. Moreover, we demonstrate that fusing our synthetic data with the original training set significantly boosts downstream model performance, particularly for class-imbalanced scenarios. Our findings validate that Genetic Prompt is an effective method for producing high-quality synthetic data for a wide range of NLP applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。