从质量、多样性和复杂性三方面评估大模型生成的合成数据,揭示其对下游模型性能的影响。
Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models
- 以质量、多样性和复杂性为维度评估合成数据生成算法
- 高质量促进分布内泛化,多样性提升分布外泛化能力
- 强调质量与多样性权衡,指导自进化算法设计
利用大语言模型生成合成数据是一种极具前景的方法,可近乎无限地扩展自然数据以应对各类任务。由于方法多样,现有研究中直接比较不同合成数据生成算法仍十分稀缺,难以明确改进来源与瓶颈所在。本文提出通过数据质量(Quality)、多样性(Diversity)和复杂性(Complexity)三个维度来评估各算法生成的合成数据构成。这三个特性在开放性任务中至关重要,且显著影响下游模型的能力。我们发现:质量对分布内泛化至关重要,多样性对分布外泛化至关重要,而复杂性则对两者均有益处。此外,我们强调训练数据中存在质量-多样性权衡,并分析了合成数据流水线中各组件对QDC特性的影响。基于此,我们构建了按组件使用及对QDC影响分类的算法谱系。该分析延伸至强化学习与自进化算法中平衡QDC的重要性。类似地,模型输出质量与多样性之间也存在权衡,影响合成数据组成。当前多数模型仅优化输出质量,限制了多样性与自进化潜力。我们认为平衡这些权衡是未来自进化算法发展的关键,并指出若干在此方向取得进展的工作。
原文摘要 · Abstract (English)
Synthetic data generation with Large Language Models is a promising paradigm for augmenting natural data over a nearly infinite range of tasks. Given this variety, direct comparisons among synthetic data generation algorithms are scarce, making it difficult to understand where improvement comes from and what bottlenecks exist. We propose to evaluate algorithms via the makeup of synthetic data generated by each algorithm in terms of data quality, diversity, and complexity. We choose these three characteristics for their significance in open-ended processes and the impact each has on the capabilities of downstream models. We find quality to be essential for in-distribution model generalization, diversity to be essential for out-of-distribution generalization, and complexity to be beneficial for both. Further, we emphasize the existence of Quality-Diversity trade-offs in training data and the downstream effects on model performance. We then examine the effect of various components in the synthetic data pipeline on each data characteristic. This examination allows us to taxonomize and compare synthetic data generation algorithms through the components they utilize and the resulting effects on data QDC composition. This analysis extends into a discussion on the importance of balancing QDC in synthetic data for efficient reinforcement learning and self-improvement algorithms. Analogous to the QD trade-offs in training data, often there exist trade-offs between model output quality and output diversity which impact the composition of synthetic data. We observe that many models are currently evaluated and optimized only for output quality, thereby limiting output diversity and the potential for self-improvement. We argue that balancing these trade-offs is essential to the development of future self-improvement algorithms and highlight a number of works making progress in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。