用嵌入空间密度优化生成数据,提升小模型推理性能。
Efficient Embedding-based Synthetic Data Generation for Complex Reasoning Tasks
- 基于嵌入空间密度选择生成数据,增强多样性。
- 在多个基准上实现稳定性能提升。
- 适合资源有限但需复杂推理的小模型训练。
合成数据生成(SDG)利用大语言模型(LLMs),近年来被广泛采用以通过微调提升小型但更高效模型的性能。SDG的关键挑战在于确保生成数据的质量与多样性。本文分析了嵌入空间中生成数据的分布与多样性,发现特定邻域内样本密度与该区域预测准确率呈强相关。基于此,提出一种基于嵌入的针对性采样流程,显著提升数据多样性,并在多个基准测试中持续改善模型性能。
原文摘要 · Abstract (English)
Synthetic Data Generation (SDG), leveraging Large Language Models (LLMs), has recently been recognized and broadly adopted as an effective approach to improve the performance of smaller but more resource and compute efficient LLMs through fine-tuning. A key challenge in SDG is ensuring the quality and diversity of the generated data. In this paper, we analyze the diversity and distribution of generated data in the embedding space, and demonstrate a strong correlation between the density of examples within a specific neighborhood and the accuracy of predictions on examples drawn from that region. Building on this insight, we present a targeted pipeline for embedding-based sampling that enhances data diversity and consistently improves performance across several benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。