用数学方法生成更丰富的数据集,无需训练
VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMs
- 基于行列式点过程优化数据多样性
- 比现有方法多样性提升1.5到3倍
- 适合想提升数据质量的研究者
大型语言模型(LLMs)正被广泛用于生成下游模型的合成数据集。然而,已有研究指出这类生成数据缺乏多样性。本文提出 Voyager,一种新颖且理论严谨的无训练数据集生成方法。该方法通过迭代优化行列式点过程(Determinantal Point Processes, DPP)所定义的数学目标函数,直接提升数据多样性。其无需训练,适用于闭源模型,且具备可扩展性。我们不仅提供了理论依据,还通过全面实验验证:Voyager 在多样性上显著优于主流基线方法,提升幅度达1.5至3倍。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly being used to generate synthetic datasets for the evaluation and training of downstream models. However, prior work has noted that such generated data lacks diversity. In this paper, we propose Voyager, a novel principled approach to generate diverse datasets. Our approach is iterative and directly optimizes a mathematical quantity that optimizes the diversity of the dataset using the machinery of determinantal point processes. Furthermore, our approach is training-free, applicable to closed-source models, and scalable. In addition to providing theoretical justification for the working of our method, we also demonstrate through comprehensive experiments that Voyager significantly outperforms popular baseline approaches by providing a 1.5-3 times improvement in diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。