用少量真实数据加权筛选高质量合成文本,提升分类模型性能
Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification
- 基于真实数据加权筛选高质多样合成文本
- 在多个分类任务中超越交叉熵和现有加权方法
- 适合数据稀缺场景下高效利用合成数据
通过大语言模型(LLMs)生成合成数据增强可提升下游任务性能,尤其在真实数据稀缺时。然而,生成数据可能偏离真实分布,导致模型应用效果下降。为此,本文提出高效加权损失方法,仅用少量真实数据即可强调高质量、多样化的合成数据,实现与真实分布对齐。我们在多个文本分类任务上验证了该方法的有效性,结果表明,在BERT级模型上,该方法显著优于标准交叉熵损失及其他数据加权策略,为利用任意合适生成器的合成数据提供了有效解决方案。
原文摘要 · Abstract (English)
Synthetic data augmentation via large language models (LLMs) allows researchers to leverage additional training data, thus enhancing the performance of downstream tasks, especially when real-world data is scarce. However, the generated data can deviate from the real-world data, and this misalignment can bring deficient outcomes while applying the trained model to applications. Therefore, we proposed efficient weighted-loss approaches to align synthetic data with real-world distribution by emphasizing high-quality and diversified data generated by LLMs with using merely a little real-world data. We empirically assessed the effectiveness of our method on multiple text classification tasks, and the results showed leveraging our approaches on a BERT-level model robustly outperformed standard cross-entropy and other data weighting approaches, providing potential solutions to effectively leveraging synthetic data from any suitable data generator for model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。