用GPT-4生成数据,低成本提升对话语义分析效果
Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis
- 混合使用人工标注与GPT-4生成数据
- 预算降低时,生成数据占比应提高
- 在多种预算下均优于纯人工或纯生成数据
近期研究显示,大模型少样本学习可低成本生成监督模型训练数据。然而,大模型生成数据的质量可能低于人工标注数据。这引出关键问题:如何权衡高质量但昂贵的人工数据与低质量但廉价的大模型生成数据?本文利用GPT-4合成对话语义框架分析的训练数据,并考察不同预算下的最优分配策略。实验在多种预算水平下进行,结果表明,在广泛预算范围内,混合使用人工与大模型生成数据可实现最佳成本效益。值得注意的是,随着预算下降,更高比例的生成数据更优。
原文摘要 · Abstract (English)
Recent studies have demonstrated that few-shot learning allows LLMs to generate training data for supervised models at a low cost. However, the quality of LLM-generated data may not entirely match that of human-labeled data. This raises a crucial question: how should one balance the trade-off between the higher quality but more expensive human data and the lower quality yet substantially cheaper LLM-generated data? In this paper, we synthesized training data for conversational semantic frame analysis using GPT-4 and examined how to allocate budgets optimally to achieve the best performance. Our experiments, conducted across various budget levels, reveal that optimal cost-efficiency is achieved by combining both human and LLM-generated data across a wide range of budget levels. Notably, as the budget decreases, a higher proportion of LLM-generated data becomes more preferable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。