少量人工数据可显著提升合成数据模型表现,性价比远超纯合成数据。
A Little Human Data Goes A Long Way
- 用逐步替换法测试合成数据对模型影响,发现90%替换仍保持性能。
- 仅需125条人工数据即可显著提升纯合成数据训练的模型效果。
- 200条人工数据的增益需10倍合成数据才能达成,成本更低。
面对昂贵的人工标注成本,自然语言处理系统开发者越来越多依赖合成数据生成。尽管该方法前景可观,但合成数据替代人工标注的程度仍不明确。本文通过在八种不同数据集上,逐步将人类生成数据替换为合成数据,研究其在事实验证(FV)与问答(QA)任务中的影响。令人意外的是,替换高达90%的训练数据仅导致性能轻微下降,但最后10%的替换却引发严重性能滑坡。我们发现,仅需125条人工数据即可显著改进纯合成数据训练的模型。此外,仅200条额外人工数据带来的性能提升,需要约10倍数量的合成数据才能实现,且估算出此时人工标注更具成本效益。结果表明,即使无法大规模人工标注,保留小比例人工数据仍具有极高价值。
原文摘要 · Abstract (English)
Faced with an expensive human annotation process, creators of NLP systems increasingly turn to synthetic data generation. While this method shows promise, the extent to which synthetic data can replace human annotation is poorly understood. We investigate the use of synthetic data in Fact Verification (FV) and Question Answering (QA) by studying the effects of incrementally replacing human generated data with synthetic points on eight diverse datasets. Strikingly, replacing up to 90% of the training data only marginally decreases performance, but replacing the final 10% leads to severe declines. We find that models trained on purely synthetic data can be reliably improved by including as few as 125 human generated data points. We show that matching the performance gain of just a little additional human data (only 200 points) requires an order of magnitude more synthetic data and estimate price ratios at which human annotation would be a more cost-effective solution. Our results suggest that even when human annotation at scale is infeasible, there is great value to having a small proportion of the dataset being human generated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。