合成数据泛滥下,传统学习方法失效,需新算法应对。
Learning from Synthetic Data: Limitations of ERM
- 用非均匀加权策略处理不同生成批次的数据,优于传统经验风险最小化。
- 在任意混合比例的自然与合成数据中,经典方法无法收敛到正确概念。
- 提出可学习任意类别且抗污染的新算法,适合高可靠性场景使用。
大型语言模型的普及导致合成内容激增,从评论网站到法庭文件,大量看似自然的数据实则由模型生成。本文重新审视这一普遍现象下的基础学习理论问题:输入为自然与合成数据混合,学习算法对数据来源一无所知。研究发现,在估计任意d维分布均值时,尽管经验风险最小化(ERM)能收敛至真实均值,但仍被一种对不同生成批次数据赋予非均匀权重的算法超越。在泛化一致性学习(PAC)框架下,差距更为显著:ERM并不总能收敛至真实概念,呼应了模型坍塌的研究。但本文证明,存在算法可在任意VC类和任意污染比例下学习出正确假设。
原文摘要 · Abstract (English)
The prevalence and low cost of LLMs have led to a rise of synthetic content. From review sites to court documents, "natural" content has been contaminated by data points that appear similar to natural data, but are in fact LLM-generated. In this work we revisit fundamental learning theory questions in this, now ubiquitous, setting. We model this scenario as a sequence of learning tasks where the input is a mix of natural and synthetic data, and the learning algorithms are oblivious to the origin of any individual example. We study the possibilities and limitations of ERM in this setting. For the problem of estimating the mean of an arbitrary $d$-dimensional distribution, we find that while ERM converges to the true mean, it is outperformed by an algorithm that assigns non-uniform weights to examples from different generations of data. For the PAC learning setting, the disparity is even more stark. We find that ERM does not always converge to the true concept, echoing the model collapse literature. However, we show there are algorithms capable of learning the correct hypothesis for arbitrary VC classes and arbitrary amounts of contamination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。