用刻意练习理念提升合成数据生成效率,少做多好。
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
- 动态生成有挑战性的合成样本,聚焦关键信息。
- 图像分类任务上样本量减少3.4至8倍,迭代次数减半以上。
- 适合追求高效训练的视觉模型研究者参考。
受人类学习中刻意练习原则的启发,我们提出一种名为刻意练习(DP)的新型合成数据生成框架,通过动态生成提升样本效率。以往研究表明,单纯增加合成数据会因收益递减而难以持续提升性能。已有工作指出剪枝是改善扩展性的重要机制,能帮助模型聚焦于最具信息量的样本。与先生成大量数据再剪枝不同,DP直接高效地生成高价值样本。理论上证明,训练于有挑战性的信息样本可改善缩放规律;实证验证表明,DP在更少样本和迭代次数下实现更优性能:在ImageNet-100上生成样本减少3.4倍,迭代次数减少6倍;在ImageNet-1k上样本减少8倍,迭代次数降低30%,且性能优于先前方法。
原文摘要 · Abstract (English)
Inspired by the principle of deliberate practice in human learning, we propose Deliberate Practice for Synthetic Data Generation (DP), a novel framework that improves sample efficiency through dynamic synthetic data generation. Prior work has shown that scaling synthetic data is inherently challenging, as naively adding new data leads to diminishing returns. To address this, pruning has been identified as a key mechanism for improving scaling, enabling models to focus on the most informative synthetic samples. Rather than generating a large dataset and pruning it afterward, DP efficiently approximates the direct generation of informative samples. We theoretically show how training on challenging, informative examples improves scaling laws and empirically validate that DP achieves better scaling performance with significantly fewer training samples and iterations. On ImageNet-100, DP generates 3.4x fewer samples and requires six times fewer iterations, while on ImageNet-1k, it generates 8x fewer samples with a 30 percent reduction in iterations, all while achieving superior performance compared to prior work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。