分离合成数据生成的两种路径,发现固定源头更高效
When Does Generating More Help? Disentangling Fixed-Source Synthesis from Source Expansion in Synthetic Data Scaling

- 固定种子问题池和教师模型,仅增加生成预算测试固定源头合成
- 低预算拟合的缩放定律可准确预测高预算性能,且大预算下扩充源头更优
- 在固定源头下,复杂合成策略不优于简单拒绝采样,适合对比生成方法
合成数据可通过源扩展(SE)和固定源合成(FSS)两条路径扩展:前者通过增加种子材料或生成器扩大来源,后者保持源不变而增加生成预算。现有研究常将两者混同,使FSS被忽视。本文固定种子问题池与教师模型,仅调整拒绝采样(RS)下的每问题响应预算,隔离出FSS。基于重复采样覆盖固定源的机制,推导出修正后的缩放律。该公式在低预算下拟合后,可准确预测所有教师-学生对在最高预算下的表现。在相同总样本量下,小预算时两者性能相当;大预算时,扩充源头优于增加响应数。在FSS内部,从已有种子生成新问题或更换合成协议,均未优于基础拒绝采样。因此FSS是有限的缩放维度,也是比较合成策略的理想控制环境。代码与数据将公开。
原文摘要 · Abstract (English)
Synthetic data can be scaled along two routes: Source Expansion (SE), which enlarges the source by adding seed materials or generators, and Fixed-Source Synthesis (FSS), which holds the source fixed and scales the generation budget. Existing scaling studies typically expand the source as the data grows, conflating SE with FSS and leaving FSS underexplored. We isolate FSS by holding the seed-question pool and teacher model fixed, varying only the per-question response budget under Rejection Sampling (RS). We adapt the rectified scaling law to FSS, deriving it from how repeated sampling covers a fixed source. Empirically, the derived form, fit on low budgets, predicts performance at the held-out highest budget for every evaluated teacher--student pair. At matched total-sample budgets, SE and FSS are comparable at small budgets; at large budgets, adding seed questions outperforms spending the same budget on more responses. Within FSS, however, neither synthesizing additional questions from the existing seeds nor varying the synthesis protocol outperforms plain RS at matched budgets. FSS is thus a bounded scaling axis and a controlled setting for comparing synthesis protocols. We will release our code and data to facilitate further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。