揭示合成数据市场中模型坍塌的经济机制,提出最优溯源补贴方案。
The Economics of Model Collapse: Equilibrium, Welfare, and Optimal Provenance Subsidies in Synthetic Data Markets
- 构建合成数据市场的微观经济模型,定义污染均衡与福利分解公式。
- 实证发现每代迭代坍塌率约18.1%,十年后模型质量提升23.1%且分布漂移下降55%。
- 给出可计算的最优溯源补贴和水印强度公式,适合政策制定者与合规研究者。
生成式人工智能正在重塑训练数据的供给端:越来越多的新文本、图像和结构化记录由前代模型生成而非人类创造。对合成内容进行递归训练会引发可度量且常不可逆的分布保真度损失,即模型坍塌。本文首次建立合成数据市场在模型坍塌下的统一微观经济理论。提出合成数据污染均衡(SDCE),证明其存在性与普遍唯一性;推导出福利分解公式 W = W_prod + W_cons - L_coll - L_info;建立基于Wasserstein梯度流的均场坍塌极限;证明信息受限下无法实现理想溯源;获得福利最大化溯源补贴 s* = KL(q||p)/(2 kappa) 与水印强度 w* = (1 - psi) KL(q||p)/(2 kappa psi) 的闭式解。证明仅基于生产方观测的溯源估计器存在信息论下界,并表明PMIR算法可逼近该下界,收敛至epsilon-SDCE需 O(epsilon^-2 log T) 次迭代。在C4合成基准上对十代重训练进行简化回归,估计坍塌率系数 b-hat = 0.181(HAC标准误0.024),与结构预测值0.183相差不足一个标准误。校准实验使第十年模型质量相比未调控基准提升23.1%,同时持留多样性探测器上的2-Wasserstein漂移从0.318降至0.142。跨世代缩放实验在 t ∈ {1,...,10} 上恢复对数坍塌律 log Q_t = log Q_0 - 0.183 t rho^2,R² = 0.962。
原文摘要 · Abstract (English)
Generative artificial intelligence is rapidly transforming the supply side of training data: an increasing share of new tokens, images, and structured records is produced by previous-generation models rather than by human originators. Recursive training on such synthetic content induces a measurable and often irreversible loss of distributional fidelity, a phenomenon known as model collapse. We develop the first unified microeconomic theory of synthetic data markets under model collapse. We introduce the Synthetic Data Contamination Equilibrium (SDCE), prove existence and generic uniqueness, derive a welfare decomposition W = W_prod + W_cons - L_coll - L_info, establish a Wasserstein-gradient-flow mean-field collapse limit, prove an impossibility of information-constrained implementation, and obtain closed-form expressions for the welfare-maximizing provenance subsidy s* = KL(q||p)/(2 kappa) and the welfare-maximizing watermark strength w* = (1 - psi) KL(q||p)/(2 kappa psi). We prove an information-theoretic Cramer-Rao lower bound on any provenance estimator using only producer-side observations and show that the Provenance-Market Iterative Retraining (PMIR) algorithm attains this bound up to constants while converging to an epsilon-SDCE in O(epsilon^-2 log T) iterations. A reduced-form OLS estimation on a C4-synthetic benchmark over ten retraining generations yields a collapse-rate coefficient b-hat = 0.181 (HAC s.e. 0.024), within one standard error of the structural prediction 0.183. Calibrated experiments raise generation-ten model quality by 23.1 percent over the unregulated benchmark while lowering the 2-Wasserstein drift on a held-out diversity probe from 0.318 to 0.142. Scaling experiments over generations t in {1,...,10} recover a logarithmic-in-t collapse law log Q_t = log Q_0 - 0.183 t rho^2 with R^2 = 0.962.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。