arXiv:2605.07724cs.LGcs.AI2026-05

用多元奖励机制可避免生成模型在合成数据重训中崩溃。

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences

论文配图:Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
图 1 · 摘自论文原文
  • 基于多个奖励函数进行数据筛选,打破单一目标导致的输出坍缩。
  • 理论证明模型能收敛到稳定分布,保留多样性的高得分区域。
  • 适合关注生成模型稳定性与对齐机制的研究者阅读。

递归重训生成模型面临关键表征挑战:当合成输出基于固定奖励信号筛选时,模型易坍缩至少数过度优化的输出。以往研究认为,必须引入真实数据才能避免此问题。本文从对齐视角重新审视该结论,表明通过多奖励函数引导的筛选可缓解坍缩。我们形式化了异质偏好下的递归训练动态,并证明在特定条件下,模型收敛至一个稳定分布,能在多个高奖励区域间分配概率质量。该极限分布保持多样性,且严格满足加权纳什讨价还价解,为合成重训循环中的价值聚合提供了形式化解释。

原文摘要 · Abstract (English)

Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed reward signal, the model tends to collapse onto a narrow set of outputs that over-optimize that objective. Prior work suggests that such collapse is unavoidable without adding real data into the mix. We revisit this conclusion from an alignment perspective and show that collapse can be mitigated through curation based on multiple reward functions. We formalize the dynamics of recursive training under heterogeneous preferences and prove that, under certain conditions, the model converges to a stable distribution that allocates probability mass across competing high-reward regions. The limiting distribution preserves diversity and provably satisfies a weighted Nash bargaining solution, offering a formal interpretation of value aggregation in synthetic retraining loops.

生成模型合成数据对齐机制多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。