提出统一解释模型自生成崩溃的熵水库投影理论
Entropy-Reservoir Bregman Projection: An Information-Geometric Unification of Model Collapse
- 将自引用学习建模为分布空间中的贝格曼投影序列
- 证明有限样本噪声导致熵指数衰减并引发模型崩溃
- 引入熵水库可稳定系统,提供可量化的修复设计原则
自引用学习——用模型自身生成的数据训练模型——虽具无限扩展潜力,却长期受模型崩溃困扰:语言模型趋于重复文本,生成对抗网络丢失模式,强化学习策略过度利用。尽管实践中采用真实数据混合、熵奖励、知识蒸馏或检索增强生成等临时手段,但缺乏统一原理来解释失败与成功。本文提出熵水库贝格曼投影(ERBP)框架,将闭环过程建模为分布空间中的随机贝格曼投影序列。无外部耦合时,有限样本噪声迫使系统投影至不断缩小的经验支持集,导致熵指数衰减并最终崩溃。引入熵水库——在每次投影中混合高熵分布——可注入可控熵流,严格保证动态稳定性。理论给出:(i) 崩溃的必要条件,(ii) 非平凡熵下界的充分条件,(iii) 仅依赖样本量及贝格曼生成器强凸性/利普希茨常数的闭式速率。在大语言模型自训练、软演员-评论家强化学习和生成对抗网络优化上的实验验证了预测,表明不同稳定化技巧对应特定水库选择与耦合系数。因此,ERBP将零散经验法则转化为单一量化设计规则:监控并预算熵流。
原文摘要 · Abstract (English)
Self-referential learning -- training a model on data it generated itself -- promises boundless scalability but chronically suffers from model collapse: language models degenerate into repetitive text, GANs drop modes, and reinforcement-learning policies over-exploit. Although practitioners employ ad~hoc fixes such as real-data mixing, entropy bonuses, knowledge distillation, or retrieval-augmented generation, a single principle that explains both the failure mode and the success of these fixes has remained elusive. We present Entropy-Reservoir Bregman Projection (ERBP), an information-geometric framework that unifies these phenomena. We model the closed loop as a stochastic Bregman projection sequence in distribution space. Without external coupling, finite-sample noise forces the system to project onto an ever-shrinking empirical support, causing exponential entropy decay and eventual collapse. Introducing an Entropy Reservoir -- a high-entropy distribution mixed into each projection -- injects a controllable entropy flux that provably stabilises the dynamics. Our theory yields (i) a necessary condition for collapse, (ii) a sufficient condition that guarantees a non-trivial entropy floor, and (iii) closed-form rates that depend only on sample size and the strong-convexity/Lipschitz constants of the Bregman generator. Experiments on large-language-model self-training, Soft Actor-Critic in reinforcement learning, and GAN optimisation validate our predictions and show that disparate stabilisation heuristics correspond to specific reservoir choices and coupling coefficients. ERBP thus transforms a collection of folk remedies into a single, quantitative design rule: monitor and budget your entropy flux.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。