arXiv:2601.22513cs.AI2026-01被引 2

首次为自奖励语言模型提供理论保障,解释其为何能克服差初始状态。

Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models

  • 通过理论分析揭示迭代更新的收敛机制
  • 性能提升速率达1/√n,且初始模型影响随迭代指数衰减
  • 适用于理解模型自我对齐原理的研究者与实践者

自奖励语言模型(SRLMs)在无外部反馈的情况下实现持续对齐优化,取得显著成果。然而,其核心机制尚缺乏理论阐释,存在关键空白。本文首次为SRLMs提供严格的理论保证:首先建立单步更新的下界,揭示其对初始模型质量的高度依赖;进而推导出完整迭代范式的有限样本误差界,表明性能以$ ilde{/mathcal{O}}(1/ ext{√}n)$速率随样本量$n$提升;关键发现是初始模型的影响随迭代次数$T$呈指数衰减。这从理论上解释了自奖励为何有效——它能稳健地克服不良初始化,引导模型动态趋向内部稳定与一致性。最后,将理论框架应用于线性Softmax模型类,得到与实际架构对应的定制化保证。

原文摘要 · Abstract (English)

Self-Rewarding Language Models (SRLMs) achieve notable success in iteratively improving alignment without external feedback. Yet, despite their striking empirical progress, the core mechanisms driving their capabilities remain unelucidated, leaving a critical gap in theoretical understanding. This paper provides the first rigorous theoretical guarantees for SRLMs. We first establish a lower bound that characterizes the fundamental limits of a single update step, revealing a critical dependence on the quality of the initial model. We then derive finite-sample error bounds for the full iterative paradigm, showing that performance improves at a rate of $\widetilde{\mathcal{O}}\left(1/\sqrt{n}\right)$ with sample size $n$. Crucially, our analysis reveals that the dependence on the initial model decays exponentially with the number of iterations $T$. This provides a formal explanation for why self-rewarding succeeds: it robustly overcomes poor initialization by steering the dynamics toward internal stability and consistency. Finally, we instantiate our theoretical framework for the linear softmax model class, yielding tailored guarantees that connect our high-level insights to practical model architectures.

语言模型自奖励理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。