揭示自回归生成中自我纠错盲点的数学根源,提出可量化预测的理论框架。
Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation
- 基于谱代数理论,定义错误传播算子并给出盲点出现的充要条件。
- 推导出纠正标记激活阈值,与实验中89.3%的纠错率提升一致。
- 统一文本、图像、视频生成的自纠正机制,适用于多种模型架构。
大规模自回归语言模型存在自我纠错盲点:当错误被归因于外部来源时能可靠修正,但对自己的输出却无法纠正。以往研究通过受控错误注入、误差深度分解、基于强化学习的验证-修正训练及内在自验证等手段实证了该现象,但缺乏对生成过程抑制错误检测机制的正式建模、纠正标记的定量激活条件以及强化学习自纠正的收敛保证。本文提出SPARC——一种自回归生成中自我纠错的谱代数理论。定义错误传播算子为残差流上每步注意力雅可比矩阵的乘积,证明盲点出现当且仅当该算子的谱半径不小于1。推导出纠正标记需超越的精确激活阈值,其函数形式与简单“等待”标记实现的89.3%盲点降低结果相符。进一步证明,强化学习验证-修正训练收敛速率与耦合强度平方成正比、样本数平方根成反比,当且仅当验证-修正耦合矩阵的谱范数小于1时成立;该条件在残差流自回归模型中保持不变,统一了文本大模型与自回归图像、视频生成。四个骨干网络及视觉自回归探针的实验验证了所有定理,谱预测与实测盲点率误差低于3.2% RMSE。
原文摘要 · Abstract (English)
Large autoregressive language models exhibit a self-correction blind spot: they reliably fix identical errors when attributed to an external source yet fail to fix the same errors in their own outputs. Prior work has documented this phenomenon empirically, through controlled error injection, error-depth decompositions, RL-based verifier-corrector training, and intrinsic self-verification, but offers no formal model of why generating a token suppresses the ability to detect its error, no quantitative activation condition for correction markers, and no convergence guarantee for reinforcement-learning-based self-correction. We close these gaps with SPARC, a spectral-algebraic theory of self-correction in autoregressive generation. We define the error-propagation operator as the product of per-step attention Jacobians on the residual stream and prove that the blind spot arises if and only if the spectral radius of this operator is at least one. We derive a sharp activation threshold, given as a function of the spectral radius, that a correction marker must exceed, recovering the 89.3\% blind-spot reduction observed with a simple ``Wait'' marker. We further prove that RL-based verifier-corrector training converges at a rate proportional to the squared coupling strength over the square root of the number of samples if and only if the verifier-corrector coupling matrix has spectral norm below one, and that this criterion is invariant across residual-stream autoregressive modalities, unifying text LLMs and autoregressive image and video generation. Experiments across four backbones and a visual autoregressive probe validate every theorem, with spectral predictions matching measured blind-spot rates within 3.2\% RMSE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。