提出数学理论解释大模型自修正过程中的准确率变化规律。
A Probabilistic Inference Scaling Theory for LLM Self-Correction
- 用概率模型推导出多轮自修正后准确率的演化公式。
- 仅需一轮实验即可预测多轮准确率,与实际结果高度吻合。
- 为理解大模型自我纠错机制提供理论依据,适合研究者参考。
大型语言模型(LLMs)展现出通过自修正不断优化生成答案的能力,实现多轮迭代下的性能提升。然而,这种迭代过程中准确率如何演变的机制仍不明确。为此,本文提出一种概率推理理论,建模准确率变化动态,并解释多轮自修正中的性能改进现象。通过数学推导,我们得出第t轮自修正后的准确率公式:Acc_t = Upp - α^t(Upp - Acc_0),其中Acc_0为初始准确率,Upp为准确率收敛的上限,α决定收敛速率。基于该理论,仅需一次自修正实验即可估算出所有轮次的准确率曲线。在多种模型和数据集上的大量实验表明,理论预测与实际准确率曲线高度一致,验证了该理论的有效性。本工作为理解大模型自修正机制提供了理论基础,推动后续探索。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated the capability to refine their generated answers through self-correction, enabling continuous performance improvement over multiple rounds. However, the mechanisms underlying how and why accuracy evolves during this iterative process remain unexplored. To fill this gap, we propose a probabilistic theory to model the dynamics of accuracy change and explain the performance improvements observed in multi-round self-correction. Through mathematical derivation, we establish that the accuracy after the $t^{th}$ round of self-correction is given by: $Acc_t = Upp - α^t(Upp - Acc_0),$ where $Acc_0$ denotes the initial accuracy, $Upp$ represents the upper bound of accuracy convergence, and $α$ determines the rate of convergence. Based on our theory, these parameters can be calculated and the predicted accuracy curve then can be obtained through only a single round of self-correction. Extensive experiments across diverse models and datasets demonstrate that our theoretical predictions align closely with empirical accuracy curves, validating the effectiveness of the theory. Our work provides a theoretical foundation for understanding LLM self-correction, thus paving the way for further explorations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。