arXiv:2511.00763cs.AI2025-11被引 2

发现大模型在重复任务中准确率呈双指数崩溃,揭示其无法独立执行每一步操作。

How Focused Are LLMs? A Quantitative Study via Repetitive Deterministic Prediction Tasks

  • 通过重复确定性任务测试模型,发现准确率随长度呈双指数下降
  • 超过特征长度后出现准确率悬崖,表明模型无法独立执行每步操作
  • 提出统计物理模型解释注意力干扰,适合研究模型推理极限的读者

我们研究大语言模型在重复确定性预测任务中的表现,考察序列准确率随输出长度的变化规律。每个任务涉及重复相同操作n次,如按规则替换字符串字母、整数加法,以及多体量子力学中的字符串算子乘法。若模型采用简单重复算法,成功率应随序列长度呈指数衰减。然而,对主流大模型的实验显示,准确率在特征长度后呈现尖锐的双指数下降,形成准确率悬崖,标志着可靠生成向不稳定生成的转变。这表明模型无法独立执行每一步操作。为此,我们提出一个受统计物理启发的模型,捕捉提示外在条件与生成标记间内部干扰的竞争关系。该模型能定量复现观察到的交叉现象,并建立注意力诱导干扰与序列级失败之间的可解释联系。将模型拟合到多个模型和任务的实证结果,得到刻画每个模型-任务对固有错误率和错误累积因子的有效参数,为理解大语言模型确定性准确性的极限提供了原则性框架。

原文摘要 · Abstract (English)

We investigate the performance of large language models on repetitive deterministic prediction tasks and study how the sequence accuracy rate scales with output length. Each such task involves repeating the same operation n times. Examples include letter replacement in strings following a given rule, integer addition, and multiplication of string operators in many body quantum mechanics. If the model performs the task through a simple repetition algorithm, the success rate should decay exponentially with sequence length. In contrast, our experiments on leading large language models reveal a sharp double exponential drop beyond a characteristic length scale, forming an accuracy cliff that marks the transition from reliable to unstable generation. This indicates that the models fail to execute each operation independently. To explain this phenomenon, we propose a statistical physics inspired model that captures the competition between external conditioning from the prompt and internal interference among generated tokens. The model quantitatively reproduces the observed crossover and provides an interpretable link between attention induced interference and sequence level failure. Fitting the model to empirical results across multiple models and tasks yields effective parameters that characterize the intrinsic error rate and error accumulation factor for each model task pair, offering a principled framework for understanding the limits of deterministic accuracy in large language models.

大模型推理准确率分析序列生成注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。