arXiv:2505.24187cs.CL2025-05被引 12

突破传统认知,发现大模型错误集中在少数关键词上。

Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models

  • 错误并非均匀分布,而是集中于5-10%的关键语义节点。
  • 长文本输出可靠性由少数决策点决定,而非逐字累积误差。
  • 适合关注高效推理、模型优化与长文本生成的研究者。

主流观点认为大语言模型(LLM)的可靠性随序列长度呈指数衰减,基于每标记独立出错概率的假设。本研究通过整合新证据,挑战这一观点:LLM错误并非均匀分布,而是集中在少数“关键标记”(占总标记数的5-10%),这些标记代表关键语义决策点。通过区分高影响标记与日益可预测的多数标记,我们提出新的可靠性公式,解释了现代LLM在数千标记下的持续连贯性。多条研究线索表明,长上下文表现主要依赖准确把握少数关键语义决策点,而非整体标记级精度,从而支持针对性策略,显著优于盲目扩展方法。因此,我们提出下一代系统框架,聚焦于选择性保留语义关键标记、在不确定决策边界动态分配计算资源、在歧义处进行多路径探索,并设计契合自然语义域的架构。这标志着从单纯扩展转向战略推理的根本转变,有望在不显著增加计算成本的情况下实现突破性性能,超越指数衰减假说,为更强大高效的语言系统开辟新路径。

原文摘要 · Abstract (English)

The prevailing assumption of an exponential decay in large language model (LLM) reliability with sequence length, predicated on independent per-token error probabilities, posits an inherent limitation for long autoregressive outputs. Our research fundamentally challenges this view by synthesizing emerging evidence that LLM errors are not uniformly distributed but are concentrated at sparse "key tokens" ($5-10\%$ of total tokens) representing critical decision junctions. By distinguishing these high-impact tokens from the increasingly predictable majority, we introduce a new reliability formula explaining the sustained coherence of modern LLMs over thousands of tokens. Converging research streams reveal that long-context performance primarily depends on accurately navigating a few crucial semantic decision points rather than on uniform token-level accuracy, enabling targeted strategies that significantly outperform brute-force approaches. We thus propose a framework for next-generation systems centered on selective preservation of semantically vital tokens, dynamic computational allocation at uncertain decision boundaries, multi-path exploration at ambiguities, and architectures aligned with natural semantic domains. This marks a fundamental shift from raw scaling to strategic reasoning, promising breakthrough performance without proportionate computational scaling and offering a more nuanced understanding that supersedes the exponential decay hypothesis, thereby opening pathways toward substantially more powerful and efficient language systems.

大模型错误分布长文本生成推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。