arXiv:2602.01120cs.LGcs.AI2026-02

提出马尔可夫模型优化大模型推理时的逐步扩展,实现精准平衡准确率与效率。

MarkovScale: Towards Optimal Sequential Scaling at Inference Time

  • 将逐步扩展建模为两状态马尔可夫过程,推导出理论最优条件。
  • 在3个大模型、5个基准上显著优于现有方法,提升稳定且可预测。
  • 适合追求高效推理的工业应用与资源受限场景使用。

逐步扩展是一种重要的推理时缩放范式,但其性能提升通常有限且难以理解,主要由于普遍采用启发式、非原理性方法,导致缺乏清晰的最优边界。为此,我们提出一个原理性框架,将逐步扩展建模为两状态马尔可夫过程。该方法揭示了逐步扩展的内在特性,并导出关键问题的闭式解,如准确率提升的具体条件,以及理论上的上限、中性与下限性能边界。基于此,我们开发了MarkovScale系统,应用这些最优性准则,在准确率与效率间实现理论支撑的平衡。在3个骨干大模型、5个基准和超过20种配置下的全面实验表明,MarkovScale持续优于最先进的并行与逐步扩展方法,标志着大模型推理向最优与资源高效迈出重要一步。源代码将在录用后公开。

原文摘要 · Abstract (English)

Sequential scaling is a prominent inference-time scaling paradigm, yet its performance improvements are typically modest and not well understood, largely due to the prevalence of heuristic, non-principled approaches that obscure clear optimality bounds. To address this, we propose a principled framework that models sequential scaling as a two-state Markov process. This approach reveals the underlying properties of sequential scaling and yields closed-form solutions for essential aspects, such as the specific conditions under which accuracy is improved and the theoretical upper, neutral, and lower performance bounds. Leveraging this formulation, we develop MarkovScale, a practical system that applies these optimality criteria to achieve a theoretically grounded balance between accuracy and efficiency. Comprehensive experiments across 3 backbone LLMs, 5 benchmarks, and over 20 configurations show that MarkovScale consistently outperforms state-of-the-art parallel and sequential scaling methods, representing a significant step toward optimal and resource-efficient inference in LLMs. The source code will be open upon acceptance at https://open-upon-acceptance.

大模型推理序列缩放马尔可夫模型效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。