用经济学方法优化大模型推理资源分配,提升效率与准确率
The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs

- 基于经济均衡原理设计资源分配策略,动态调整推理预算
- 在资源受限时,相比均匀分配,全局准确率最高提升3倍
- 适合部署大模型且计算资源紧张的生产环境使用
推理阶段的扩展已成为提升大语言模型性能的关键途径,但实际部署受限于严格的计算预算。本文将推理预算分配建模为受经济原则支配的全局约束优化问题,通过引入偏移脉冲函数刻画单次查询的推理效用,推导出基于全局影子价格的最优分配策略,实现资源稀缺下的边际效用均衡。据此提出受限潜伏效用均衡推理分配方法(CLEAR),能够理性放弃无法求解的查询,并将资源重新分配给接近求解阈值的可解查询。在多个推理任务和不同流量场景下的大量实验表明,CLEAR显著提升了总令牌成本与平均准确率之间的帕累托前沿。在资源极度匮乏的情况下,相较均匀分配,全局准确率最高提升3倍。
原文摘要 · Abstract (English)
Inference-time scaling has emerged as a critical avenue for enhancing Large Language Models' performance, yet real-world deployment is constrained by strict computational budgets. In this work, we formulate inference budget allocation as a global constrained optimization problem governed by economic principles. By modeling per-query reasoning utility with a shifted-surge function, we derive an optimal allocation policy based on a global shadow price that equilibrates marginal utility under resource scarcity. Based on this theory, we propose Constrained Latent-utility Equilibrium Allocation for Reasoning (CLEAR). It performs rational abandonment and reallocates resources from insolvent queries to solvable queries near their emergence thresholds. Extensive experiments on several reasoning tasks with different traffic streams demonstrate that CLEAR significantly improves the Pareto frontier of total token cost versus mean accuracy. In resource-scarce regimes, CLEAR achieves up to a 3x improvement in global accuracy compared to uniform allocation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。