用经济成本评估语言模型,看谁更划算。
Cost-of-Pass: An Economic Framework for Evaluating Language Models
- 定义‘单次通过成本’,综合准确率与推理开销
- 轻量模型最便宜处理简单计算,大模型擅长知识题,推理模型解复杂题
- 发现模型进步显著,复杂任务成本每几个月减半,适合关注部署性价比的团队
AI系统的大规模应用取决于其生成的经济价值是否超过推理成本。为此,我们基于生产理论构建了一个经济驱动的框架,通过结合准确率和推理成本来评估语言模型的生产力。我们提出“成本-通过”(cost-of-pass):生成一个正确答案的预期货币成本。进而定义前沿成本-通过——在现有模型或人类专家中可达到的最低成本,以雇佣专家的成本近似估算。分析揭示:第一,轻量模型在基础量化任务中最具成本效益,大模型适用于知识密集型任务,推理模型在复杂量化问题上表现最优,尽管每词成本更高;第二,过去一年中前沿成本-通过显著下降,尤其在复杂量化任务中,成本约每几个月减半;第三,通过反事实前沿分析发现,轻量、大模型和推理模型的创新分别推动了基础量化、知识密集和复杂量化任务的进步;第四,评估常见推理优化技术(多数投票、自精炼)和预算感知方法(TALE-EP)发现,性能提升微小的方法通常不值得额外开销,而TALE-EP展现一定潜力。总体而言,互补性的模型创新是成本效率提升的主要驱动力,本框架为衡量进展和指导部署提供了原则性工具。
原文摘要 · Abstract (English)
Widespread adoption of AI systems hinges on their ability to generate economic value that outweighs their inference costs. Evaluating this tradeoff requires metrics accounting for both performance and costs. Building on production theory, we develop an economically grounded framework to evaluate language models' productivity by combining accuracy and inference cost. We formalize cost-of-pass: the expected monetary cost of generating a correct solution. We then define the frontier cost-of-pass: the minimum cost-of-pass achievable across available models or the human-expert(s), using the approx. cost of hiring an expert. Our analysis reveals distinct economic insights. First, lightweight models are most cost-effective for basic quantitative tasks, large models for knowledge-intensive ones, and reasoning models for complex quantitative problems, despite higher per-token costs. Second, tracking the frontier cost-of-pass over the past year reveals significant progress, particularly for complex quant. tasks where the cost roughly halved every few months. Third, to trace key innovations driving this progress, we examine counterfactual frontiers -- estimates of cost-efficiency without specific model classes. We find that innovations in lightweight, large, and reasoning models have been essential for pushing the frontier in basic quant., knowledge-intensive, and complex quant. tasks, respectively. Finally, we assess the cost-reductions from common inference-time techniques (majority voting and self-refinement), and a budget-aware technique (TALE-EP). We find that performance-oriented methods with marginal performance gains rarely justify the costs, while TALE-EP shows some promise. Overall, our findings underscore that complementary model-level innovations are the primary drivers of cost-efficiency and our framework provides a principled tool for measuring this progress and guiding deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。