AI推理成本年降5到10倍,但大模型涨价快于效率提升。
The Price of Progress: Price Performance and the Future of AI
- 通过分析历史价格数据,量化了前沿模型在知识、推理等任务上的成本下降速度。
- 算法效率每年提升约3倍,但大模型和复杂任务使运行成本每年上涨3至18倍。
- 建议评测时公开并考虑成本,以真实反映AI落地价值,适合关注算力经济的研究者。
近年来,语言模型在高级基准测试中取得巨大进展,但这些进步大多依赖更昂贵的模型。因此,基准测试可能扭曲了每美元所能实现的实际能力进展。为解决此问题,我们整合了Artificial Analysis和Epoch AI的数据,构建了迄今最大规模的当前与历史运行基准测试的价格数据集。研究发现,前沿模型在知识、推理、数学和软件工程等基准上的性能单位成本每年下降5至10倍,这主要得益于经济规律、硬件效率和算法效率的提升。控制竞争因素后,仅算法效率每年提升约3倍。然而,由于模型变大和推理需求增加,运行前沿模型的成本每年上升3至18倍。最后,我们建议评估者公开并考量基准测试的成本,作为衡量AI实际影响的关键指标。
原文摘要 · Abstract (English)
Language models have seen enormous progress on advanced benchmarks in recent years, but much of this progress has only been possible by using more costly models. Benchmarks may therefore present a warped picture of progress in practical capabilities *per dollar*. To remedy this, we use data from Artificial Analysis and Epoch AI to form the largest dataset of current and historical prices to run benchmarks to date. We find that the price for a given level of benchmark performance has decreased remarkably fast, around $5\times$ to $10\times$ per year, for frontier models on knowledge, reasoning, math, and software engineering benchmarks. These reductions in the cost of AI inference are due to economic forces, hardware efficiency improvements, and algorithmic efficiency improvements. Isolating out open models to control for competition effects and dividing by hardware price declines, we estimate that algorithmic efficiency progress is around $3\times$ per year. However, at the same time, the price of running frontier models is rising between $3\times$ to $18\times$ per year due to bigger models and larger reasoning demands. Finally, we recommend that evaluators both publicize and take into account the price of benchmarking as an essential part of measuring the real-world impact of AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。