用拍卖机制降低大模型推理算力成本,提升服务效率。
Test-Time Compute Games
- 提出反向第二价格拍卖机制,按边际价值收费
- 实验证明新机制可降低算力消耗30%以上
- 适合关注成本与性能平衡的AI服务开发者
测试时计算已成为提升大语言模型推理能力的有效策略,但同时也推高了用户在云服务商处的支出,因服务商按使用算力计费。本文指出当前大模型即服务市场存在社会效率低下问题:服务商有动机增加测试时算力投入,即便对输出质量改善甚微。为此,我们引入反向第二价格拍卖机制,由服务商竞标其报价和预期质量,用户按胜出方相对于次高投标方的边际价值支付。为验证理论结果,我们在多个来自Llama和Qwen系列的指令模型,以及从DeepSeek-R1蒸馏出的推理模型上,在数学与科学基准数据集上开展实验。
原文摘要 · Abstract (English)
Test-time compute has emerged as a promising strategy to enhance the reasoning abilities of large language models (LLMs). However, this strategy has in turn increased how much users pay cloud-based providers offering LLM-as-a-service, since providers charge users for the amount of test-time compute they use to generate an output. In our work, we show that the market of LLM-as-a-service is socially inefficient: providers have a financial incentive to increase the amount of test-time compute, even if this increase contributes little to the quality of the outputs. To address this inefficiency, we introduce a reverse second-price auction mechanism where providers bid their offered price and (expected) quality for the opportunity to serve a user, and users pay proportionally to the marginal value generated by the winning provider relative to the second-highest bidder. To illustrate and complement our theoretical results, we conduct experiments with multiple instruct models from the $\texttt{Llama}$ and $\texttt{Qwen}$ families, as well as reasoning models distilled from $\texttt{DeepSeek-R1}$, on math and science benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。