arXiv:2608.26235cs.AIcs.PF2026-08中稿 · the 2026 TPCTC Con…

评估大模型推理的性价比,发现并非所有任务都值得多思考。

The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts

  • 提出新指标TES,衡量推理带来的准确率提升与token消耗的比值。
  • 发现顺序推理任务如AIME 2025 TES高,而知识回忆任务如MMLU-Pro TES低。
  • 提醒按任务类型、思考程度和部署环境决定是否开启推理,避免浪费。

仅以准确率为标准评估具备推理能力的大模型,忽略了关键的部署问题:延长思考所消耗的token是否值得?我们引入边际评估指标Token Economy Score(TES),衡量推理模型相比非推理基线的准确率提升,除以生成的token倍增系数。针对具有推理开关的模型族和无直接非推理对照的前沿模型,分别定义了配对与近似TES版本。在7个基准测试上,对151次模型-基准评估进行了实证分析,涵盖数学、代码生成、科学推理、指令遵循、专家知识、知识回忆和研究级物理等领域。分析考察了三个部署相关维度:何种任务结构能带来正向边际推理效率、模型族内增加推理投入如何影响TES、部署环境如何改变经济可行性。结果表明,任务结构比名义难度更能预测推理效率:如AIME 2025和LiveCodeBench等序列推理任务拥有高TES,而尽管困难的MMLU-Pro等知识回忆任务反而呈现低TES。我们还发现,在更高推理投入水平下存在系统性边际递减,甚至出现额外思考降低准确率的情况。此外,推理成本占比(RCS)显示内部思考常主导推理支出,而部署成本倍增器(DCM)揭示本地部署可显著改善原本昂贵的推理负载经济性。这些发现支持一种由基准驱动的模型选择规则:应根据任务类型、投入程度和部署环境选择性启用推理,而非将其视为普遍有益模式。

原文摘要 · Abstract (English)

Accuracy-only benchmarking of reasoning-capable large language models misses a central deployment question: when do extended thinking tokens earn their cost? We introduce the Token Economy Score (TES), a marginal benchmarking metric that measures the accuracy gain of a reasoning model over a non-reasoning baseline, normalized by the generated-token multiplier. We define paired and approximated TES variants for model families with reasoning toggles and frontier models without direct non-reasoning counterparts. We then conduct an empirical benchmarking analysis across 151 model-benchmark evaluation runs on seven benchmarks spanning mathematics, code generation, science reasoning, instruction following, expert knowledge, knowledge recall, and research-level physics. The analysis examines three deployment-facing dimensions: which task structures yield positive marginal reasoning efficiency, how increasing reasoning effort changes TES within model families, and how deployment context changes economic viability. Results show that task structure predicts reasoning efficiency better than nominal difficulty: sequential inferencechain tasks such as AIME 2025 and LiveCodeBench show high TES, while knowledge-recall tasks such as MMLU-Pro show low TES despite their difficulty. We also find systematic diminishing returns at higher reasoning effort levels, including cases where additional thinking reduces accuracy. Finally, Reasoning Cost Share (RCS) shows that inference spend is often dominated by internal thinking, while Deployment Cost Multiplier (DCM) shows how on-premises deployment can change the economics of otherwise costly reasoning workloads. These findings support a benchmarking-driven model-selection rule: enable reasoning selectively by task type, effort level, and deployment context rather than treating it as a universally beneficial mode.

大模型推理成本评估任务适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。