arXiv:2605.00300cs.AIcs.DC2026-05

用终端粒度评估AI推理,综合速度、成本与质量。

Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference

论文配图:Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
图 1 · 摘自论文原文
  • 从终端角度测量推理表现,覆盖速度、延迟、价格等五维指标
  • 同一模型在不同终端上准确率差12.5点,能耗差6.2倍
  • 支持按任务类型动态调整排名,适合部署决策者使用

现有推理基准多在模型和提供商层面比较AI系统,但实际部署决策的最小单元是终端:即(提供商、模型、库存单位)组合,包含特定量化方式、解码策略、区域和服务栈。我们提出TokenArena,一个连续性基准,以终端粒度测量五个核心维度(输出速度、首字延迟、工作负载融合价格、有效上下文长度、实时终端质量),并结合建模的能耗估算,合成三个核心指标:每正确答案所需焦耳数、每正确答案美元数、终端保真度(输出分布与自有参考的相似性)。该框架在方法论上具有创新性。在78个提供12个模型族服务的终端中,同一模型在不同终端上平均准确率差异达数学与代码任务12.5分,指纹相似性最高相差12分,尾部延迟相差一个数量级,建模能耗比达6.2倍。我们进一步发现,工作负载感知的融合定价会显著改变排名:在聊天设定(3:1输入输出比)下排名前10的7个终端,在检索增强设定(20:1)下跌出前十;而推理设定(1:5)则提升了被聊天设定高估价格的前沿闭源模型。我们已开源框架、数据模式、探测工具与评测工具链,并发布v1.0榜单快照,采用CC BY 4.0许可。TokenArena是一种方法论,而非单一排名;我们公开完整溯源信息与局限性,欢迎外部复现。

原文摘要 · Abstract (English)

Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the endpoint: the (provider, model, stock-keeping-unit) tuple at which a specific quantization, decoding strategy, region, and serving stack is exposed. We introduce TokenArena, a continuous benchmark that measures inference at endpoint granularity along five core axes (output speed, time to first token, workload-blended price, effective context, and quality on the live endpoint) and synthesizes them, together with a modeled energy estimate, into three headline composites: joules per correct answer, dollars per correct answer, and endpoint fidelity (output-distribution similarity to a first-party reference). The framework's novelty is empirical and methodological. Across 78 endpoints serving 12 model families, the same model on different endpoints differs in mean accuracy by up to 12.5 points on math and code, in fingerprint similarity to first party by up to 12 points, in tail latency by an order of magnitude, and in modeled joules per correct answer by a factor of 6.2. We further show that workload-aware blended pricing reorders the leaderboard substantially: 7 of 10 top-ranked endpoints under the chat preset (3:1 input:output) fall out of the top 10 under the retrieval-augmented preset (20:1), and the reasoning preset (1:5) elevates frontier closed models that the chat preset penalizes on price. We release the framework, schema, probe and eval harness, and a v1.0 leaderboard snapshot under CC BY 4.0. TokenArena is a methodology, not a single ranking; we publish full provenance and limitations and welcome external replication.

推理优化终端评估能效分析性能基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。