arXiv:2604.07035cs.CL2026-04被引 2

评测7个开源推理模型在4个基准上的表现,强调部署实际考量。

Unified Deployment-Aware Evaluation of Open Reasoning Language Models

论文配图:Unified Deployment-Aware Evaluation of Open Reasoning Language Models
图 1 · 摘自论文原文
  • 统一测试7种模型配置在4个数据集上,每组用相同238样本和三种提示策略。
  • Gemma-4-26B-A4B零样本提示得分最高(加权0.794),但资源消耗高。
  • 模型排名受提示方式影响大,适合按实际需求选型而非只看单一分数。

开源推理语言模型常在不同样本量、部分标准化提示和以准确率为中心的总结下进行比较,导致实际模型选择难以解读。本文对七种开源推理模型配置在四个基准(ARC-Challenge、GSM8K、MATH 1-3、TruthfulQA MC1)上进行统一评估。所有模型在相同238例子集上测试零样本、思维链(CoT)及少量样本CoT提示,形成7×4×3共84种条件,总计评估19,992个样本。除准确率外,还报告威尔逊置信区间、延迟、峰值显存(VRAM)、加权综合性能、帕累托高效操作点、提示敏感度指标与兼容性诊断。结果显示,Gemma-4-26B-A4B在零样本提示下取得最高加权得分0.794;而Gemma-4-E4B在多种提示设置下表现接近顶尖,且延迟和内存占用显著更低,是理想的实用操作点。自举与配对置换分析表明,领先配置差距极小,部署权衡依然关键。提示策略改变会重排模型排名,而非整体提升或下降。基准间的互补性带来路由空间,全知任务感知选择器可达0.825的加权得分。兼容性诊断显示,如Phi-4-Reasoning在GSM8K上的表现不佳,实为鲁棒性与接口适配问题,非能力缺陷。结论支持:开源模型评估应视为部署导向的多目标操作点问题,而非单一分数排行榜。

原文摘要 · Abstract (English)

Open reasoning language models are often compared under mixed sample sizes, partially standardized prompts, and accuracy-centered summaries, which makes practical model selection difficult to interpret. We present a unified evaluation of seven open reasoning language model configurations across four benchmarks: ARC-Challenge, GSM8K, MATH levels 1 to 3, and TruthfulQA MC1. We test zero-shot, chain-of-thought (CoT), and few-shot CoT prompting on the same 238-example subset for every model--dataset--strategy condition, yielding a complete 7 x 4 x 3 design with 84 conditions and 19,992 evaluated examples. Beyond accuracy, we report Wilson confidence intervals, latency, peak video random access memory (VRAM), weighted aggregate performance, Pareto-efficient operating points, prompt-sensitivity metrics, and compatibility diagnostics. Gemma-4-26B-A4B with zero-shot prompting achieves the highest weighted score at 0.794. Gemma-4-E4B remains close to the top across prompting settings while using substantially lower latency and memory, making it a strong practical operating point. Bootstrap and paired-permutation analyses show that the leading configurations are close enough that deployment tradeoffs remain important. We also find that prompting strategy changes model rankings rather than shifting all models uniformly. Benchmark-specific complementarity creates routing headroom, with an oracle task-aware selector reaching a weighted score of 0.825. Compatibility diagnostics show that some apparent failures, especially Phi-4-Reasoning on GSM8K, reflect robustness and interface-adherence problems under the shared evaluation pipeline. These results support a central claim: open-model evaluation should be framed as a deployment-aware, multi-objective operating-point problem rather than as a single-score leaderboard exercise.

模型评测部署优化多目标评估推理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。