arXiv:2506.04734cs.AIcs.CL2025-06被引 3

评测设计差异导致大模型推理能力夸大,真实性能存疑。

Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design

  • 通过调整评测条件,发现模型性能波动剧烈。
  • 相同模型在不同评测环境下得分相差可达数十个百分点。
  • 提醒研究者警惕评测方法对结果的误导,适合关注模型评估可靠性的读者。

以 Deepseek-R1-Distill 系列为代表的推理模型因其在数学、科学、编程等领域的优异表现,被开源社区广泛采用。然而,我们的研究表明,其基准评测结果受多种因素影响,存在显著波动。评测条件的细微差异即可导致结果大幅变化。类似现象也出现在基于 Deepseek-R1-Distill 微调的其他开源推理模型以及 QwQ-32B 模型中,使其宣称的性能提升难以稳定复现。因此,我们呼吁建立更严格的模型性能评估范式,并提供了对 Deepseek-R1-Distill 系列模型的实证评估。

原文摘要 · Abstract (English)

Reasoning models represented by the Deepseek-R1-Distill series have been widely adopted by the open-source community due to their strong performance in mathematics, science, programming, and other domains. However, our study reveals that their benchmark evaluation results are subject to significant fluctuations caused by various factors. Subtle differences in evaluation conditions can lead to substantial variations in results. Similar phenomena are observed in other open-source inference models fine-tuned based on the Deepseek-R1-Distill series, as well as in the QwQ-32B model, making their claimed performance improvements difficult to reproduce reliably. Therefore, we advocate for the establishment of a more rigorous paradigm for model performance evaluation and present our empirical assessments of the Deepseek-R1-Distill series models.

模型评测大模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。