提出新评估指标,精准衡量大模型推理时的计算扩展能力
ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
- 引入样本级感知机制,识别计算增加却性能下降的反向缩放行为
- 动态采样降低准确率波动与词元数不稳对评估的影响
- 实测发现Claude Opus在推理扩展性上优于其他主流模型
测试时扩展已成为提升大模型推理性能的关键范式,可在推理阶段动态分配计算资源。然而随着推理模型迅速发展,如何系统比较不同模型的测试时扩展能力仍是一个关键问题。本文提出ARISE(自适应分辨率感知扩展评估),一种专门用于评估大模型测试时扩展有效性的新指标。不同于现有方法,ARISE包含两项创新:(1) 样本级感知机制,能有效惩罚计算增加却导致性能下降的负向扩展行为;(2) 动态采样机制,减轻准确率波动和词元数不稳对最终评估结果的影响。我们在数学推理、代码生成及代理任务等多个领域对前沿推理模型进行了全面实验。结果表明,ARISE能提供可靠且细粒度的测试时扩展能力测量,揭示出不同模型间显著的扩展效率差异。值得注意的是,评估发现Claude Opus在扩展特性上优于其他当代推理模型。
原文摘要 · Abstract (English)
Test-time scaling has emerged as a transformative paradigm for enhancing the performance of large reasoning models, enabling dynamic allocation of computational resources during inference. However, as the landscape of reasoning models rapidly expands, a critical question remains: how can we systematically compare and evaluate the test-time scaling capabilities across different models? In this paper, we introduce ARISE (Adaptive Resolution-aware Scaling Evaluation), a novel metric specifically designed to assess the test-time scaling effectiveness of large reasoning models. Unlike existing evaluation approaches, ARISE incorporates two key innovations: (1) sample-level awareness that effectively penalizes negative scaling behaviors where increased computation leads to performance degradation, and (2) a dynamic sampling mechanism that mitigates the impact of accuracy fluctuations and token count instability on the final assessment. We conduct comprehensive experiments evaluating state-of-the-art reasoning models across diverse domains including mathematical reasoning, code generation, and agentic tasks. Our results demonstrate that ARISE provides a reliable and fine-grained measurement of test-time scaling capabilities, revealing significant variations in scaling efficiency across models. Notably, our evaluation identifies Claude Opus as exhibiting superior scaling characteristics compared to other contemporary reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。