新模型更聪明不是靠多想,而是想得更准。
The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer
- 对比不同版本模型,发现性能提升不依赖更长推理链。
- 推理越长准确率反而下降,尤其在弱模型上更明显。
- 高效模型能用更少算力达成更好结果,适合效率优化研究者。
大型语言模型在数学推理方面取得显著进展,依赖思维链和强化学习。然而,推理词元使用量与准确率提升之间的关系仍不明确。我们系统分析了o1-mini与o3-mini变体在Omni-MATH基准上的推理链长度,发现o3-mini (m) 在无需更长推理链的情况下实现了更高准确率。此外,我们发现所有模型和计算设置下,推理链越长,准确率普遍下降,即使控制题目难度也是如此。这一下降幅度在更高效的模型中显著减小,表明新一代推理模型能更有效地利用测试时计算资源。最后,我们指出,虽然o3-mini (h) 相较于o3-mini (m) 有微弱准确率提升,但其通过在所有问题上大幅增加推理词元实现,包括o3-mini (m) 已能解决的问题。这些发现揭示了模型能力与推理长度的新关系,对效率、扩展性及评估方法具有启示。
原文摘要 · Abstract (English)
Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and reinforcement learning. However, many open questions remain regarding the interplay between reasoning token usage and accuracy gains. In particular, when comparing models across generations, it is unclear whether improved performance results from longer reasoning chains or more efficient reasoning. We systematically analyze reasoning chain length across o1-mini and o3-mini variants on the Omni-MATH benchmark, finding that o3-mini (m) achieves superior accuracy without requiring longer reasoning chains than o1-mini. Moreover, we show that accuracy generally declines as reasoning chains grow across all models and compute settings, even when controlling for difficulty of the questions. This accuracy drop is significantly smaller in more proficient models, suggesting that new generations of reasoning models use test-time compute more effectively. Finally, we highlight that while o3-mini (h) achieves a marginal accuracy gain over o3-mini (m), it does so by allocating substantially more reasoning tokens across all problems, even the ones that o3-mini (m) can already solve. These findings provide new insights into the relationship between model capability and reasoning length, with implications for efficiency, scaling, and evaluation methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。