arXiv:2606.21972cs.AImath.HO2026-06

测试AI解数学科普题的难度适应性,发现机器表现不差于人类。

Human vs Machine Mathematical Difficulty on Project Euler: An Experimental Analysis

  • 用人类解题时间衡量难度,检验机器耗时与难度的幂律关系
  • 20个模型显示机器耗时增长慢于人类,支持机器更稳定
  • 揭示强模型成功概率随难度呈指数下降,适合评估AI数学能力

我们研究前沿AI系统在Project Euler平台上的计算数学问题中,其努力程度与成功概率如何随人类难度变化。基于MathArena基准数据集,包含50个问题、26种模型配置的3840次尝试,以网站公开的人类解题时间为难度指标。受蒂莫西·高华斯启发,我们检验生成词元成本与人类时间之间的幂律关系 $t_{\text{machine}} = a \cdot t_{\text{human}}^b$,发现25个可用拟合中20个模型满足 $b < 1$,包括最强基线模型,说明机器并未比人类更差地随难度恶化。我们还考察成功概率是否符合简单指数衰减模型 $p_{\text{success}} = e^{c t_{\text{human}}}$,通过分箱聚合数据,发现中位数分箱级 $R^2 = 0.92$,对22个覆盖最好的配置提供适度支持。采用METR方法拟合逻辑回归成功曲线,提取50%任务完成阈值 $h_{50}$;2026年4月20日快照中最强配置在最快五名人类基准下约为2.5–4.3小时,状态最先进模型的 $h_{50}$ 对数线性拟合显示约75天的翻倍周期。

原文摘要 · Abstract (English)

We study how the effort and success probability of frontier AI systems scale with human difficulty on problems from Project Euler, an online platform of computational mathematics problems. Our dataset, from the MathArena benchmark, consists of 3840 attempts across 50 problems and 26 model configurations, with problem difficulty measured by the site's public human solve times. Motivated by a proposal of Timothy Gowers, we test a power-law relation $t_{\text{machine}} = a \cdot t_{\text{human}}^b$ between generated-token cost per successful answer and human time, and find $b < 1$ for 20 of the 25 models with usable fits, including the strongest base models; this operationalization therefore does not support an earlier prediction that machines scale worse than humans with difficulty. We also investigate whether success probability on the tested problems can be modeled by a simple exponential decay $p_{\text{success}} = e^{c t_{\text{human}}}$, predicting a linear relation between $\log p_{\text{success}}$ and $t_{\text{human}}$. Using a binning approach for data aggregation we find moderate empirical support (median bin-level $R^2 = 0.92$ across the 22 best-covered configurations) for this model. Following METR, we also fit logistic success curves and extract 50\% task-length horizons $h_{50}$; the strongest configurations in our 20 April 2026 snapshot reach roughly $2.5$--$4.3$ hours on our fastest-five human baseline, with a log-linear fit through the state-of-the-art frontier giving a descriptive doubling time of about $75$~days for the SOTA $h_{50}$.

数学推理机器评估难度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。