arXiv:2605.09292cs.AIcs.CY2026-05被引 1

模型答案正确但解题思路单一,揭示了推理多样性的重要性。

Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning

论文配图:Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning
图 1 · 摘自论文原文
  • 构建策略级评估框架,通过双AI标注+人工仲裁识别解题路径
  • 四款前沿模型在多解提示下仅恢复71%人类参考策略,几何与数论差距最大
  • 发现50种新有效解法,说明模型有创新潜力但覆盖不全

大型语言模型在数学推理基准上已达到高准确率,但准确率无法反映推理灵活性。我们在80道AMC 10/12和AIME题目上构建策略级评估框架,基于217个来自AoPS的参考解法族进行标注。通过双AI编码加人工仲裁,对模型输出的解法类型、有效性与正确性进行评估。四款前沿模型在单解提示下准确率达95%-100%,但在多解提示下恢复的策略远少于人类参考集:Gemini、DeepSeek、GPT和Claude分别生成184、152、151、110种有效策略,几何与数论领域差距显著。模型共发现50种基准中未见的有效新解法,表明其既存在策略覆盖不足,也具备一定创造性推理能力。20道题重复实验显示策略发现呈边际递减,最强模型三次运行后仅恢复55种参考策略中的39种(71%)。结果表明,策略多样性应作为超越答案正确的数学推理评估新维度。

原文摘要 · Abstract (English)

Large language models now achieve high final-answer accuracy on mathematical reasoning benchmarks, but accuracy alone does not capture reasoning flexibility. We introduce a strategy-level evaluation framework instantiated on 80 AMC 10/12 and AIME problems with 217 AoPS-derived reference strategy families. Model outputs are annotated for strategy identity, validity, and correctness using dual-AI coding with human adjudication. Across four frontier models, we find a pronounced decoupling between answer accuracy and strategy diversity. Under a single-solution prompt, all models achieve high accuracy (95%-100%), but under a multiple-strategy prompt they recover substantially fewer strategies than the human reference set. Gemini, DeepSeek, GPT, and Claude generate 184, 152, 151, and 110 distinct valid strategies, respectively, with the largest gaps in Geometry and Number Theory. The models collectively produce 50 benchmark-novel valid strategies, indicating both incomplete coverage of human strategies and some capacity for alternative reasoning. A repeated-run robustness check on 20 problems shows diminishing gains in discovered strategies, with the strongest model recovering only 39 of 55 AoPS-reference strategies (71%) after three runs. These findings position strategy diversity as a complementary dimension for evaluating mathematical reasoning beyond answer correctness.

数学推理策略多样性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。