提出方法级多样性新度量,揭示大模型解题策略差异比表面表达更关键
Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning

- 区分表面多样性和解题策略多样性,引入方法级多样性评估
- 现有度量无法反映真实策略差异,导致强化学习优化方向偏离
- 策略多样性提升推理性能,但直接优化仍会受评分模型偏好干扰
大模型数学推理中的多样性对探索至关重要,但现有度量大多仅捕捉表面形式变化,而非解题策略差异。本文提出方法级多样性:同一问题不同正确解法间的策略差异。通过人工校准的LLM评判框架,发现先前度量无法可靠反映方法级多样性,且该偏差在多样性感知的强化学习中持续存在——目标度量被保持,但方法级多样性反而下降。进一步研究发现,方法多样性的候选解集能提升测试时缩放性能;然而,在训练中直接优化评分手册多样性奖励,会使策略利用评判模型的特定偏好,而非真正拓展解题路径。因此,直接诱导方法级多样性仍是开放问题。本工作首次定义方法级多样性,揭示了表面与策略层面信号间的系统性偏差,推动大模型向更类人、真正的多样化推理迈进。
原文摘要 · Abstract (English)
Diversity in LLM mathematical reasoning is critical for exploration, but common diversity metrics mostly capture surface-level variation rather than differences in how a problem is solved. We address this gap by introducing approach-level diversity: variation in strategies across correct solutions to the same problem. Using a human-calibrated LLM judge framework, we show that prior diversity measures are unreliable proxies for approach-level diversity, and this mismatch carries over to diversity-aware RLVR, where target metrics are preserved while approach-level diversity declines. Investigating when approach-level diversity helps and whether it can be directly induced, we find that approach-diverse candidate sets improve test-time scaling. However, optimizing an LLM judge diversity reward during training causes the policy to exploit judge-specific preferences rather than broaden its approaches, leaving direct optimization of approach-level diversity as an open problem. Together, our work introduces the notion of approach-level diversity and uncovers a systematic divergence between surface- and approach-level signals, marking a step toward LLMs that reason in genuinely diverse, human-like ways.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。