arXiv:2509.22480cs.CLcs.AI2025-09被引 3

发现大模型解题分歧越大,解决能力越强,提出新评估指标提升训练效果。

Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving

  • 用解题方案的多样性作为新评估指标
  • 在三个领域中提升解题成功率,验证有效性
  • 适合改进微调与强化学习的训练策略

大语言模型(LLMs)广泛应用于问题求解任务。现有工作主要通过有标签数据的监督微调(SFT)或基于任务反馈的强化学习(RL)提升性能。本文从新视角出发,研究单个问题下模型生成解法的分歧程度。实验表明,更高的解法分歧与更强的问题求解能力正相关。基于此发现,我们提出将解法分歧作为新型评估指标,可支持SFT与RL策略。我们在三个典型问题领域验证该方法,结果一致显示使用解法分歧能提升成功解题率。这些结果表明,解法分歧是推动大模型训练与评估的简单而有效的工具。

原文摘要 · Abstract (English)

Large language models (LLMs) have been widely used for problem-solving tasks. Most recent work improves their performance through supervised fine-tuning (SFT) with labeled data or reinforcement learning (RL) from task feedback. In this paper, we study a new perspective: the divergence in solutions generated by LLMs for a single problem. We show that higher solution divergence is positively related to better problem-solving abilities across various models. Based on this finding, we propose solution divergence as a novel metric that can support both SFT and RL strategies. We test this idea on three representative problem domains and find that using solution divergence consistently improves success rates. These results suggest that solution divergence is a simple but effective tool for advancing LLM training and evaluation.

大模型解题能力评估指标训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。