用验证器提升数学难题生成质量,让大模型自动生成有效且有挑战性的问题。
Verifier-Backed Hard Problem Generation for Mathematical Reasoning

- 引入独立验证器约束生成过程,确保问题有效且难。
- 在不定积分和通用数学推理任务上显著优于基线方法。
- 适合想自研数学题的大模型训练或自主科研的团队使用。
大型语言模型在解决科学与数学问题方面表现出强大能力,但在生成有效、具有挑战性且新颖的问题上仍存在困难,而这是推动大模型训练和实现自主科学研究的关键。现有方法要么依赖昂贵的人类专家,要么采用简单的自对弈范式,常因奖励劫持导致生成无效问题。本文提出VHG框架,基于三方自对弈机制,将独立验证器融入传统的出题者-解题者二元结构中,使出题者的奖励由验证器评估的问题有效性与解题者评估的问题难度共同决定。我们实现了两种验证器变体:硬符号验证器和软基于LLM的验证器,并在不定积分任务与通用数学推理任务上进行评估。实验结果表明,VHG在各方面均显著优于所有基线方法。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate strong capabilities for solving scientific and mathematical problems, yet they struggle to produce valid, challenging, and novel problems - an essential component for advancing LLM training and enabling autonomous scientific research. Existing problem generation approaches either depend on expensive human expert involvement or adopt naive self-play paradigms, which frequently yield invalid problems due to reward hacking. This work introduces VHG, a verifier-enhanced hard problem generation framework built upon three-party self-play. By integrating an independent verifier into the conventional setter-solver duality, our design constrains the setter's reward to be jointly determined by problem validity (evaluated by the verifier) and difficulty (assessed by the solver). We instantiate two verifier variants: a Hard symbolic verifier and a Soft LLM-based verifier, with evaluations conducted on indefinite integral tasks and general mathematical reasoning tasks. Experimental results show that VHG substantially outperforms all baseline methods by a clear margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。